Spatial-TTT: An Open-Source Spatial Intelligence Model at 2B Parameters
Tsinghua's open-source Spatial-TTT makes ECCV 2026: at just 2B parameters it beats GPT-5 and Gemini-3-pro on multiple spatial intelligence benchmarks, handling 120-minute streaming video while updating its spatial memory as it watches.


Spatial-TTT: An Open-Source Spatial Intelligence Model at 2B Parameters
Tsinghua's open-source Spatial-TTT makes ECCV 2026: at just 2B parameters it beats GPT-5 and Gemini-3-pro on multiple spatial intelligence benchmarks, handling 120-minute streaming video while updating its spatial memory as it watches.
In real settings like robotics, autonomous driving, and AR, spatial understanding has never been solvable with "one glance at an image." The camera moves, the viewpoint shifts, targets flicker in and out of view — spatial information is scattered across long video streams, and a model must not only "see" but "remember, connect, and keep updating." This is a key threshold on multimodal models' road to the real world.
Spatial-TTT, a spatial intelligence model with Tsinghua University PhD student Fangfu Liu as first author, has been officially accepted by the top computer vision conference ECCV 2026. Its answer in one sentence: the model shouldn't just watch video — it should watch, update, and "grow" a spatial memory as it goes.
With only 2B parameters, it beats closed-source models like GPT-5 and Gemini-3-pro on multiple specialized spatial intelligence benchmarks tested in the paper, handles streaming video up to 120 minutes long, and is already open-sourced.
What Spatial-TTT Is
Think of it as working memory for multimodal models: a traditional VLM (vision-language model) looks at one frame or a fixed-length clip at a time, and spatial information either gets crammed into context or discarded outright. Spatial-TTT's core is continuously writing to and updating an external spatial memory module during inference, so the model's spatial understanding accumulates with viewing time instead of relying on an ever-inflating context window.
This is closer to how humans understand space — you don't take in an entire room in one shot; you build a stable spatial memory gradually through moving, observing, forgetting, and correcting.

Spatial-TTT's core idea at ECCV 2026: update spatial memory while watching streaming video.
Core Capabilities
1. Streaming 120-Minute Video
Instead of cramming every frame into context, a test-time spatial memory update mechanism "digests" up to two hours of video into a compact, queryable spatial state. Long videos no longer get brutally truncated, and VRAM doesn't blow up.
2. 2B Parameters Beating Closed-Source Giants
On multiple specialized benchmarks including SpatialBench and video spatial reasoning, the 2B Spatial-TTT outperforms closed-source models like GPT-5 and Gemini-3-pro that carry dozens of times its parameter count. It's a rare open-source lead in the "spatial intelligence" niche.
3. Continuous Updates, Not Static Snapshots
The real selling point isn't strong single-shot inference — it's that the model keeps up as the world changes. When objects in the same scene are moved, occluded, or reappear, Spatial-TTT's spatial memory corrects dynamically instead of sticking to the first snapshot it saw.
Use Cases
Clear target users:
- Robotics and embodied AI teams: as a robot's spatial perception module, it handles long-horizon visual input while maintaining a stable spatial memory — lighter on VRAM and steadier than a general-purpose VLM
- Autonomous driving perception engineers: process long-horizon driving footage for scene-level spatial understanding and target tracking
- AR / VR app developers: spatial anchoring and semantic mapping under continuously shifting camera viewpoints
- Spatial intelligence researchers: a strong baseline to reproduce and improve; paper and code are open-sourced
A clear-eyed caveat: Spatial-TTT is currently a specialized capability model, not a general VLM replacement — its strength is spatial understanding; its conversational text ability is unremarkable, so it's a poor choice for a general chatbot.
Practical Tips
- As a long-video spatial reasoning baseline: benchmark it against GPT-5 and Gemini-3-pro on your own datasets, focusing on 120-minute-class long streaming scenarios — where its lead is most obvious
- As a robot perception module: when combining with a VLA framework, treat spatial memory as a queryable state interface — it saves tokens over directly concatenating image embeddings
- Mind the input assumptions: it assumes continuous, timestamped video streams; its edge shrinks on image collections (no temporal order)
💡 Tip: The ECCV 2026 paper and open-source code will keep updating; keep an eye on the SpatialBench benchmark itself — it is becoming the "de facto standard" evaluation set for spatial intelligence, and getting familiar with it early helps you benchmark your own model against the field.
Related articles

Google Gemini Live Translate: Listen and Translate Across 70+ Languages
Google launches Gemini 3.5 Live Translate for real-time speech translation across 70+ languages, preserving your pace and tone with only seconds of delay, now live in Google Translate and Meet.

Claude Fable 5: A Hands-On Guide to Anthropic's Strongest Model
Anthropic ships Claude Fable 5 and Mythos 5 in dual editions — 80.3% on SWE-bench Pro, API pricing at $10 per million input tokens, free for a limited time until June 22.

OpenAI's Official Codex Workflow Guide: From Screenshots to Web Pages to AI-Run Research
OpenAI updates a dozen-plus official Codex real-world workflow cases covering Computer Use, /goal long-horizon objectives, PPT generation, game development, and other practical scenarios — a step-by-step guide to using Codex efficiently.

7 Field-Tested Lessons for Writing Great Claude Skills
Skill-writing lessons straight from Anthropic: trim the context, accumulate a pitfall list, script the stable steps — double your AI collaboration efficiency.

How to Run Claude Code Inside Codex at the Same Time
One setup to run Codex and Claude Code side by side: GPT plans on the left, Claude works on the right, each serving as the other's fallback — their refusal boundaries don't get in each other's way.

Kimi K2.7 Code Released: Token Consumption Down 30%
Moonshot AI has released and open-sourced the Kimi K2.7 Code coding model: 1.1 trillion parameters, 256K context, a big fix for overthinking on long-horizon tasks, plus a Speed variant with 6x the speed at 2x the price.