LingBot-World 2.0: An Open-Source, Playable Interactive World Model with Effectively Unlimited Duration
Ant Group's Robbyant has open-sourced LingBot-World-Infinity: hour-level real-time generation at 720p/60fps, with a built-in agent proposing events automatically — turning watching video into entering a world.


LingBot-World 2.0: An Open-Source, Playable Interactive World Model with Effectively Unlimited Duration
Ant Group's Robbyant has open-sourced LingBot-World-Infinity: hour-level real-time generation at 720p/60fps, with a built-in agent proposing events automatically — turning watching video into entering a world.
Previous-generation world models were like The Truman Show — they could only play back, not react. The moment a user strayed off script, the model fell apart. In July 2026, Robbyant (under Ant Group) open-sourced LingBot-World 2.0 (also known as LingBot-World-Infinity), pushing world models from "watching a video" to "entering a world." It is currently the only open-source model to achieve hour-level, near-unlimited generation duration in the general domain, combining long-duration continuous generation, high dynamism, semantic interaction, solid visual quality, real-time performance, and full open source in one system. If you work on games, embodied AI, autonomous driving simulation, or world-model research, this is an open-source foundation you can pick up, try out, and reproduce right away.
What LingBot-World 2.0 is
LingBot-World 2.0 is a real-time interactive world model open-sourced by Robbyant. From the user's perspective, what's in front of you is a world with an initial frame and a backstory; every action you take (move, attack, cast a spell) generates the next stretch of footage in real time, genuinely pushing the world forward. From the system's perspective, a built-in agent is also observing the world, proactively proposing new events based on the current state — making the sky rain, sending enemies reinforcements — so the environment keeps evolving without any player intervention.

LingBot-World-Infinity is the only model in the lineup to achieve hour-level, near-unlimited generation duration in the general domain.
Three headline upgrades
Compared with version 1.0 from six months ago, 2.0 mainly did three things:
- Hour-level real-time generation: stable 720p / 60fps output even over long continuous runs
- Richer actions and events: attacking, archery, spell-casting, shooting, plus weather and environmental changes can all become part of the interaction
- A built-in agent proposing new events in real time: the world is no longer static — it continuously evolves on its own and changes with the user
Hands-on results
Silicon Star Pro's hands-on testing showed several typical scenarios:
A medieval fantasy valley: sweeping grassland, wooden signposts, village huts, flags, distant mountains, and mist, with complete spatial layering. The character can cross the valley with WASD, summon a horse and ride it kicking up dust, with the cloak hem and camera following the motion — a genuine open-world sense of speed. In combat, the musket brings muzzle flash and smoke, while magic shows each skill's area of effect and hit feedback through glowing particles, energy trails, and shockwaves.
From shopping mall to Van Gogh: following the moving viewpoint, the frame contains lanterns, stalls, smoke, pedestrians, and damp cobblestones. Clicking "Van Gogh-style hallucination" in the event proposals immediately brings swirling blue vortexes and Starry Night-style brushstrokes into the sky. Then entering a greengrocer, a fruit shop, and an upscale mall, the shot transitions and smoothness are excellent, with lighting, materials, and shelf structures holding up with strong consistency — showing the model's world-knowledge understanding of "what makes different stores different."
The technical foundation: three advantages
1. Causal generation (solving long-run degradation)
Traditional video generation models use bidirectional attention, where every frame can "see" both past and future — stunning for short clips. The cost is that the model never truly learns causality: it only knows that statistically "gunshot" is often followed by "glass shattering," not that the former causes the latter. Over long autoregressive runs, tiny errors compound repeatedly, and the frame gradually blurs, deforms, and finally collapses.
LingBot-World 2.0 introduces MoBA (hybrid bidirectional and autoregressive attention): the autoregressive part ensures the model generates forward in time, while the bidirectional part retains its grasp of overall frame relationships and visual quality. The current frame depends only on historical context, the current state, and user input — future information cannot leak ahead.
2. Real-time performance (few-step generation + streaming deployment)
High-quality video generation usually requires multi-step sampling — the more complex the scene, the longer the wait — but an interactive world can't have users stopping to wait after every keypress. The team's approach:
- Train a high-quality base model first, then use consistency distillation and DMD to compress multi-step diffusion into few-step generation
- On the deployment side, combine parallel inference, asynchronous VAE decoding, streaming transfer, and dynamic KV cache management to cut the delay from user input to on-screen feedback
The significance of 720p/60fps isn't just a sharper picture — it means feedback after each keypress is fast enough, scene continuity is smooth enough, and event changes don't feel jarring. Meanwhile, the 1.3B lightweight model lowers the deployment barrier, giving the system a chance to run on a single consumer-grade GPU.
The 1.3B lightweight model lowers the barrier to trying it locally.
3. Agentic harness (brain-cerebellum coordination)
LingBot-World 2.0 wraps a "brain-cerebellum" coordination framework around the video generation model:
- A VLM acts as the brain: continuously observing the frame, understanding user actions, and proposing what might happen next
- The underlying video generation model acts as the cerebellum: turning those events into continuous, believable footage
When you don't know what to do but still want to experience the unknown, you can click an event proposal in the right-side panel, or press U / O to get a next-step choice randomly generated from the current world state.
Known limitations (frankly listed by the team)
The tech report devotes a section to the current boundaries — hurdles the whole field still needs to clear together:
- Long-term memory remains the biggest problem: the model can stay visually stable for a long time, but when an area leaves the context window and you come back to it later, it feels more like a similar area regenerated from scratch — it doesn't necessarily remember "that door was opened a moment ago." The team sums it up as "the world persists in appearance but not in identity"
- Consistency and physics understanding: identity and style can slowly drift over very long explorations; physics understanding is far from the determinism of a real engine — characters and objects occasionally clip through each other and collision relationships get confused. The model has learned the visual appearance of physics, but not physics itself
- Compute requirements: the 14B main model suits high-quality experiences and research validation, and the 1.3B lightweight model lowers the local barrier, but getting a world to run smoothly on your own hardware is still a systems engineering project
How to get started
The main model is released under a non-commercial open-source license. Multiple entry points:
- Website: technology.robbyant.com/lingbot-world-v2
- Code (GitHub): github.com/Robbyant/lingbot-world-v2
- Model (HuggingFace): huggingface.co/collections/robbyant/lingbot-world-v2
- Model (ModelScope): modelscope.cn/collections/Robbyant/LingBot-World-V2
- Try online (Reactor): reactor.inc/lingbot-world-v2
- Mobile: the "World Model" feature in the Lingguang app
- Tech report: github.com/Robbyant/lingbot-world-v2/blob/main/paper.pdf
💡 Tip: note that the open-source license is non-commercial. Research, trying it out, and reproduction are fine; confirm the licensing terms before any commercial use.
Use cases
- Games / interactive content: turn script-driven into world-driven, with story growing naturally out of player behavior and agent event proposals
- Creators: shift from crafting frames to steering the direction the world evolves through actions and events — "from painting every leaf to planting a tree that grows by itself"
- Embodied AI / autonomous driving: provides a continuously changing simulation environment with a constant stream of unexpected events — far closer to real-world complexity than a fixed-scene replay buffer
- Multiplayer interaction: supports multiple people entering the same world, offering a working prototype for AI-native multiplayer interaction
- World-model research: provides an event-driven, sustainably running public experimental foundation on which researchers can test key questions like long-horizon causal consistency, open-domain dynamics, and agent coordination with environmental feedback
