Alibaba's HappyOyster 1.0: Create a World You Can Walk Into with a Single Sentence
Alibaba has released the world model product HappyOyster 1.0, which generates open worlds with real-time exploration and physics interaction from a single sentence, offering World Exploration and Real-Time Directing modes; the API is expected to open in early July.


Alibaba's HappyOyster 1.0: Create a World You Can Walk Into with a Single Sentence
Alibaba has released the world model product HappyOyster 1.0, which generates open worlds with real-time exploration and physics interaction from a single sentence, offering World Exploration and Real-Time Directing modes; the API is expected to open in early July.
HappyOyster 1.0 is a world model product officially launched by Alibaba on June 17, 2026. Its core capability: one sentence or one image is enough to generate a complete, explorable, interactive digital world. Unlike text-to-video, where "the generation is the final cut", with HappyOyster the experience is only beginning the moment generation completes.
The product's name draws inspiration from Shakespeare's famous line "The world is your oyster".

The core version of HappyOyster 1.0.
Two Core Features
World Exploration (Adventure)
You become part of the world as a character. Switch freely between first-person and third-person, with over a minute of real-time movement and camera control supported.
New interactive actions include: dashing/sprinting, crouching, attacking, and jumping, plus more complex environmental interactions — riding in and driving vehicles, and fighting with various weapons.
The key experiential difference is physical interaction feedback. For example: after a landed punch, the opponent triggers a staggered recoil reaction; a character can use a torch item and the scene lighting switches over sensibly; an explorer crossing a thick snow-covered ridge leaves footprints with every step and kicks up powder from the snow they compress.
Whatever art style the world uses (realistic, claymation, anime), anyone can walk into it the same way and issue real-time commands.
Real-Time Directing
You become the director standing above the world. Streaming generation, speak and it performs, with commands injected at any time to change where things go. Three key features:
- Pause: freeze the world at any moment and resume once you've thought it through
- Rewind: fold back to any earlier node mid-performance and start over, with the original version preserved
- Story branching: fork completely different directions from the same node
Multimodal references (using @image to lock in a character's appearance) support consistency over 3-minute-long sequences.
How It Differs from Text-to-Video
Text-to-video learns a one-way mapping of "text -> video", and it ends when generation ends. A world model learns:
the transition pattern of current state + your action -> next state
The model must understand the current scene structure, entity attributes, and physical relationships, and accurately predict and render the world's next state even as you throw commands at it at any moment.
Core Technology
HappyOyster 1.0's technical advantages come down to four points:
1. World state modeling: the world's current state is compressed into a latent state summary, updated and recursively passed along with each generated segment. The state summary is serializable and archivable — this is what makes pause, rewind, and story branching possible.
2. Endogenous consistency: when you enter a world, every character and every key prop is issued an "identity card", which the model carries throughout generation. When a character turns around, gets occluded, or walks off-screen and reappears minutes later, their face and clothing never change.
3. Open causal action space: action commands and natural language share the same semantic interface, with no predefined action set required. The model learns causal chains on its own through large-scale causal training (punch thrown -> hit lands -> NPC knocked to the ground -> dust kicks up off the floor -> a wine glass is shaken loose).
4. Long-sequence audio-video coordination: audio and visuals are jointly decoded and generated from the same world state. Gravel crunches underfoot, engines roar as they accelerate — sound is not dubbed in afterward; it is part of the world itself.
Use Cases
| Scenario | Value |
|---|---|
| Interactive games | Generate open-world prototypes with real-time physics feedback from a single sentence, cutting the cycle from weeks to hours |
| Real-time virtual companionship | Generate virtual characters you can interact with at any time — they listen, talk, and stay consistent over long sessions |
| Interactive short dramas | The pause-rewind-branch trio lets one opening fork into multiple storylines |
| Live streaming | A single viewer command instantly changes where the scene goes |
| Virtual travel experiences | Places no camera can reach — the lunar surface, undersea palaces — continuously simulated in pixel space |
How to Use It
- China site: www.happyoyster.cn (live now; register with a phone number and receive free creation credits for logging in daily)
- API: expected to open in early July 2026