Oneira

Oneira From Open-Ended Generation to Open-World Interaction in Video World Models

Xindi Yang1 Baolu Li3 Liam Lee4 Zhenfei Yin5 Songxin Zhang4 Zhuoyang Song4 Xu Jia3 Jianfei Cai1 Tien-Tsin Wong1 Bingyi Jing4 Mengyue Yang2
1Monash University 2University of Bristol 3Dalian University of Technology 4The Chinese University of Hong Kong, Shenzhen 5Oxford University

Corresponding author

What appears in the generated world becomes actionable, and what the agent changes remains changed.

The inset in each video is the coarse conditioning video that the engine renders from the world state table.

Abstract

Generative video world models can now synthesize increasingly open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should also expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are dynamically incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state continuously managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent dynamically incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not explicitly represented in the state. Experiments show that Oneira enables direct and logically consistent interaction with newly generated objects during open-world exploration, while preserving the effects of prior interactions consistently over long horizons.

Full interaction in a generated world

Open-ended generation does not imply full interaction. A fully interactive generative world requires two properties.

Open-World Interactivity

As exploration reveals or generates new entities, the interaction space should expand with the world, so that newly created content can become explicit targets of subsequent actions.

Persistent State

When an agent changes the world, the outcome of that interaction should become a durable part of the world state and continue to constrain future observations and interactions.

Oneira meets both requirements with an explicit world state table that a coding agent updates. As shown in Fig. 1, it supports navigation and interaction beyond the input image (a), instance-level control (b), generalization to novel objects (c), and effects beyond the target object (d).

Four rows of frames. (a) Explore, then interact. (b) Instance-level interaction. (c) Novel objects. (d) Effects beyond the target. Each frame carries the conditioning video as an inset, and a strip under each row shows the world state table.
Fig. 1. Oneira conditions video generation on an explicit world state table that a coding agent updates. The small image in the corner of each frame shows the coarse conditioning video rendered from the table. Purple marks the default state, yellow a changed state, and a dashed outline a box removed from the table.

Method

Oneira separates world evolution from visual realization. The explicit world state determines what exists and what changes, while the video model determines how those changes look and move.

Pipeline of Oneira. The coding agent reads the world state table and plans an action for the goal of picking up the lower left apple. The engine writes the action into the table and renders the coarse conditioning video along the camera path. The video generation model generates the segment from the first frame, optional memory frames, the conditioning video and the chunk captions.
Fig. 2. Pipeline of Oneira for one segment. The sphere is a schematic, and the conditioning and generated frames come from one segment produced by Oneira.

1Coding agent $A$

Given the input image and a goal, the coding agent $A$ reads the previous world state table $\mathcal{T}_{k-1}$ and plans each action $\mathcal{P}_k$. For the goal in Fig. 2, $\mathcal{P}_k$ picks up the lower left apple in chunk 6 and records that apples 2 and 3 leave as a side effect.

2World state table $\mathcal{T}_k$

The world state table $\mathcal{T}_k$ keeps only the interaction-relevant state of the scene. It records the walkable ground, and each object entry holds a semantic label, a 3D box, a presence flag, and a state per chunk. When the pear in Fig. 2 appears in the generated frames, the coding agent registers it in the table.

3Engine $E$

The engine $E$ writes each planned action into the table and renders the updated table $\mathcal{T}_k$ along the camera path $C_k$ as the coarse conditioning video $V_k$. An interaction appears in $V_k$ as the removal or recoloring of an object box. In Fig. 2, the lower left apple loses its box at t2, and the two tumbling apples lose theirs at t3.

4Video generation model $G$

The video model $G$ generates the segment from the first frame, optional memory frames, $V_k$, and the chunk captions $c_k$. It fills in the appearance, motion, and interaction details that the world state leaves unspecified. The last frame and the updated table $\mathcal{T}_k$ start the next segment.

$$\begin{aligned} \mathcal{P}_k &= A\big(\mathcal{T}_{k-1}, \hat{I}^{(k)}_0, a_k\big) && \text{plan} \\ (\mathcal{T}_k, V_k, c_k) &= E\big(\mathcal{T}_{k-1}, \mathcal{P}_k\big) && \text{update and render} \\ \hat{I}^{(k)} &\sim G\big(\hat{I}^{(k)}_0, V_k, c_k\big) && \text{generate} \end{aligned}$$

Here $a_k$ is the action input of segment $k$ and $\hat{I}^{(k)}_0$ is its first frame. This loop follows a standard world model with a transition $s_k = f(s_{k-1}, a_k)$ and an observation model $o_k \sim p(o \mid s_k)$. The world state table $\mathcal{T}_k$ acts as the explicit state $s_k$, the coding agent $A$ and the engine $E$ jointly instantiate $f$, and the video model $G$ serves as $p$.

Results

Our baselines are the MiniMax-H3 Ref2VA backbone and three interactive video world models, LingBot-World-V2, YUME 1.5, and AlayaWorld. The first two parts below compare Oneira with them on the same first frame and the same chunk captions, and the last two parts show Oneira on its own.

Each case shows the first frame, the conditioning video of Oneira, and the chunk captions next to the generated videos. Orange marks the chunks whose captions carry an interaction. All videos of a case play in sync while the case is on screen, and the caption panel highlights the current chunk during playback. Click a video to play or pause, and double-click it to view it large.

Quantitative results

We evaluate interaction control, memory of changed states, and visual quality on InteractionBench. Bold marks the best entry in each column based on unrounded scores.

MethodInteraction ↑Timing ↑Target ↑
MiniMax-H3 Ref2VA0.300.570.53
LingBot-World-V20.490.400.39
YUME 1.50.050.430.22
AlayaWorld0.080.160.26
Oneira (ours)0.780.760.82

Pass rates on the 100 interaction cases, checking the action, its chunk, and the changed instance.