WorldToken

WorldToken

Time-First Sequence Modeling
for Robotic Imitation Learning

Chunkai Yang1,*Andong Yang2,*Di Huang1,†Chao Gao2,†Guyue Zhou2

1 Wuhan University     2 Tsinghua University

* Equal contribution   ·   † Corresponding authors

Can a robot read the world as a language model reads text?

Paper Figure 1 compares a language model that reads text tokens with WorldToken, which reads a sequence of robot observations to generate actions.
A language model reads a sequence of text tokens to predict what comes next. WorldToken combines what the robot sees and senses at each step into one world token, then reads the sequence to plan its next movements.
1World token per stepWorldToken combines what the robot sees and senses at each step into one compact representation.
59.4%Success in kitchen tasksAn 85-million-parameter model learns to perform 23 tasks in the RoboCasa kitchen simulator.
95%Success in block rearrangementThe robot uses more than two minutes of past observations to keep track of a sequence of rearrangements.
01 / Method

WorldToken connects what the robot sees over time.

At each step, WorldToken combines the camera views, the robot’s body position, and its task instruction into a compact representation called a world token. It uses the sequence of these tokens to decide how to move.

WorldToken combines what the robot sees and senses, connects it with earlier observations, and plans the next movements.
The robot observes its surroundings, plans and carries out a few movements, then observes again.
01 — PERCEIVE

The robot observes its surroundings.

WorldToken brings together the robot’s camera views, body position, and task instruction so it can use them together.

02 — REMEMBER

It uses what happened earlier.

The model reads the world tokens in time order, connecting the current situation with earlier observations.

03 — ACT

It plans its next movements.

Using the current observation and its history, the model generates a short sequence of movements for the robot to carry out.

02 / Why history matters

A longer memory helps the robot complete the task.

The robot must rearrange three colored blocks through a sequence of swaps. Both videos use the same model and start with the same block arrangement. We change how far back the model can look.

Longer memory

About 146 seconds of history
The robot completes the sequence in about 2 min 23 s.

Shorter memory

About 15 seconds of history
The robot does not finish within 3 min 30 s.
0:00 / 3:30 simulated

Remembering more of the past improves task completion.

When the robot can look further back, it is more likely to finish the block-rearrangement task. The bars show how often it succeeds with different amounts of memory.

How much of the past the robot can see

The robot can keep rearranging blocks for many minutes.

In this longer demonstration, the robot continues after its first success and makes 31 correctly ordered swaps.

Watch the full recording below, or increase the playback speed to see how the sequence unfolds.

Watch more block-rearrangement videos
Full recording: 14 min 36 s
03 / Multitask control

WorldToken learns to handle everyday kitchen tasks.

Watch all 22 kitchen demonstrations

Place a mango in the cabinet

Open the cabinet door

Open the left drawer

Turn on the faucet

Turn on the stove

Press the coffee-machine button

04 / Learning from demonstrations

More demonstrations help the robot complete more tasks.

WorldToken learns by following examples of how a task should be done. Giving it more examples helps it perform kitchen tasks more successfully.

The robot learns from examples.

We trained an 85-million-parameter model with different numbers of demonstrations for each task. As the number of examples grew, it succeeded more often across 23 kitchen tasks.

The chart shows how often the robot completes a task successfully.

Read more about the experiments in the paper
Number of demonstrations for each task

The scaling study also compares models of different sizes.

The left plot shows how adding demonstrations reduces errors in predicted movements. The right plot shows how those errors change with model size.

Paper Figure 4 compares action prediction errors across different numbers of demonstrations and different model sizes.

More demonstrations help the robot predict movements more accurately, while the benefits of increasing model size eventually level off.

Model size (N). N1–N5 denote models with 44.3M, 85.3M, 218.8M, 648.9M, and 1.49B parameters, respectively.

Training data (D). D denotes the number of training demonstrations per task. For example, D2900 means 2,900 demonstrations for each of the 23 RoboCasa tasks.

05 / Open research

You can explore the work and try it yourself.

We share our code, trained models, and experiment records so you can try WorldToken and learn more about how it works.

Citation

@article{yang2026worldtoken,
  title   = {WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning},
  author  = {Yang, Chunkai and Yang, Andong and Huang, Di and Gao, Chao and Zhou, Guyue},
  journal = {arXiv preprint arXiv:2608.22591},
  year    = {2026},
  url     = {https://arxiv.org/abs/2608.22591}
}