Odyssey-3 Takes the Top Spot as World Models Advance

Odyssey released the Odyssey-3 series on October 8, with the Pro version topping the Physics-IQ Verified video physics benchmark with a score of 66.1. Beyond generating beautiful visuals, its more noteworthy advance is enabling environments to evolve continuously in response to user actions.
Odyssey-3 Takes the Top Spot, Advancing World Models
Odyssey has pushed competition among world models into a more readily testable phase: when shown a real-world physics experiment, can a model correctly predict what happens next?
According to an October 9 report by IT Home, Odyssey announced yesterday (October 8) the launch of the Odyssey-3 series of foundational world models. Among them, Odyssey-3 Pro scored 66.1 points on the Physics-IQ Verified video-to-video benchmark, setting a new record for the highest score on the leaderboard. The series includes a Standard version and a Pro version: the former balances physical accuracy and generation costs, while the latter focuses on stronger physical prediction capabilities.
This release deserves attention. Video models have become highly capable of producing visual convincingness, but attractive visuals and credible physical processes remain two separate things that need to be verified independently. How one ball moves after colliding with another, how liquid diverts when it encounters an obstacle, and whether an object is still in its original position after the camera turns away and then back again—these details determine whether the generated environment is merely footage to watch or a space that users and agents might repeatedly interact with.
The results reported for Odyssey-3 provide additional evidence for the latter direction. But “topping a physics benchmark” is still not the same as “understanding the real world”; the boundary between the two needs to be carefully distinguished.
The 66.1 Score Measures Whether a Model Can Continue a Physical Process
Physics-IQ Verified was jointly developed by Anates Labs and DeepMind. According to the report, it asks models to continue real-world physics experiment videos and then compares their predictions with the actual experimental results. The benchmark covers fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.
The value of this test lies in shifting the evaluation target from visual quality to the process of change.
For an easy-to-understand example: given a video of liquid flowing toward an obstacle, the model must determine where the liquid will flow around it, how it will disperse, and whether these changes conform to the actual experiment. This example explains the testing approach; it is not a specific test question disclosed in the report. A model may be able to generate clear water surfaces and attractive reflections while still predicting the subsequent flow incorrectly.
Looking real visually does not mean predicting the dynamics correctly. The latter is particularly important for world models, because one task they may eventually perform is helping systems determine “what will happen after taking this action.”
A score of 66.1 indicates that Odyssey-3 Pro leads in the physical prediction capabilities measured by this benchmark. However, the score must be interpreted according to its original meaning: the report does not provide the scoring formula, category-level results, the complete list of competitors, or the margin of victory. Therefore, 66.1 cannot be directly interpreted as “the model can solve 66.1% of real-world physics problems,” nor can this figure alone establish that it leads in every category of physical phenomenon.
A more accurate assessment is: On the Physics-IQ Verified video-continuation evaluation, Odyssey-3 Pro delivered the best result on the current leaderboard. This is a clear improvement, but it is also a conditional conclusion.

From Generating Videos to Letting Users Change What Happens Next
Another central theme of Odyssey-3 is interaction.
According to the information disclosed this time, it uses an autoregressive diffusion transformer architecture and can generate embodied environments from prompts. During generation, users can move the viewpoint, take actions, or introduce events, while the model continuously predicts how the environment will change based on previous observations and the latest input.
Placed in a usage scenario, the distinction becomes fairly clear.
In common video-generation tasks, the scene and action are specified first, and the model then produces a video. The direction demonstrated by Odyssey-3 is different: users first enter a generated environment, change their position or trigger events along the way, and then observe how the environment responds. The next segment must maintain the existing state while also incorporating the changes just introduced by the user.
This places higher demands on the model. It must not only preserve the visual style, but also maintain spatial relationships, object states, and the consequences of actions as much as possible. Rotating the camera should not make objects on a table disappear out of nowhere; knocking over a cup should be reflected in subsequent frames. This describes the problems an interactive world model is expected to solve and should not be taken as proof that Odyssey-3 has already achieved stable performance in all scenarios.
“Autoregressive” helps explain why the model is suited to continuous generation: it advances through time, using existing information to predict what comes next. Diffusion generation and transformers form part of its architecture. However, the public summary does not explain how the model preserves environmental state, how much historical observation it can process, or how it corrects errors after extended operation. The architectural name indicates a technical direction; actual interactive quality still needs to be demonstrated through runtime results.
The current preview supports three distinct forms of operation:
- First-person navigation: moving through the environment from the participant’s perspective, suitable for observing continuity and responsiveness during movement.
- Third-person navigation: observing the relationship between the character and the environment from an external viewpoint, making it easier to see whether actions affect the surrounding state.
- Independent camera movement: allowing the camera to move separately, providing more viewing angles for checking spatial consistency.
For developers, these capabilities mean that evaluation methods also need to change. Looking only at the most impressive ten or so seconds in an official clip makes it difficult to determine whether an environment is usable. Repeatedly following the same route, inspecting the same object from different angles, and triggering events continuously are often more effective ways to expose the model’s limitations.
The Two Sets of Leaderboard Results Answer Different Questions
In addition to physics prediction, the report also disclosed Odyssey-3’s performance in the WorldMark evaluation. It ranked first in three environment categories, with the following scores:
- First-person stylized environments: 77.2.
- Third-person realistic environments: 79.0.
- Third-person stylized environments: 76.3.
These results supplement the picture of its performance across different viewpoints and visual styles, but they cannot be directly compared with the 66.1 score on Physics-IQ Verified. The two evaluations address different questions, and their scores do not necessarily use the same scale.
The materials released this time do not disclose WorldMark’s complete scoring criteria. Therefore, the more cautious interpretation is that Odyssey-3 performed best in the three environment categories listed above; these results should not be further interpreted as meaning that it has solved all problems of interactive consistency. The report also does not provide a result for first-person realistic environments, so it would be inappropriate to fill in that gap speculatively.
Taken together, the two sets of results are still meaningful: Odyssey is seeking to demonstrate capabilities involving both physical-process prediction and the generation of interactive environments. The former concerns what will happen next, while the latter concerns whether the environment can continue operating after the user intervenes. For world models to move toward practical applications, both lines of capability need to advance.
How Developers Should View the Standard and Pro Versions
Launching both Standard and Pro versions of Odyssey-3 is a pragmatic product decision.
The cost of using a world model is not limited to the price of generating a single piece of content. Continuous interaction creates ongoing computational demands, while latency directly affects user operation: how long the environment takes to respond after the user moves, and whether the model can promptly incorporate a newly introduced event. These factors will affect which model tier developers ultimately choose.
According to the disclosed positioning, the Standard version balances physical accuracy and generation costs, while the Pro version offers stronger physical prediction capabilities. This suggests an initial selection strategy: if the goal is to validate an interaction flow or build an environmental prototype, cost and responsiveness may matter more; if the task depends heavily on accurately predicting the consequences of actions, stronger physical prediction capabilities will be more valuable.
But this is only a judgment based on product positioning. The materials released this time do not provide the prices, response latency, output specifications, concurrency limits, or detailed comparisons between the two versions on identical tasks. It is not yet possible to calculate which version offers better value, let alone treat the Pro version’s leaderboard results as a performance guarantee for the Standard version.
Teams that have already integrated models such as GPT, Claude, or Gemini also need to pay attention to differences at the interface layer. Continuous interactive environments may involve video observations, action inputs, and state updates; their integration requirements cannot be inferred simply from the word “model.” Existing reports have not disclosed Odyssey-3’s API protocol, nor have they provided information that it has been integrated into OpenAI Hub. There is therefore currently no basis for providing an example of a call using the OpenAI format.
Robots and Agents Will Need It, but Validation Cannot Be Skipped
The appeal of world models lies in their potential to provide agents with an environment in which they can trial and error.
If a system can predict the changes caused by an action, robots may be able to try different strategies in simulated environments before acting in the real world. Developers may also construct training scenarios and observe how agents respond to moving objects, changing routes, or unexpected events. For games and interactive content, continuously generated environments could reduce reliance on pre-scripted scenes.
These are all reasonable application directions, but they do not automatically count as capabilities already delivered by Odyssey-3.
Robotic control in particular requires a careful distinction between “looks reasonable” and “is sufficient to support decision-making.” During robotic-arm grasping, errors in object position, contact relationships, and action timing can all change the outcome. A predicted video that appears sufficiently natural does not necessarily mean that its information is accurate enough to control hardware directly. The materials released this time do not disclose results from real-world robot deployments or the success rate of transferring simulated strategies to real-world environments.
For teams preparing to evaluate this type of model, the most useful tests often center on their own tasks:
- After continuous interaction for a period of time, do object positions and the environmental layout remain consistent?
- Under similar initial conditions, can repeated operations produce reasonable and interpretable results?
- When multiple events occur in succession, can subsequent generations incorporate the preceding state changes?
- Does the model’s output contain the information required by the task, and can that information be read reliably?
These checks turn the claim that “world models have great potential” into concrete engineering questions. Leaderboards provide clues for screening; specific tasks determine whether a model can enter a product.
What Deserves Attention Is That World Models Are Beginning to Face More Concrete Tests
As of October 9, Odyssey-3’s clearest progress is that the Pro version set a new record on the Physics-IQ Verified leaderboard, the series demonstrated multiple navigation modes and event-interaction capabilities, and it ranked first in three WorldMark environment categories.
This makes it a new model worth continued attention from developers. Especially as video generation becomes increasingly adept at producing compelling visuals, evaluations such as physics-experiment continuation bring the discussion back to a more concrete question: how closely do the changes predicted by the model match the changes that actually occur?
The factors that will determine Odyssey-3’s practical value next will be data beyond the leaderboards: whether long-term interaction remains stable, whether action response is fast enough, how large the cost gap between the two versions is, and whether developers can connect these capabilities to their own systems through clear interfaces.
The 66.1 score deserves attention, but whether the environment can continue to withstand interaction is even more worth tracking. For a world model, the most convincing demonstration will ultimately be whether, after a user takes a new action, it can still correctly continue the world that follows.
Sources
- IT Home: Scoring 66.1 and Taking the Top Spot, the Odyssey-3 Foundational World Model Debuts: Reported on October 9, 2026, providing information about Odyssey’s model launch on October 8, as well as the Physics-IQ Verified score, the positioning of the Standard and Pro versions, the architecture, preview-version interaction capabilities, and WorldMark data. The publication facts and leaderboard figures in this article are based on that report; the scenario examples and engineering evaluation recommendations are analytical.



