Odyssey unveils Odyssey-3 for robots, cars and game agents

The model used 20 hours of simulated driving data, but Odyssey's broad physical-intelligence claims still rest on its own evaluations.

By · Published

Primary source: X

Why it matters

Odyssey is betting that one visual foundation model can cut the task-specific data needed for robotics and autonomy. Public access will test whether its demos transfer beyond company-run conditions.

Odyssey unveils Odyssey-3 for robots, cars and game agents

Oliver Cameron and Jeff Hawke unveiled Odyssey-3 on September 15th, presenting a single foundation world model that Odyssey says can underpin control systems for robot arms, humanoids, cars, drones and video-game agents.

https://x.com/odysseyml/status/2099900067356586276

poster=/api/storage/public-objects/tweet-videos/odyssey-3-world-model-robots-cars-game-agents-poster-a667bdde.jpg|Video from @odysseyml on X

The release is the clearest expression yet of the founders' original bet. Cameron previously co-founded autonomous-driving company Voyage and became a vice president of product at Cruise after its 2021 acquisition. Hawke worked on research and engineering at autonomous-driving developer Wayve. They founded Odyssey in 2023 around the idea that a model trained to predict how the world changes could supply reusable physical knowledge across machines, rather than forcing each robot or vehicle to learn mainly from task-specific demonstrations.

Odyssey-3 is an autoregressive diffusion transformer trained on visual observations, according to Cameron and Hawke's technical overview. Odyssey attaches smaller action decoders to the frozen foundation model, then trains those components to translate its internal representations into controls for a particular machine or software environment.

That architecture is the product bet: expensive, broad pretraining at the center, followed by relatively small amounts of specialized experience for each application. Odyssey says robot-arm policies trained with tens of hours of demonstrations recovered from missed grasps and objects dropped in unusual positions, including behaviors that were not present in the demonstrations.

Twenty hours of simulation, then Indian roads

Odyssey's most concrete test involved autonomous driving. Odyssey says it trained a driving policy with 20 hours of simulated data while keeping the world model frozen. The policy then generated waypoints in real time during closed-loop driving on public roads in India.

In Odyssey's evaluation, policies trained only in simulation traveled about 77% as far between safety-driver interventions as policies trained on real driving footage. Both sets of policies were tested on busy roads involving overtaking vehicles, bends and junctions, Odyssey said.

The figure is narrower than the headline claim that Odyssey-3 can drive cars. It compares two Odyssey-trained policies rather than measuring Odyssey-3 against a commercial autonomous-driving system, and Odyssey did not publish a common benchmark that would make the result directly comparable with work from other labs. The test does support Cameron and Hawke's underlying thesis that a pretrained visual model can reduce the amount of application-specific data required for a physical task.

Odyssey used a similar recipe for aerial navigation. A drone policy trained on tens of hours of simulated flight data demonstrated obstacle avoidance and stable flight in a simulated indoor environment, according to Odyssey. The drone demonstrations were not flights in a physical room.

A foundation for humanoid control

Odyssey also disclosed a research collaboration with Flexion, a robotics developer led by co-founder and CEO Nikita Rudin. Flexion built humanoid control policies on Odyssey-3 using tens of hours of teleoperation data. Demonstrations published by Odyssey show a humanoid opening containers and moving household objects.

Odyssey says the resulting policies continued operating through lighting changes that caused the vision-language-action baselines it tested to fail. The announcement does not identify those baselines or provide enough detail to reproduce the comparison. Flexion supplied substantial robot-learning and control engineering, making the demonstrations evidence for Odyssey-3 as a useful backbone rather than a robot controller that works without adaptation.

For robot arms, Odyssey is working with Poke & Wiggle to evaluate how performance transfers across bodies, camera viewpoints and control systems. Those evaluations matter because Odyssey's general-purpose claim depends on consistent transfer, rather than a collection of separate demonstrations tuned for favorable settings.

The same backbone enters video games

Odyssey trained game-control policies on recordings paired with keyboard and mouse inputs. Odyssey says policies trained in Grand Theft Auto V produced movement in Red Dead Redemption 2 and motorcycle riding in Sleeping Dogs without further policy training on those games. In one experiment, about two hours of Grand Theft Auto footage was enough for a mobility policy to control horseback movement in Red Dead Redemption 2.

The gaming work serves two purposes. It offers a cheaper environment for testing whether learned behaviors transfer between bodies and control schemes, and it gives Odyssey simulated worlds in which other AI agents can take actions and learn from the consequences. Odyssey has previously explored that feedback loop through PROWL, a reinforcement-learning framework in which agents search simulated environments for failures that can improve the world model.

Odyssey plans to make Odyssey-3 publicly available in the coming weeks. Tuesday's announcement is therefore an unveiling and a set of company-run demonstrations, rather than a public model release that outside developers can test today.

The timing follows a large expansion of Odyssey's resources. In June, Odyssey raised a $310 million Series B at a $1.45 billion valuation, led by Natural Capital with participation from Amazon, AMD Ventures, GV, EQT and In-Q-Tel. The financing brought Odyssey's disclosed funding to $337 million and made AWS its preferred cloud provider, with plans to use Amazon's Trainium chips alongside other hardware.

Cameron and Hawke are using that capital to pursue a model category that remains far smaller than frontier language models by their own account. Odyssey-3 moves Odyssey beyond generating interactive scenes and into policies that act inside physical and virtual environments. Public access will determine whether the broad demonstrations share a genuinely reusable representation of the world or a common model surrounded by substantial task-specific engineering.

Reader comments

Conversation for this story loads after sign-in.