Yang Likun's vision from 4 years ago delivers results: H-JEPA maze success rate rises from 18% to 73%.
动察 Beating AI News Flash: AMI Labs, founded by Yann LeCun, together with multiple universities, has unveiled H-JEPA, a world model capable of hierarchical prediction of the future and action planning. Higher layers handle long-term goals, while lower layers handle specific actions. The paper, code, and pre-trained weights have all been made public.
H-JEPA continues the JEPA approach. It first converts scenes into abstract environmental features, then combines actions to predict future states, without the need to generate video frame by frame. However, long-term planning and immediate control require different information. For example, when a quadruped robot navigates a maze, choosing a route mainly depends on position and passages, while moving its legs cannot do without body posture. A single-layer model uses the same set of features to handle both tasks, and the farther the planning horizon, the more strained it becomes.
H-JEPA organizes multiple JEPAs into different planning levels. Higher layers predict a more distant future, pass intermediate goals to lower layers, and gradually refine them into specific actions. Unlike the previous HWM, in which all layers share the same set of features, H-JEPA lets each layer learn its own environmental features. In maze experiments, higher layers retained positional information and gradually ignored short-term details such as leg posture. Yann LeCun proposed a similar concept as early as 2022, and now the team has provided an implementation that can be trained end to end.
In the Visual AntMaze simulated maze, the three-level H-JEPA achieved a success rate of 73.3%, while the single-layer LeWorldModel reached 18.0%, and the former required less planning computation. However, more planning levels are not always better. In the Push-T task of simulating object pushing, the three-layer model was actually inferior to the two-layer model. The team also conducted offline tests using the real robot video dataset DROID, where predicted action trajectories were closer to expert demonstrations than those of the single-layer model.