Figure Released The VLA Model Helix: A One-sentence Instruction, Humanoid Robots To Do Housework Collaboratively

Feb 27, 2025 Leave a message

Recently, Figure AI, an innovative company in the field of robotics in the United States, released a major breakthrough: a general-purpose visual language action (VLA) model called Helix. For the first time, this model realizes high-speed continuous control of the complete upper body of a humanoid robot, and perfectly integrates perception, language understanding, and learning control.

 

The advent of the Helix model marks an important step forward in the operational flexibility of humanoid robots. With simple natural language commands, the robot can easily grasp almost any small household object, even those that have never been touched during training, without any prior demonstration or custom programming. This capability is due to the powerful generalization ability of the Helix model.

humanoid robot

Figure AI highlighted that the Helix model has created a number of industry firsts. For the first time, it enables high-speed continuous control of the entire upper body of a humanoid robot, including the flexible control of the wrist, torso, head and each finger. In testing, the robot successfully processed thousands of new items that were cluttered with disorganization, from glassware and toys to tools and clothes, without prior demonstration or programming.

 

What's even more amazing is that the Helix model also has multi-robot collaboration capabilities. In the test, the two robots were able to work together on long-term, complex tasks, working together on never-before-seen items, such as sorting out unfamiliar groceries together. This capability opens up more possibilities for the practical application of robots in the home environment.

 

The Helix model also demonstrates excellent scene understanding and semantic parsing capabilities. When prompted to "pick up a desert object", the robot is not only able to recognize that the toy cactus fits this abstract concept, but also selects the nearest hand and performs a precise grasping action. This universal gripping function, from language to motion, provides greater convenience for the deployment of humanoid robots in unstructured environments.

 

The Helix model was able to achieve these breakthroughs thanks to its groundbreaking dual-system architecture. The architecture consists of System 1 and System 2, which are responsible for high-speed precise control, scene understanding, and semantic parsing, respectively. System 2 is based on the open-source VLM with 7B parameters, which operates at a frequency of 7-9Hz to ensure generalization across objects and scenarios. System 1 is an 80M parameter visual motor strategy model, which converts the semantic representation of System 2 into continuous action instructions at a frequency of 200Hz to achieve millisecond-level real-time response. This decoupled architecture enables the two systems to perform their respective functions and work together to achieve efficient humanoid robot control.

 

Helix models use very few resources during training. Using only about 500 hours of high-quality supervised data, the team was able to achieve a robust generalization of objects. This data represents less than 5% of the size of previously collected VLA datasets and does not rely on multi-bot entity collection or multi-stage training. This achievement not only demonstrates the efficiency of the Helix model, but also provides more possibilities for the development of humanoid robots in the future.