HEX is a whole-body vision-language-action framework for full-sized humanoid robots. It combines a Qwen-VL backbone, a Unified Proprioceptive Predictor (UPP), and a flow-matching action head to predict continuous future actions. The key idea of HEX is to align heterogeneous humanoid states into shared body-part slots and learn predictive body dynamics from cross-embodiment humanoid data. This enables the policy to transfer across different humanoid platforms and perform long-horizon whole-body manipulation. During deployment, HEX directly predicts arm, hand, and waist actions, while providing high-level commands to a low-level RL-based whole-body controller for generating leg actions. This design enables coordinated and stable humanoid manipulation.
- ✅ 2026/10/04 Optimize the pretraining and fine-tuning code.
- ✅ 2026/10/01 Release improved model checkpoints with better performance.
- ✅ 2026/09/18: All pretraining and fine-tuning datasets for HEX have been released.
- ✅ 2026/05/17: The pretraining and fine-tuning code for HEX has been released.
First, git clone this repo and cd into it.
# clone project
git clone https://github.com/Open-X-Humanoid/HEX.git
cd HEXThen create python/pytorch env.
# crerate conda environment
conda create -n hex python=3.10 -y
conda activate hex
# Install env dependencies
sudo apt update
sudo apt install libegl1-mesa-dev libglu1-mesa
# Install requirements
pip install -r requirements.txt
# Install FlashAttention2
pip install flash-attn --no-build-isolation
# Install HEX
pip install -e .If flash-attn fails to install correctly, you can run
python hex/utils/test_flash_attn.pyto check the versions of PyTorch, CUDA, and the libstdc++ ABI. Then, manually download a compatible wheel from the flash-attn release. We use version 2.7.3. However, for newer GPUs (e.g., NVIDIA RTX 5090), you should install the latest available release (e.g., version 2.8.3) to ensure compatibility. Example:
wget https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.3/flash_attn-2.7.3+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl
pip install flash_attn-2.7.3+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whlWe release the pretrained HEX checkpoint and provide an improved checkpoint trained with a refined data mixture, where lower-quality data sources are down-weighted, on Hugging Face.
| Description | Params | Link |
|---|---|---|
| HEX | 2.4B | 🤗 HEX-model |
To download the HEX checkpoint, first modify the target download path in hex/utils/download_model_hex.py, and then run:
python hex/utils/download_model_hex.pyBefore running inference, please also download the Qwen3-VL base model:
python hex/utils/download_model_qwen.pyAfter downloading Qwen3-VL, update the framework.qwenvl.base_vlm field in the config.yaml file of the downloaded HEX checkpoint to your local Qwen3-VL path.
Once both the HEX checkpoint and the Qwen3-VL model are prepared, follow notebooks/eval_model.ipynb to run model inference.
We release all processed datasets used by HEX on 🤗 Hugging Face. The released data have already been converted into the format used by HEX and can be directly used for pretraining, fine-tuning, and evaluation without additional preprocessing.
The dataset repository is organized into two main subsets:
pretrain/: processed multi-embodiment datasets used for HEX pretraining.eval/: real-world task datasets used for fine-tuning and evaluation.
The overall structure is:
HEX-Datasets/
├── pretrain/
│ ├── agibot/
│ ├── g1/
│ ├── h1/
│ ├── leju/
│ ├── tiangong2/
│ ├── tiangong3/
│ └── tianyi/
├── eval/
│ ├── dvt217_carry_boxes_and_avoid_obstacles/
│ ├── dvt217_carry_boxes_follow_human/
│ ├── dvt217_imitate_posture/
│ ├── dvt217_pour_wine_follow_the_finger/
│ ├── dvt217_turn_around_and_carry_boxes/
│ ├── evt12_carry_box_and_tidy_table/
│ ├── evt12_put_cube_in_box/
│ ├── evt12_tidy_table/
│ ├── evt2_40_pick_up_box/
│ ├── evt2_40_pick_up_toy/
│ └── ...
└── eval_others/ # deprecated
Note:
eval_others/is a legacy directory and is no longer used in the current HEX evaluation pipeline.
To download the released datasets, run:
bash scripts/download_datasets.shSee the README files under each active subset for more detailed dataset descriptions.
Original data sources and preprocessing
HEX is trained on data collected from multiple humanoid embodiments and public datasets. The original data sources are listed below.
| Embodiment / Platform | Source | Dataset |
|---|---|---|
| Tiangong Series | HEX | 🤗 HF Link |
| Unitree G1 | Humanoid Everyday | 🤗 HF Link |
| AgiBot-to-Unitree G1 | AgiBot World Colosseo & TrajBooster | 🤗 HF Link |
| Unitree H1 | Humanoid Everyday | 🤗 HF Link |
| Leju Kuavo | RoboCOIN | 🤗 HF Link |
The released HEX datasets follow the LeRobot v2.1 data format. Each dataset therefore requires a corresponding modality.json.
These preprocessing steps are only required when reconstructing the datasets from the original sources. The processed datasets released in 🤗 X-Humanoid/HEX-Datasets can be used directly.
Due to commercial restrictions, we are unable to release the data collection pipeline used for the Tiangong series robots.
For users interested in collecting data on Unitree G1, we recommend referring to the following open-source data collection pipelines:
- OpenTrajBooster, which uses a VR headset and handheld joysticks for full-body teleoperation.
- Psi0: uses a PICO VR headset with controllers, along with a waist tracker and foot trackers for full-body teleoperation.
You can download our pretrained HEX model and skip this step if you only want to run inference, fine-tuning, or evaluation.
Before pretraining, download the Qwen3-VL backbone:
bash scripts/download_models.shThen, configure the following fields in scripts/pretrain_hex.sh:
base_vlm: path to the downloaded Qwen3-VL backbone.data_root_dir: path to the local pretraining dataset directory.dataset_name: dataset mixture used for pretraining.
The dataset root is automatically exposed through HEX_PRETRAIN_DATA_ROOT, so no source-code modification is required.
Finally, start pretraining with:
bash scripts/pretrain_hex.shOther training settings can be directly adjusted in scripts/pretrain_hex.sh.
After obtaining the pretrained HEX model, you can further fine-tune HEX on downstream tasks using the released evaluation datasets.
Configure the following fields in scripts/fine_tune_hex.sh:
base_vlm: path to the Qwen3-VL backbone.data_root_dir: path to the local evaluation dataset directory.dataset_name: downstream task used for fine-tuning.pretrained_models_path: path to the pretrained HEX checkpoint.
Then, start fine-tuning with:
bash scripts/fine_tune_hex.shOther training settings can be directly adjusted in scripts/fine_tune_hex.sh.
Due to commercial restrictions, the RL-based low-level whole-body controller used for the Tiangong series robots is not open-sourced. However, we provide a sample real-world deployment interface in examples/real_world, together with the corresponding deployment scripts:
Deploy the HEX policy on the server:
bash scripts/deploy_server.shRun the HEX client on the robot side:
bash scripts/deploy_client.shIf you want to deploy your own model on Unitree G1, you may refer to the following open-source projects:
- OpenTrajBooster: uses HOMIE as the low-level RL-based whole-body controller.
- Psi0: uses AMO as the low-level RL-based whole-body controller.
When training your own low-level controller, please make sure that the command space output by the high-level VLA policy matches the input space expected by the low-level controller. The dataset construction process should also follow the same interface for consistent training and deployment.
Thanks to the cross-embodiment capability of VLA models, HEX can also be evaluated in simulation environments such as LIBERO.
First, download the LIBERO datasets:
python hex/utils/download_dataset_libero.py --base_dir /your/dataset/pathThen, replace the modality.json file for each LIBERO suite with the provided template in examples/LIBERO/modality.json.
Next, modify the following fields in scripts/libero/train_hex_libero.sh:
base_vlm: path to your Qwen3-VL backbonedataset_name: name of the LIBERO dataset mixturedata_root_dir: path to your local LIBERO dataset directory
Then start training with:
bash scripts/libero/train_hex_libero.shFor evaluation, modify the following fields in scripts/libero/eval_libero.sh:
ckpt_root: root directory of the trained checkpointckpt_path: relative path to the checkpoint file
Then run:
bash scripts/libero/eval_libero.sh@article{bai2026hex,
title={HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation},
author={Bai, Shuanghao and Li, Meng and Lv, Xinyuan and Wang, Jiawei and Wang, Xinhua and Liao, Fei and Hou, Chengkai and Gu, Langzhe and Zhou, Wanqi and Wu, Kun and others},
journal={arXiv preprint arXiv:2604.07993},
year={2026}
}
This project draws inspiration from and builds upon several notable open-source projects, including: StarVLA, Isaac-GR00T, HiMoE-VLA, LeRobot, Humanoid Everyday, RoboCOIN, AgiBot-World, and OpenTrajBooster.
