ModelNVIDIANVIDIApublished Jul 1, 2026seen 2w

nvidia/ARDY-G1-RP-25FPS-Horizon52

Open original ↗

Captured source

source ↗
published Jul 1, 2026seen 2wcaptured 2whttp 200method plainlicense otherlibrary ardydownloads 344likes 8

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

[Paper](https://research.nvidia.com/labs/sil/projects/ardy/assets/ardy_paper.pdf), [Project Page](https://research.nvidia.com/labs/sil/projects/ardy/)

Description:

ARDY is an autoregressive diffusion model designed for interactive motion generation, supporting online text prompting and flexible long-horizon kinematic constraints (root paths/waypoints, full-body keyframes, and sparse joint positions/rotations) with real-time responsiveness.

ARDY-G1-RP-25FPS-Horizon52 was developed by NVIDIA as a part of the ARDY project. It was trained on the Bones Rigplay 1 dataset with the 34-joint Unitree G1 robot skeleton at 25 fps. See [below](#model-versions) for other model variants.

This model is ready for commercial or non-commercial use.

License/Terms of Use:

Use of this model is governed by the NVIDIA Open Model Agreement

Deployment Geography:

Global

Use Case:

Developers and researchers with any level of animation experience can use ARDY to generate controllable humanoid motions in their real-time applications. This could include motion planning for humanoid robots, character movement in digital twin and industrial simulations, digital human motion for synthetic data, and animations for games and other interactive applications.

Release Date:

HuggingFace: 07/10/2026 via HuggingFace

Reference:

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

Model Architecture:

Architecture Type: Diffusion Model

Network Architecture: Novel Two-Stage Transformer

Number of model parameters: 326 M

Input:

Input Type(s): Text, Other: Pose Constraints, History Poses

Input Format(s): String, Tensor

Input Parameters: One-Dimensional (1D), N-Dimensional (ND)

Other Properties Related to Input: History pose duration is max 8 sec.

Output:

Output Type(s): Other: Pose Sequence

Output Format: Tensor

Output Parameters: N-Dimensional (ND)

Other Properties Related to Output: Pose sequence contains global root translation and joint rotations. Output poses have max duration of 8 sec.

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration:

Runtime Engines: PyTorch

Supported Hardware Microarchitecture Compatibility:

  • NVIDIA Ampere
  • NVIDIA Blackwell
  • NVIDIA Hopper

Supported Operating System(s): Linux

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Model Versions:

This repo corresponds to the ARDY-G1-RP-25FPS-Horizon52 model variant. Please refer to the codebase for installation and usage instructions.

Training, Testing, and Evaluation Datasets:

The model was trained and evaluated using the Bones Rigplay 1 dataset.

Training Dataset:

Data Modality:

  • Text
  • Other: Human Motion Capture

Text Training Data Size: Less than a Billion Tokens

Other Training Data Size: 630 hours of human motion captures

Data Collection Method by dataset: Automatic/Sensors

Labeling Method by dataset: Hybrid: Automated/Human

Properties: Contains optical motion capture data with corresponding text descriptions covering a diverse range of behaviors such as locomotion, everyday activities, and gestures. Motions are clipped to 10 sec long and resampled to the desired FPS for training. An LLM is used to augment the dataset with diverse paraphrases of text labels.

Testing Dataset:

Data Collection Method by dataset: Automatic/Sensors

Labeling Method by dataset: Hybrid: Automated/Human

Properties: 70 hours of motion data held out from training. The test split contains motions from content categories not seen in training.

Evaluation Dataset:

Benchmark Score: See codebase for evaluation results.

Data Collection Method by dataset: Automatic/Sensors

Labeling Method by dataset: Hybrid: Automated/Human

Properties: Same as test dataset.

Inference:

Acceleration Engine: TensorRT

Test Hardware:

  • NVIDIA A100
  • NVIDIA RTX 4090

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

For more detailed information on ethical considerations for this model, please see the Model Card++ Bias, Explainability, Safety & Security, and Privacy Subcards below.

Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.

Bias

Field | Response :---------------------------------------------------------------------------------------------------|:--------------- Participation...

Excerpt shown — open the source for the full document.