Skip to main content

Brachiosaurus Model

A sauropod-inspired quadrupedal simulation model designed for four-legged locomotion and head-positioning tasks.

Generated specifications

These values are generated from the environment and compiled MJCF model. “Dynamic mass” is a simulator property, not an anatomical weight estimate.

GaitQuadrupedal
Observation dimension86
Action dimension / actuators30
Generalized coordinates (nq)38
Generalized velocities (nv)37
Compiled dynamic model mass175.3 kg
Plant contract revisionsPolicy r8 · physics r4 · visual r2
Modelenvironments/brachiosaurus/assets/brachiosaurus.xml

Current curriculum configuration

The Stable-Baselines3 budgets and early-advancement gates below come from the current TOML files and can differ from older published runs. A stage also ends when its configured budget is exhausted. The recipe column names the behavior a stage belongs to and whether its checkpoint is a published deliverable; the warm-start column names the stage it starts from.

StageRecipeWarm-start fromObjectiveSB3 configured budgetSB3 early-advancement gate
1 — Balancestand (deliverable)—Learn to stand on four legs without falling6Mreward ≥ 1,040; episode length ≥ 950; ≥ 10 episodes/evaluation; 3 consecutive passes
2 — Locomotionwalk (deliverable)1 — BalanceLearn coordinated quadrupedal walking16Mreward ≥ 100; episode length ≥ 750; avg. velocity ≥ 0.75 m/s; ≥ 10 episodes/evaluation; 3 consecutive passes
3 — Food Reachhunt (deliverable)2 — LocomotionMove the head tip within the configured distance threshold of food12Mreward ≥ 100; task success ≥ 50.0%; ≥ 10 episodes/evaluation; 3 consecutive passes

Backend-specific task-success semantics

  • Stable-Baselines3 / JAX/MJX — Head-tip distance-threshold success: The head-tip site comes within the configured food-reach threshold of the food target while the food-reach bonus is enabled.

Per-deliverable success semantics

  • stand (1 — Balance) · Stable-Baselines3 — Reward-gated stance (reward_and_length/v1): The stance checkpoint clears the reward_and_length/v1 gate: mean evaluation reward at or above the statue-derived collapse rail and a near-full-horizon mean episode length over the required consecutive evaluations. The zero-action statue clears this gate, so stand is labelled by its gate kind here rather than claimed as certified stance quality; stance_quality/v1 waits on shin instrumentation (a kneeling pose currently reads identically to airborne) (plan §4.8).
  • walk (2 — Locomotion) · Stable-Baselines3 — Gated forward velocity (reward_and_length/v1): The locomotion checkpoint clears the reward_and_length/v1 gate: mean forward velocity at or above the stage's configured minimum, with its reward and episode-length floors, over the required consecutive evaluations.
  • hunt (3 — Food Reach) · Stable-Baselines3 / JAX/MJX — Head-tip distance-threshold success: The head-tip site comes within the configured food-reach threshold of the food target while the food-reach bonus is enabled.

Published run summaries

These are historical experiment records. A result is not evidence for the current model revision unless its provenance is explicitly marked current and verified.

PPO · Stable-Baselines3 · 2026-07-18

Provenance: Historical model · unverified · 30 evaluation episodes · Stable-Baselines3 (version not recorded)

StageTrained stepsBest eval rewardAvg. forward velocityTask success (Stable-Baselines3)Passed
1 — balance6M1740.530.01 m/s—passed retired gate (reward gate)
2 — locomotion16M6634.601.42 m/s3.3%passed retired gate (reward gate)
3 — food reach12M1368.130.71 m/s100.0%passed retired gate (reward gate)

Current Stable-Baselines3 definition for this task label: The head-tip site comes within the configured food-reach threshold of the food target while the food-reach bonus is enabled.

Features​

  • First quadrupedal species in the project
  • Articulated 4-segment neck with head control (6 actuators)
  • Front legs longer than rear, producing the model's raised-front posture
  • Simplified column-like legs for support
  • Head-positioning behavior using neck articulation
  • Three published behaviors on one trunk (configs/brachiosaurus/stages.toml): stand (stance), walk (locomotion, warm-started from stance) and hunt (behavior — the food-reach task, warm-started from locomotion); see Behavior Recipes

Unique Characteristics​

Brachiosaurus differs from the bipedal species in the project:

  • Quadrupedal gait instead of bipedal
  • Target-proximity task using the head tip instead of a predatory contact task

Diagnostic Metrics (behavior node, hunt)​

After each training stage the eval command reports:

MetricDescription
mean_forward_velocityAverage forward speed (m/s)
gait_symmetryLeft/right stride symmetry ∈ [0, 1]
stride_frequencyStep frequency (Hz)
cost_of_transportEnergy efficiency (lower is better)
mean_pelvis_heightTorso height stability (m)
success_rateFraction of episodes where the head tip enters the configured distance threshold around the food target
min_prey_distanceClosest approach to food (m, head_food_distance)

Training Videos​

Published videos for Brachiosaurus Altithorax, where available. Artifact provenance is shown on each stage.

STAGE 1
stand · deliverable

Balance

Learn to stand on four legs without falling

PPO · Stable-Baselines3 artifact · historical model · unverified · backend version not recorded. It is not evidence for the current model or stage config.

STAGE 2
walk · deliverable

Locomotion

Learn coordinated quadrupedal walking

PPO · Stable-Baselines3 artifact · historical model · unverified · backend version not recorded. It is not evidence for the current model or stage config.

STAGE 3
hunt · deliverable

Food Reach

Move the head tip within the configured distance threshold of food

No published video is available for this stage.

Usage​

cd environments/brachiosaurus

# View the model
python scripts/view_model.py

# Train the stance node using its current TOML-configured budget
python scripts/train_sb3.py train --stage 1

# The advancing stages (stance, locomotion, behavior) in manifest order
python scripts/train_sb3.py curriculum --algorithm ppo

# Reuse an earlier run's certified stance and walk; train hunt in a fresh directory
python scripts/train_sb3.py curriculum --algorithm ppo \
--trunk-from logs/<earlier_run> --output-dir logs/<new_run>

See Behavior Recipes for the notebook workflow.