Skip to main content

Velociraptor Model

A dinosaur-inspired bipedal simulation model with actuated sickle-claw geoms, trained for locomotion and target-contact tasks. Its manifest (configs/velociraptor/stages.toml) declares three published behaviors on one trunk: stance (stand), locomotion (walk, warm-started from stance) and behavior (hunt — the strike task, warm-started from locomotion). See Behavior Recipes.

Generated specifications

These values are generated from the environment and compiled MJCF model. “Dynamic mass” is a simulator property, not an anatomical weight estimate.

GaitBipedal
Observation dimension70
Action dimension / actuators22
Generalized coordinates (nq)31
Generalized velocities (nv)30
Compiled dynamic model mass13.5 kg
Plant contract revisionsPolicy r10 · physics r2 · visual r3
Modelenvironments/velociraptor/assets/raptor.xml

Current curriculum configuration

The Stable-Baselines3 budgets and early-advancement gates below come from the current TOML files and can differ from older published runs. A stage also ends when its configured budget is exhausted. The recipe column names the behavior a stage belongs to and whether its checkpoint is a published deliverable; the warm-start column names the stage it starts from.

StageRecipeWarm-start fromObjectiveSB3 configured budgetSB3 early-advancement gate
1 — Balancestand (deliverable)—Learn to stand and balance without falling6Mreward ≥ 1,050; episode length ≥ 950; ≥ 10 episodes/evaluation; 3 consecutive passes
2 — Locomotionwalk (deliverable)1 — BalanceLearn forward walking/running8Mreward ≥ 100; episode length ≥ 750; avg. velocity ≥ 2 m/s; ≥ 10 episodes/evaluation; 3 consecutive passes
3 — Strikehunt (deliverable)2 — LocomotionSprint and strike prey with sickle claw12Mreward ≥ 100; task success ≥ 50.0%; ≥ 10 episodes/evaluation; 3 consecutive passes

Backend-specific task-success semantics

  • Stable-Baselines3 — Sickle-claw contact success: A left or right sickle-claw geom contacts the prey geom while the strike reward is enabled.
  • JAX/MJX — Sickle-claw proximity success: Either claw-tip site comes within 0.20 m of the prey target position while the strike bonus is enabled; physical geom contact is not required.

Per-deliverable success semantics

  • stand (1 — Balance) · Stable-Baselines3 — Reward-gated stance (reward_and_length/v1): The stance checkpoint clears the reward_and_length/v1 gate: mean evaluation reward at or above the statue-derived collapse rail and a near-full-horizon mean episode length over the required consecutive evaluations. The zero-action statue clears this gate, so stand is labelled by its gate kind here rather than claimed as certified stance quality; stance_quality/v1 waits on a foot-sensor repair (the single toe site reads about 55% of true load) (plan §4.8).
  • walk (2 — Locomotion) · Stable-Baselines3 — Gated forward velocity (reward_and_length/v1): The locomotion checkpoint clears the reward_and_length/v1 gate: mean forward velocity at or above the stage's configured minimum, with its reward and episode-length floors, over the required consecutive evaluations.
  • hunt (3 — Strike) · Stable-Baselines3 — Sickle-claw contact success: A left or right sickle-claw geom contacts the prey geom while the strike reward is enabled.
  • hunt (3 — Strike) · JAX/MJX — Sickle-claw proximity success: Either claw-tip site comes within 0.20 m of the prey target position while the strike bonus is enabled; physical geom contact is not required.

Published run summaries

These are historical experiment records. A result is not evidence for the current model revision unless its provenance is explicitly marked current and verified.

PPO · Stable-Baselines3 · 2026-03-15

Provenance: Historical model · unverified · evaluation episode count not recorded · Stable-Baselines3 (version not recorded)

StageTrained stepsBest eval rewardAvg. forward velocityTask success (Stable-Baselines3)Passed
1 — balance6M1964.430.11 m/s—passed retired gate (reward gate)
2 — locomotion8M2678.683.47 m/s—passed retired gate (reward gate)
3 — strike8M1366.192.02 m/s93.3%passed retired gate (reward gate)

Current Stable-Baselines3 definition for this task label: A left or right sickle-claw geom contacts the prey geom while the strike reward is enabled.

SAC · Stable-Baselines3 · 2026-03-21

Provenance: Historical model · unverified · evaluation episode count not recorded · Stable-Baselines3 (version not recorded)

StageTrained stepsBest eval rewardAvg. forward velocityTask success (Stable-Baselines3)Passed
1 — balance6M970.19-0.64 m/s—passed retired gate (reward gate)
2 — locomotion8M2078.622.91 m/s—passed retired gate (reward gate)
3 — strike8M1195.431.63 m/s90.0%passed retired gate (reward gate)

Current Stable-Baselines3 definition for this task label: A left or right sickle-claw geom contacts the prey geom while the strike reward is enabled.

Anatomy​

  • Torso - Elongated body with simplified head/neck
  • Legs - Digitigrade hind limbs (hip pitch/roll, knee, ankle, 2 toe digits per leg)
  • Sickle claws - Actuated claw geom on digit 2 of each foot
  • Arms - Stub forelimbs with shoulder pitch/roll actuators
  • Tail - 5-segment counterbalance (4 actuated segments)

Diagnostic Metrics (behavior node, hunt)​

After each training stage the eval command reports:

MetricDescription
mean_forward_velocityAverage forward speed (m/s)
gait_symmetryLeft/right stride symmetry ∈ [0, 1]
stride_frequencyStep frequency (Hz)
cost_of_transportEnergy efficiency (lower is better)
mean_pelvis_heightUpright stability (m)
mean_heading_alignmentcos θ toward prey ∈ [-1, 1]
success_rateFraction of episodes with successful sickle-claw target contact
min_prey_distanceClosest approach to prey (m)

Training Videos​

Published videos for Velociraptor Mongoliensis, where available. Artifact provenance is shown on each stage.

STAGE 1
stand · deliverable

Balance

Learn to stand and balance without falling

PPO · Stable-Baselines3 artifact · historical model · unverified · backend version not recorded. It is not evidence for the current model or stage config.

STAGE 2
walk · deliverable

Locomotion

Learn forward walking/running

PPO · Stable-Baselines3 artifact · historical model · unverified · backend version not recorded. It is not evidence for the current model or stage config.

STAGE 3
hunt · deliverable

Strike

Sprint and strike prey with sickle claw

PPO · Stable-Baselines3 artifact · historical model · unverified · backend version not recorded. It is not evidence for the current model or stage config.

Usage​

cd environments/velociraptor

# View the model
python scripts/view_model.py

# Train the stance node using its current TOML-configured budget
python scripts/train_sb3.py train --stage 1

# The advancing stages (stance, locomotion, behavior) in manifest order
python scripts/train_sb3.py curriculum --algorithm ppo

# Reuse an earlier run's certified stance and walk; train hunt in a fresh directory
python scripts/train_sb3.py curriculum --algorithm ppo \
--trunk-from logs/<earlier_run> --output-dir logs/<new_run>

See Behavior Recipes for the notebook workflow.