Skip to main content

Hyperparameters

Guide to tuning hyperparameters for training experiments.

Config Files​

All hyperparameters are defined in TOML config files under configs/<species>/. Every species commits a stages.toml manifest (schema mesozoic.stage-manifest/v2) that names its stage config files and declares the recipe graph over them. Velociraptor, Brachiosaurus and Dibothrosuchus have three stage configs named by stage number; T-Rex, Compsognathus and the Compsognathus robot have four, named by stage id:

configs/
├── velociraptor/
│ ├── stage1_balance.toml
│ ├── stage2_locomotion.toml
│ ├── stage3_strike.toml
│ └── stages.toml
├── trex/
│ ├── stance.toml
│ ├── recovery.toml
│ ├── locomotion.toml
│ ├── behavior.toml
│ └── stages.toml
├── brachiosaurus/
│ ├── stage1_balance.toml
│ ├── stage2_locomotion.toml
│ ├── stage3_food_reach.toml
│ └── stages.toml
├── dibothrosuchus/
│ ├── stage1_balance.toml
│ ├── stage2_locomotion.toml
│ ├── stage3_snap.toml
│ └── stages.toml
├── compsognathus/
│ ├── stance.toml
│ ├── recovery.toml
│ ├── locomotion.toml
│ ├── behavior.toml
│ └── stages.toml
└── compsognathus_robot/
├── stance.toml
├── recovery.toml
├── locomotion.toml
├── behavior.toml
└── stages.toml

Each stage TOML file contains [stage], [env] and [curriculum] sections plus the algorithm sections it supports: [ppo] and [sac] in every stage file except the T-Rex recovery.toml, which is PPO-only. The loader rejects any other table, [jax] included: the JAX/MJX backend was retired (decision D-D17). stages.toml carries no hyperparameters. It records the manifest schema and, per [[stages]] entry in manifest order: the stage id, its config file, an optional legacy_number (how integer stage references resolve), warm_start_from (the id of an earlier entry the node initialises from; absent means root), deliverable = true (its certified checkpoint is a published policy) and a recipe label (stand, walk or hunt; a label resolves to its deepest deliverable in manifest order). Recipes are derived from the edges, never declared in a second table. Ids are an open vocabulary matching ^[a-z][a-z0-9_]*$, with stance, recovery, locomotion and behavior reserved. See Behavior Recipes.

Per-Stage Hyperparameters​

Each stage has its own [ppo] and [sac] sections. When the curriculum command advances, it loads the next TOML file and re-initialises the algorithm with that stage's settings; the next node's weights come from its declared warm_start_from parent's handoff checkpoint, loaded under initialize_next_stage. Values differ across species and change as experiments evolve, so the stage TOML files are the authoritative source.

Every trained node records a hyperparameters_sha256 in the run block of its stage_config.json: a digest over the stage's [ppo] or [sac] table plus its warmup_ / ramp_ shaping keys, independent of key order and untouched by env kwargs or gate thresholds. A label is recorded beside it when --label TEXT (train or curriculum) or the notebook's RUN_LABEL is set. Both are copied into the run's provenance.deliverables records and tagged onto the W&B run as hp:<12 hex> and label:<text>, which is how two variants of one node are told apart later.

The broad curriculum intent is balance first, locomotion second, and a species-specific simulator task third. Reward weights shift with that intent; do not infer current values from a copied range in this guide. The generated model pages show the current stage names, budgets, and advancement gates read from the same configs.

PPO Parameters​

ParameterDescription
learning_rateInitial network learning rate
learning_rate_endOptional end value for a linear learning-rate schedule
n_stepsSteps collected per rollout buffer
batch_sizeMinibatch size for gradient updates
n_epochsOptimisation epochs per PPO update
gammaReward discount factor
gae_lambdaGeneralized advantage-estimation factor
clip_rangePPO surrogate-objective clipping range
ent_coef / ent_coef_endInitial and optional scheduled entropy coefficients

SAC Parameters​

ParameterDescription
learning_rateNetwork learning rate
batch_sizeReplay-sample batch size
gammaReward discount factor
tauSoft target-update coefficient
ent_coefEntropy coefficient or automatic-tuning mode
buffer_sizeReplay-buffer capacity
train_freqEnvironment-data collection frequency between updates
gradient_stepsGradient updates per training interval

Curriculum Thresholds​

Stage transitions are controlled by the [curriculum] section in each config:

ParameterDescription
timestepsMaximum configured stage budget before advancing
min_avg_rewardMinimum evaluation-window mean reward for early advancement
min_avg_episode_lengthMinimum evaluation-window mean episode length for early advancement
min_avg_forward_velOptional minimum mean forward velocity; enabled when greater than zero (reward_and_length/v1 only)
min_success_rateOptional minimum episode success rate; enabled when greater than zero (reward_and_length/v1 only; retired for the trex hunt, which is gated on the bound below)
min_success_lcbtask_success/v1 only: the one-sided 95% Clopper-Pearson lower bound on per-episode task success must clear this bar (trex hunt: 0.5, provisional — 20/30 clears it, 19/30 does not); min_avg_reward is then a collapse rail, never the gate
min_eval_episodesMinimum episodes required in an evaluation window; defaults to 10 in StageThreshold except for task_success/v1, where it is REQUIRED (the bound's power is a function of the declared n; trex hunt: 30)
required_consecutiveConsecutive evaluations that must satisfy every enabled criterion

For the SB3 curriculum, every enabled criterion must pass in the same evaluation window, the window must contain at least min_eval_episodes episodes (10 by default, the StageThreshold.min_eval_episodes implementation default; the declared value for a task_success/v1 stage), and this result must repeat for required_consecutive evaluations before the stage advances early. If the timesteps budget runs out first, the node's gate_verdict.json records a failure and the curriculum command stops before the next node rather than training it from an uncertified parent. Consult the generated model pages or the TOML files for current values rather than relying on a static table here.

Overriding Hyperparameters from the CLI​

Use --override to change TOML values without editing files — useful for hyperparameter sweeps. Keys use dot notation; values are auto-cast to int, float, or str:

# Override learning rate and entropy coefficient for all stages
python scripts/train_sb3.py train --stage 1 \
--override ppo.learning_rate=1e-3 ppo.ent_coef=0.02 env.alive_bonus=5.0

# Works with curriculum too — applies to ALL stages
python scripts/train_sb3.py curriculum \
--override ppo.learning_rate=2e-4

Supported key prefixes:

PrefixOverrides
ppo.Xppo_kwargs[X]
sac.Xsac_kwargs[X]
env.Xenv_kwargs[X] (reward weights, episode settings)

For stage-scoped overrides, prefix with the stage's legacy number or its id: 1.ppo.learning_rate=3e-4 2.ppo.learning_rate=1e-4, or recovery.ppo.learning_rate=1e-4 on a train --stage recovery run. Plain section.key=value still applies to all stages.

Experimenting on a node that already passes​

When a run is trunked (curriculum --trunk-from, or the notebook's TRUNK_FROM), whether an edited node is reused or retrained depends on which block you edited. The task digest a verdict is bound to covers the effective [env] block, the plant and the push schedule — not [ppo] / [sac], the timesteps budget, or the gate thresholds.

  1. An algorithm-block edit ([ppo], [sac], or a warmup_ / ramp_ shaping key) leaves the digest unchanged, so a trunked run reuses the old certified checkpoint unless the node is the run's target or --retrain-from <id> / RETRAIN_FROM covers it. The reuse prints a warning naming the differing dotted keys (ppo.learning_rate, shaping.warmup_timesteps) and pointing at the retrain knob; it is never a refusal.
  2. An [env] edit changes the digest: the old checkpoint is refused by the reuse rule and the node and everything below it retrain.
  3. A gate-threshold edit leaves the digest unchanged but changes the gate digest the verdict records (gate_sha256, decision D-A22): reuse rule 7 refuses the old verdict naming the thresholds that differ, and a trunked run then trains the node itself. The remedy is a re-judge under the current gate — the notebook JUDGE branch / generate_stage_artifacts, or backfill_gate_verdict.py --force --gate current for a reward_and_length/v1 directory — never a retrain.

A variant is a new run directory: writing into a stage directory that already holds stage_config.json or gate_verdict.json is refused unless the load is an explicit same-stage resume. It is also a trunk for everything it holds: --trunk-from follows a run's own stage directories first and, for a node it only reused, its ancestors/<stage_id>/ record to the run that certified the node, so a variant that reused stance hands stance on to a later run as the original run's checkpoint. See Behavior Recipes.

Tips​

  1. Start from the committed config — Treat it as a reproducible baseline, not a validated optimum
  2. Compare algorithms under the same protocol — The published historical runs do not provide a controlled PPO-versus-SAC comparison
  3. Monitor with W&B — Use --wandb to track per-component rewards across stages
  4. Measure throughput — Compare CPU and GPU performance for your algorithm and parallel-environment count
  5. Increase timesteps for stage 3 — The sparse terminal reward (strike/bite/food) often needs more samples to converge

Tuning Playbook​

Use this section as a symptom-driven reference when a run is misbehaving. Each entry lists the most effective knobs first. Adjust one group at a time so you can attribute outcomes to changes.

Stage 1 — Balance​

SymptomMost likely causeWhat to change
Agent falls immediately (ep length ~30–80, reward < 50)Over-aggressive actions or wobbly initial poseLower ppo.learning_rate (try 3e-5), raise env.posture_weight, raise env.nosedive_weight
Agent hops / drifts across the arena to stay "alive"alive_bonus too large relative to drift/spin penaltiesLower env.alive_bonus to 1.0–1.75, raise env.drift_penalty_weight and env.speed_penalty_weight
Spins in place to maintain upright torsoSpin is cheaper than true balanceRaise env.spin_penalty_weight to 0.1+, add small env.heading_weight (0.1)
Reward plateaus at a mediocre value, no further progressEntropy too low or LR too high for fine-tuningLower LR to 3e-5, raise ppo.ent_coef to 0.01, or switch to cosine schedule via learning_rate_end
Reward is highly variable across seedsInit noise dominatesLower env.reset_noise_scale (default 0.05, try 0.02)
Jerky, unstable joint motionSmoothness under-weightedRaise env.smoothness_weight and env.energy_penalty_weight

Stage 2 — Locomotion​

SymptomMost likely causeWhat to change
Forward velocity stuck near zeroforward_vel_weight too low or alive_bonus dominantRaise env.forward_vel_weight to 1.0–2.0, reduce env.alive_bonus
Crab-walks sideways toward targetNo lateral or heading constraintSet env.lateral_penalty_weight to 0.1, raise env.heading_weight
Walks forward but falls frequentlyBalance reward zeroed too aggressivelyKeep env.posture_weight ≥ 0.3, re-enable mild env.alive_bonus (0.3–0.5)
Unrealistic "ice-skating" gaitSymmetry and smoothness under-weightedRaise env.gait_symmetry_weight, raise env.smoothness_weight
Max forward speed exceeds physical reasonablenessReward uncappedSet env.forward_vel_max (e.g. 3.0 for raptor, 1.5 for brachio)

Stage 3 — Behavior (Strike / Bite / Food Reach)​

The names "Bite" and "Food Reach" refer to simulation success proxies. In the Gym/SB3 environments, T-Rex uses contact from a fixed head geom and has no articulated jaw; Brachiosaurus uses a head-tip distance threshold and does not require physical food contact.

SymptomMost likely causeWhat to change
Never triggers terminal event (strike/bite/food)Sparse-reward exploration stalledRaise ppo.ent_coef to 0.005–0.01, widen ppo.clip_range to 0.15, tighten env.prey_distance_range so the target spawns closer
Lingers near target without triggeringProximity bonuses rewarding hoveringZero env.strike_proximity_weight (or env.food_head_proximity_weight), keep the terminal bonus dominant per the current stage config, and use *_approach_weight as the approach gradient
Forgets locomotion during Stage 3 warm-upReward schedule shift too abruptUse curriculum.warmup_timesteps = 300000, curriculum.ramp_timesteps = 500000, curriculum.warmup_clip_range = 0.02 to anneal changes
Learns the behavior but then regressesOver-entropy or critic driftLower ppo.ent_coef after convergence, narrow ppo.clip_range to 0.1
strike_bonus signal not dominatingDiscounted future alive-reward too largeEnsure env.alive_bonus = 0 in stage 3 and that strike_bonus >> gamma^H · per_step_reward

Algorithm-specific notes​

PPO. Treat the learning_rate / learning_rate_end schedule as a primary tuning lever. Choose n_steps, n_envs, and batch_size together so rollout batches divide cleanly into minibatches; Stable-Baselines3 warns when they do not. Tune clip_range and entropy against the stage's observed stability rather than assuming one fixed progression across species.

SAC. Automatic entropy tuning is available through ent_coef = "auto". Tune train_freq, gradient_steps, and buffer_size together: their useful values depend on the species, stage, parallel-environment count, and memory budget.

A minimal tuning workflow​

  1. Baseline. Run the stage with committed defaults for 2–3 seeds. Record best reward, mean episode length, and any behavioral metrics (success rate, velocity).
  2. Diagnose. If the run fails, match symptoms against the tables above. Do not change more than one group of knobs per run.
  3. Narrow. For promising directions, train 3–5 candidate values as separate runs, each changing one value with --override (see Overriding Hyperparameters from the CLI), and compare them with the baseline under the same protocol.
  4. Promote. Commit the winning values back to the TOML with a trailing comment explaining why (see existing configs for the house style — e.g. # Setting 4 sweep: ...).
  5. Regress-test. Re-run everything below the edited node on the certified trunk before committing: curriculum --trunk-from <certified run> --retrain-from <edited node> --output-dir <new run> retrains that node and every node after it. An [env] edit retrains from that node down automatically, since it changes the task digest. A stance change often degrades the behavior node. The regress-test run reuses the trunk's ancestors above the edited node and can pass them on: a later run trunked from it resolves each reused node through its ancestor record to the run that certified it, one machine-visible run directory away, so later behaviors can trunk from the regress-test run directly.