MolmoAct2 SO-101 rig fine-tune — anchor rungs (rig_ft_r1, pre-reg PASS)

run rig_ft_r1 (fontaine_so101_rig_ae_r1) — AE-only fine-tune of allenai/MolmoAct2-SO100_101 on the 2 SO-101 rig repos
trainer: their train_lerobot.py (branch fontaine-so101-rig), 2000 steps, global batch 64, AE lr 5e-5, --ft_vlm=false --ft_embedding=none (577M trainable / 5.5B)
data: mcobzarenco/so101_pick_place_clean (7 ep) + _v2 (50 ep), LeRobot v3.0 end-to-end, rig-only q01/q99 norm stats
launched 2026-08-10 17:48:18Z, rc=0 20:27:44Z, ~2.7 GPU-h (gate 12); pre-reg posts/2026-08-10-prereg-molmoact2-rig-finetune.md (+ Amendment 1)
reads: 240 evenly strided rig frames, matched 30-step / 1.0 s window, identical rows every rung (frozen preflight instrument)
CAVEAT (pre-registered): train-frame sanity reads, contaminated by construction — the real eval is on-rig rollouts (runbook sections 3-4)

Rung curve — matched-window MAE vs fine-tune step

Rung summary

checkpointmatched-window MAEmotion corr (min … max)max |step-0 offset|
zero-shot28.9454+0.124 … +0.44779.00
step 5006.7561+0.619 … +0.8511.71
step 10004.6600+0.785 … +0.9231.21
step 15003.5871+0.864 … +0.9571.06
step 20003.2301+0.885 … +0.9650.63

MAE by chunk timestep (log scale) — every rung + anchors

Per-joint motion correlation across rungs (zero line marked; at zero-shot joint 1 sat at +0.22 with a +79-unit step-0 offset — the posture-collapse finding of pre-reg Amendment 1)

Final rung (step 2000) per joint

jointmotion corrstep-0 offsetstep-0 err stdrig q01–q99 span
0 — shoulder_pan+0.8853+0.0431.9190.0
1 — shoulder_lift+0.9649+0.6332.85152.3
2 — elbow_flex+0.9567+0.1452.65140.0
3 — wrist_flex+0.9494+0.2743.34189.0
4 — wrist_roll+0.9465+0.5364.34280.3
5 — gripper+0.8971+0.0542.1136.9

Sample trajectories — 8 of 240 anchor frames, evenly strided. Legend below applies to every panel.