Task Sample · Signature · Pretrain

Marin-Scaling-Ladder

Design one optimizer update rule and run it unchanged across six Marin models, from 550M to 2.545B parameters, against Marin’s AdamH baseline at every rung.

The results

Codex’s PSPR beat AdamH at all six rungs, +0.60% on the combined metric. Claude Code’s RMBT beat it at five, +0.35%, and regressed 0.21% at 837M.

RungModelParamsTokensScoring updateGPUsGPU-h
E0d1152-L12550M2.904B44,317817.6
E1d1408-L15837M3.613B55,125827.6
E2d1536-L16998M4.983B38,014831.0
E3d1792-L181.385B10.560B40,28332302.5
E4d2048-L211.935B14.805B30,00032286.2
E5d2304-L232.545B18.617B35,0001282,932.1
Marin: an open-source framework for the research and development of foundation modelsStanford
Improvement over AdamH, per ladder rungGeometric mean of two ratios: Paloma bits-per-byte and log PPL.Each rung is a separate model, 550M to 2.5B. Higher is better.-1%-0.5%+0.5%+1%0AdamH control−1%: a rung below this fails the submissionCodex · GPT-5.6 — PSPR, +0.60% on averageClaude Code · Opus 5 — RMBT, +0.35% on averageE0550ME1837ME2998ME31.39BE41.94BE52.54BLadder rung and model sizeBetter than AdamH (%)
Combined improvement over AdamH at each model size, using Paloma bits-per-byte and fixed-window training loss. Higher is better; zero matches the control.
Best six-rung improvement during the search● new best · × rejected / not retained · ○ trial · ◇ partial result · ▪ stopped0+0.2%+0.4%+0.6%01020304050607080Cumulative runtime (hours)Best complete six-rung improvement (%)Codex · GPT-5.6 · FCPR · 3.60 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 3.6h 'FCPR passes its second preregistered screen' → full ladder frozenCodex · GPT-5.6 · TARF · 10.90 h Plan not launched Best complete six-rung improvement at this time: +0.00% 9.2h preregistered contingent on T-FCPR passing; T-FCPR rejected 10.9h → never implemented or launched Nearby events on the same curve: Codex · GPT-5.6 · T-FCPR · 10.91 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E2, update 5,000: geometric gain 0.99955754 versus FCPR, despite passing the macro gate. E0 had narrowly passed at 1.00010275. Trajectory steps: 1527, 1679, 1680Codex · GPT-5.6 · FCPR · 12.30 h Paused Best complete six-rung improvement at this time: +0.00% 12.3h 'FCPR E0 25k comparison is negative … the pause is operational, not a substituted final verdict'; never resumed, never formally rejected Nearby events on the same curve: Codex · GPT-5.6 · TA-FCPR · 12.22 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus FCPR: geometric gain 0.98204696 on E0 and 0.98710723 on E2; both macro safety gates fail. Trajectory steps: 1852, 1856, 1912, 1915Codex · GPT-5.6 · OEPR · 15.40 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 15.4h 'OEPR is promoted to a frozen full ladder' Nearby events on the same curve: Codex · GPT-5.6 · SCPR · 15.49 h Not selected Best complete six-rung improvement at this time: +0.00% Not selected for a full ladder: both 5,000-update screens improved RPM (geometric gain 1.00074440 on E0; 1.00252280 on E2), but OEPR improved more. This is selection rejection, not a negative score. Trajectory steps: 2393, 2470, 2471 Codex · GPT-5.6 · MPR · 15.73 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99291949 on E0 and 0.99619747 on E2. E0 also fails the macro safety gate. Trajectory steps: 2470, 2471, 2538, 2539, 2548Codex · GPT-5.6 · AGPR · 17.00 h Full ladder launched Best complete six-rung improvement at this time: +0.29% 16.9h passes E2 gate; 17.0h 'AGPR's six-rung ladder is submitted'Codex · GPT-5.6 · PSPR · 26.20 h Full ladder launched Best complete six-rung improvement at this time: +0.29% 26.1h E2 screen passes; 26.2h 'launching the six scored contracts as opt-014-pspr-full'Codex · GPT-5.6 · OEPR · 36.60 h Partial / projected result Best complete six-rung improvement at this time: +0.29% Partial / projected score: +0.35% (projected, never completed); the completed best is unchanged. 36.6h 'OEPR projects to reward 1.00354 with every safety gate passing—stronger than the staged RPM—but it is not stageable until its exact-final checkpoints and reload audit finish'; never mentioned after the 52.1h restartCodex · GPT-5.6 · CPSR · 55.20 h Full ladder launched Best complete six-rung improvement at this time: +0.60% 55.1h passes E2; 55.2h 'frozen CPSR full ladder is launched as opt-017-cpsr-full'Codex · GPT-5.6 · AGPR → control · 56.10 h Reduced control Best complete six-rung improvement at this time: +0.60% 56.1h AGPR E5 completes; 55.8h 'exact-35k PSPR-vs-AGPR gain 1.0003199' — AGPR is PSPR's reduced control; no six-rung reward ever statedCodex · GPT-5.6 · CPSR · 66.40 h Stopped Best complete six-rung improvement at this time: +0.60% 66.4h 'CPSR is now formally recorded as incomplete due to infrastructure and trusted launch cutoff'Codex · GPT-5.6 · AGST · 0.81 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E0, update 5,000: geometric gain 0.97835421 versus reduced AdamH (Paloma micro BPB 1.46582524 vs 1.43076432; fixed-window loss 4.07710351 vs 3.99814064). Its six-rung ladder was stopped; no complete-ladder score. Trajectory steps: 119, 120, 143Codex · GPT-5.6 · CA-TFCPR · 9.76 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E0, update 5,000: geometric gain 0.97556756 versus T-FCPR; macro BPB fails the 1% non-regression gate. E2 evaluation was interrupted, so it has no completed E2 screen. Trajectory steps: 1527, 1528, 1556, 1561Codex · GPT-5.6 · T-FCPR · 10.91 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E2, update 5,000: geometric gain 0.99955754 versus FCPR, despite passing the macro gate. E0 had narrowly passed at 1.00010275. Trajectory steps: 1527, 1679, 1680 Nearby events on the same curve: Codex · GPT-5.6 · TARF · 10.90 h Plan not launched Best complete six-rung improvement at this time: +0.00% 9.2h preregistered contingent on T-FCPR passing; T-FCPR rejected 10.9h → never implemented or launchedCodex · GPT-5.6 · TA-FCPR · 12.22 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus FCPR: geometric gain 0.98204696 on E0 and 0.98710723 on E2; both macro safety gates fail. Trajectory steps: 1852, 1856, 1912, 1915 Nearby events on the same curve: Codex · GPT-5.6 · FCPR · 12.30 h Paused Best complete six-rung improvement at this time: +0.00% 12.3h 'FCPR E0 25k comparison is negative … the pause is operational, not a substituted final verdict'; never resumed, never formally rejectedCodex · GPT-5.6 · SV-FCPR · 13.40 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus FCPR: geometric gain 0.99957022 on E0 and 0.99866946 on E2. Both macro safety gates pass, but neither screen improves the scored pair. Trajectory steps: 2005, 2105, 2122Codex · GPT-5.6 · GPRM · 14.32 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99965743 on E0 and 0.99957609 on E2; both macro gates pass. Trajectory steps: 2214, 2218, 2282, 2283Codex · GPT-5.6 · SCPR · 15.49 h Not selected Best complete six-rung improvement at this time: +0.00% Not selected for a full ladder: both 5,000-update screens improved RPM (geometric gain 1.00074440 on E0; 1.00252280 on E2), but OEPR improved more. This is selection rejection, not a negative score. Trajectory steps: 2393, 2470, 2471 Nearby events on the same curve: Codex · GPT-5.6 · OEPR · 15.40 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 15.4h 'OEPR is promoted to a frozen full ladder' Codex · GPT-5.6 · MPR · 15.73 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99291949 on E0 and 0.99619747 on E2. E0 also fails the macro safety gate. Trajectory steps: 2470, 2471, 2538, 2539, 2548Codex · GPT-5.6 · MPR · 15.73 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99291949 on E0 and 0.99619747 on E2. E0 also fails the macro safety gate. Trajectory steps: 2470, 2471, 2538, 2539, 2548 Nearby events on the same curve: Codex · GPT-5.6 · OEPR · 15.40 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 15.4h 'OEPR is promoted to a frozen full ladder' Codex · GPT-5.6 · SCPR · 15.49 h Not selected Best complete six-rung improvement at this time: +0.00% Not selected for a full ladder: both 5,000-update screens improved RPM (geometric gain 1.00074440 on E0; 1.00252280 on E2), but OEPR improved more. This is selection rejection, not a negative score. Trajectory steps: 2393, 2470, 2471Codex · GPT-5.6 · CTPR · 34.90 h Rejected / failed Best complete six-rung improvement at this time: +0.29% Rejected at both 5,000-update screens versus PSPR. The later exact re-audit reports geometric gain 0.99970355 on E0 and 0.99900770 on E2. Earlier commentary used slightly different comparisons; these are the re-audited screen-to-screen values. Trajectory steps: 5258, 5326, 5881, 5885Codex · GPT-5.6 · PSGR · 53.51 h Rejected / failed Best complete six-rung improvement at this time: +0.60% Rejected by the two-screen promotion rule: at update 5,000 versus PSPR, E0 passes narrowly (1.00005941), but E2 is below one (0.99989665); macro safety passes. Trajectory steps: 6034, 6084, 6086Codex · GPT-5.6 · RPM · 16.18 h New measured best Best complete six-rung improvement at this time: +0.29% 16.2h 'RPM is now staged as the first eligible incumbent: exact-step reward 1.0029320791' Trajectory steps: 2612Codex · GPT-5.6 · PSPR · 52.19 h New measured best Best complete six-rung improvement at this time: +0.60% 52.2h 'PSPR's direct-step reward is 1.005988, a substantial improvement over the staged 1.002932 incumbent'; 56.1h 'PSPR is now the staged incumbent at reward 1.0059877526, replacing RPM' Trajectory steps: 5879RPM +0.29% — polar residual mixingPSPR +0.60% — principal-sine polar transportClaude Code · Opus 5 · RMBT v1 · 2.20 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 2.2h 'Ladder launched from a single frozen source' Nearby events on the same curve: Claude Code · Opus 5 · Cautious gating · 1.77 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: fixed-window training loss 3.84154 versus 3.75773 for the mechanism-off control (+2.231%, worse). This is a single-rung screen, not a six-rung score. Trajectory steps: 142 Claude Code · Opus 5 · RMBT v1: cost overrun · 2.78 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Initial ladder stopped before completion: E0/E1/E2 projected costs were 4.9%, 10.5%, and 9.1% above the AdamH manifest. E0 reached 5,000 updates and E5 reached zero; no complete six-rung score was measured. Trajectory steps: 196, 200, 201, 202Claude Code · Opus 5 · RMBT v2 · 3.00 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 3.0h 'Ladder relaunched with the optimized source' (retraction rewritten as two Triton kernels) Nearby events on the same curve: Claude Code · Opus 5 · RMBT v1: cost overrun · 2.78 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Initial ladder stopped before completion: E0/E1/E2 projected costs were 4.9%, 10.5%, and 9.1% above the AdamH manifest. E0 reached 5,000 updates and E5 reached zero; no complete six-rung score was measured. Trajectory steps: 196, 200, 201, 202Claude Code · Opus 5 · E5 relaunch · 6.80 h Infrastructure retry Best complete six-rung improvement at this time: +0.00% 6.7h 'E5 failed with zero updates' (NCCL mixed RoCE/IB); 6.8h relaunched as opt-002-rmbt-e5bClaude Code · Opus 5 · preview · 17.80 h Partial / projected result Best complete six-rung improvement at this time: +0.00% Partial / projected score: +0.29% (E5 still at 15k); the completed best is unchanged. 17.8h 'Six-rung preview reward: 1.00292' (E5 not yet at its 35,510 scoring update)Claude Code · Opus 5 · exhaustion · 59.23 h Stopped Best complete six-rung improvement at this time: +0.35% 59.2h 'Result — the optimizer works, but the ladder cannot be staged … No submission.json is producible — I am reporting exhaustion' Trajectory steps: 2297 Nearby events on the same curve: Claude Code · Opus 5 · RMBT final validation failed · 59.21 h Rejected / failed Best complete six-rung improvement at this time: +0.35% The handoff staging check again failed endpoint noninferiority and mechanism ablation on the same measured ladder. The session ended at 59.22549 hours (step 2297, 2026-09-03T05:06:48.820Z) reporting exhaustion and no submission.json. Trajectory steps: 2294, 2297Claude Code · Opus 5 · Cautious gating · 1.77 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: fixed-window training loss 3.84154 versus 3.75773 for the mechanism-off control (+2.231%, worse). This is a single-rung screen, not a six-rung score. Trajectory steps: 142 Nearby events on the same curve: Claude Code · Opus 5 · RMBT v1 · 2.20 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 2.2h 'Ladder launched from a single frozen source'Claude Code · Opus 5 · RMBT v1: cost overrun · 2.78 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Initial ladder stopped before completion: E0/E1/E2 projected costs were 4.9%, 10.5%, and 9.1% above the AdamH manifest. E0 reached 5,000 updates and E5 reached zero; no complete six-rung score was measured. Trajectory steps: 196, 200, 201, 202 Nearby events on the same curve: Claude Code · Opus 5 · RMBT v1 · 2.20 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 2.2h 'Ladder launched from a single frozen source' Claude Code · Opus 5 · RMBT v2 · 3.00 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 3.0h 'Ladder relaunched with the optimized source' (retraction rewritten as two Triton kernels)Claude Code · Opus 5 · Vocabulary redistribution · 3.93 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: applying trust redistribution to the vocabulary matrix raised fixed-window training loss to 3.82344 versus 3.75708 for the same-kernel mechanism-off control (+1.766%, worse). No six-rung score. Trajectory steps: 259 Nearby events on the same curve: Claude Code · Opus 5 · Overcorrection (alpha=2) · 4.05 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: alpha=2 raised fixed-window training loss to 3.77349 versus 3.75708 for the same-kernel mechanism-off control (+0.437%, worse). No six-rung score. Trajectory steps: 269Claude Code · Opus 5 · Overcorrection (alpha=2) · 4.05 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: alpha=2 raised fixed-window training loss to 3.77349 versus 3.75708 for the same-kernel mechanism-off control (+0.437%, worse). No six-rung score. Trajectory steps: 269 Nearby events on the same curve: Claude Code · Opus 5 · Vocabulary redistribution · 3.93 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: applying trust redistribution to the vocabulary matrix raised fixed-window training loss to 3.82344 versus 3.75708 for the same-kernel mechanism-off control (+1.766%, worse). No six-rung score. Trajectory steps: 259Claude Code · Opus 5 · RMBT staging failed · 41.11 h Rejected / failed Best complete six-rung improvement at this time: +0.35% The already measured RMBT ladder failed two staging guards: E1 cost-matched endpoint ratio 1.00614 exceeded the 1.005 limit, and E5 cost-matched mechanism-ablation gain 0.99908 fell below 1. This did not produce a new model score. Trajectory steps: 1589, 1590Claude Code · Opus 5 · RMBT staging retry failed · 53.51 h Rejected / failed Best complete six-rung improvement at this time: +0.35% After supplementary controls, the same frozen RMBT ladder still failed E1 endpoint noninferiority and mechanism ablation. The full candidate checkpoints and measured score did not change; no submission.json was produced. Trajectory steps: 2057, 2058, 2062Claude Code · Opus 5 · RMBT final validation failed · 59.21 h Rejected / failed Best complete six-rung improvement at this time: +0.35% The handoff staging check again failed endpoint noninferiority and mechanism ablation on the same measured ladder. The session ended at 59.22549 hours (step 2297, 2026-09-03T05:06:48.820Z) reporting exhaustion and no submission.json. Trajectory steps: 2294, 2297 Nearby events on the same curve: Claude Code · Opus 5 · exhaustion · 59.23 h Stopped Best complete six-rung improvement at this time: +0.35% 59.2h 'Result — the optimizer works, but the ladder cannot be staged … No submission.json is producible — I am reporting exhaustion' Trajectory steps: 2297Claude Code · Opus 5 · RMBT · 29.18 h New measured best Best complete six-rung improvement at this time: +0.35% 29.2h 'All six rungs are now at their exact scoring updates. Step-matched reward = 1.00355'; report.html: 'Exhaustion reported; no staging … two failed cost-based guards' Trajectory steps: 1170RMBT +0.35% — radius-matched block trustCodex · GPT-5.6Claude Code · Opus 5Marin Adam Baseline
Research history: the best complete six-rung score reached so far. Hover over a point for the experiment and its outcome.

The task

Invent an explicit new update rule: its equations, state and parameter grouping. Retuning learning rates does not count. The same implementation runs at every rung.

Reference baseline

Marin’s locked AdamH control at each rung, with the same model, data order and recipe. E5 trains on 128 H100s.

What the agent may change

  • Gradient transformation, preconditioner and update geometry.
  • Optimizer state, parameter grouping and internal schedules.
  • Optimizer kernels, communication and checkpoint serialization.

What stays fixed

  • Model, initialization, objective and tokenizer.
  • Data and order, tokens, batch and precision.
  • Evaluation and launchers.

Evaluation

At each rung’s scoring update, the geometric mean of the AdamH-to-candidate ratios for Paloma bits-per-byte and log PPL over a fixed 16.8M-token window, then the geometric mean across the six rungs. A regression over 1% at any rung fails the submission.

Agent runs

Codex · GPT-5.6 Sol

17 hypotheses. PSPR, principal-sine polar transport, beat AdamH at every rung. At 2.545B, Paloma went from 0.9088 to 0.9018 bits-per-byte and training loss from 2.6040 to 2.5852.

PSPR measurements by rung
RungPaloma BPBTraining loss
E01.06813.0573
E11.03542.9755
E21.00312.8727
E30.95132.7270
E41.03662.9435
E50.90182.5852

Claude Code · Opus 5

RMBT, radius-matched block trust, rescales each hidden unit’s update relative to its weights. At 2.545B it reached 0.9020 bits-per-byte and 2.5847 training loss; extending the rule to vocabulary rows made loss worse.

RMBT measurements by rung
RungPaloma BPBTraining loss
E01.07493.0734
E11.03862.9841
E21.00572.8788
E30.95352.7319
E41.03802.9477
E50.90202.5847