Design one optimizer update rule and run it unchanged across six Marin models, from 550M to 2.545B parameters, against Marin’s AdamH baseline at every rung.
The results
Codex’s PSPR beat AdamH at all six rungs, +0.60% on the combined metric. Claude Code’s RMBT beat it at five, +0.35%, and regressed 0.21% at 837M.
| Rung | Model | Params | Tokens | Scoring update | GPUs | GPU-h |
|---|---|---|---|---|---|---|
| E0 | d1152-L12 | 550M | 2.904B | 44,317 | 8 | 17.6 |
| E1 | d1408-L15 | 837M | 3.613B | 55,125 | 8 | 27.6 |
| E2 | d1536-L16 | 998M | 4.983B | 38,014 | 8 | 31.0 |
| E3 | d1792-L18 | 1.385B | 10.560B | 40,283 | 32 | 302.5 |
| E4 | d2048-L21 | 1.935B | 14.805B | 30,000 | 32 | 286.2 |
| E5 | d2304-L23 | 2.545B | 18.617B | 35,000 | 128 | 2,932.1 |
The task
Invent an explicit new update rule: its equations, state and parameter grouping. Retuning learning rates does not count. The same implementation runs at every rung.
Reference baseline
Marin’s locked AdamH control at each rung, with the same model, data order and recipe. E5 trains on 128 H100s.
What the agent may change
- Gradient transformation, preconditioner and update geometry.
- Optimizer state, parameter grouping and internal schedules.
- Optimizer kernels, communication and checkpoint serialization.
What stays fixed
- Model, initialization, objective and tokenizer.
- Data and order, tokens, batch and precision.
- Evaluation and launchers.
Evaluation
At each rung’s scoring update, the geometric mean of the AdamH-to-candidate ratios for Paloma bits-per-byte and log PPL over a fixed 16.8M-token window, then the geometric mean across the six rungs. A regression over 1% at any rung fails the submission.
Agent runs
Codex · GPT-5.6 Sol
17 hypotheses. PSPR, principal-sine polar transport, beat AdamH at every rung. At 2.545B, Paloma went from 0.9088 to 0.9018 bits-per-byte and training loss from 2.6040 to 2.5852.
PSPR measurements by rung
| Rung | Paloma BPB | Training loss |
|---|---|---|
| E0 | 1.0681 | 3.0573 |
| E1 | 1.0354 | 2.9755 |
| E2 | 1.0031 | 2.8727 |
| E3 | 0.9513 | 2.7270 |
| E4 | 1.0366 | 2.9435 |
| E5 | 0.9018 | 2.5852 |
Claude Code · Opus 5
RMBT, radius-matched block trust, rescales each hidden unit’s update relative to its weights. At 2.545B it reached 0.9020 bits-per-byte and 2.5847 training loss; extending the rule to vocabulary rows made loss worse.
RMBT measurements by rung
| Rung | Paloma BPB | Training loss |
|---|---|---|
| E0 | 1.0749 | 3.0734 |
| E1 | 1.0386 | 2.9841 |
| E2 | 1.0057 | 2.8788 |
| E3 | 0.9535 | 2.7319 |
| E4 | 1.0380 | 2.9477 |
| E5 | 0.9020 | 2.5847 |