Design verifiable RL tasks that improve Qwen3.5-122B-A10B across math, multiple-choice reasoning, instruction following, long context and coding at once.
Results
Five-benchmark geometric mean 71.51 with the agent’s tasks against 71.37 with the baseline tasks. Coding gains the most, +2.45 points.
| Benchmark | Original | Baseline | Agent | Δ |
|---|---|---|---|---|
| PolyMath Math | 68.06 | 68.40 | 68.21 | +0.16 |
| MMLU-Pro MCQA | 86.30 | 86.29 | 86.44 | +0.14 |
| IFBench Instruction following | 67.01 | 67.69 | 67.35 | +0.34 |
| LongBench V2 Long context | 64.02 | 65.21 | 64.02 | 0.00 |
| LiveCodeBench V6 Coding | 72.06 | 71.09 | 73.54 | +1.48 |
| Geometric mean of all five | 71.09 | 71.37 | 71.51 | +0.42 |
The task
Write prompts with checkable answers, instruction constraints or executable tests: 2,560 RL tasks, 512 per capability, replacing the baseline task pool. The shared base, the starting model and the training recipe stay fixed.
Reference baseline
A shared base plus 2,560 baseline tasks drawn from DAPO, Nemotron, multi-constraint instructions, HotpotQA and Open-R1 Codeforces. Both arms train 40 waves from the same post-trained checkpoint.
What the agent may change
Prompts and labels, task generators, selection and deduplication.
What stays fixed
Base data, starting model, training recipe, reward rules, benchmarks and scorers.
Evaluation
PolyMath, MMLU-Pro, IFBench, LongBench V2 and LiveCodeBench V6, summarized as their geometric mean.
Research record
Research Agent · synthetic task design
The agent probed the starting model to find prompts at the edge of its ability, kept groups with both successes and failures, and froze 2,560 records, 512 per capability, at the 24-hour deadline.
Examples of generated tasks
| Capability | Example |
|---|---|
| Math | Add sixteen 13-digit numbers; the answer is recomputed from the prompt. |
| Multiple-choice reasoning | Choose among arithmetic answers that share modular checks and their final six digits. |
| Instruction following | Write with an exact paragraph count, a required first word and repeated-keyword constraints. |
| Long context | Combine sixteen target readings scattered across a 300-entry station log. |
| Coding | Implement nine ordered string-normalization operations, checked by 24 tests. |
Sources: release contract · operator excerpts · recorded measurements