Task Sample · Signature · Posttrain

Qwen-122B-RL-Merge

Design verifiable RL tasks that improve Qwen3.5-122B-A10B across math, multiple-choice reasoning, instruction following, long context and coding at once.

Results

Five-benchmark geometric mean 71.51 with the agent’s tasks against 71.37 with the baseline tasks. Coding gains the most, +2.45 points.

Change the training data; keep the experiment fixed Both runs start from Qwen3.5-122B-A10B. The task package supplies 2,560 shared base tasks and 2,560 baseline tasks. The agent replaces the baseline tasks with 2,560 newly generated, verifiable RL tasks designed in 24 hours, 512 per capability. Both models train on 5,120 tasks with the same RL settings, then take the same five held-out benchmarks. Synthesized RL Data Post-trained Qwen3.5-122B-A10B Baseline Agent RL tasks Ref task samples Ref task samples + 2,560 human baseline tasks + 2,560 newly synthesized RL tasks Both pools from the task package 512 per capability Train with the same RL settings 5 held-out benchmarks
The intended experiment: replace the intervention data while holding the shared base, starting model and RL recipe fixed.
BenchmarkOriginalBaselineAgentΔ
PolyMath
Math
68.0668.4068.21+0.16
MMLU-Pro
MCQA
86.3086.2986.44+0.14
IFBench
Instruction following
67.0167.6967.35+0.34
LongBench V2
Long context
64.0265.2164.020.00
LiveCodeBench V6
Coding
72.0671.0973.54+1.48
Geometric mean of all five71.0971.3771.51+0.42
Release reference scores (%), using the same table as the homepage. Δ is Agent minus Original, in percentage points. These are reported reference results, not a new reproduced training run.

The task

Write prompts with checkable answers, instruction constraints or executable tests: 2,560 RL tasks, 512 per capability, replacing the baseline task pool. The shared base, the starting model and the training recipe stay fixed.

Reference baseline

A shared base plus 2,560 baseline tasks drawn from DAPO, Nemotron, multi-constraint instructions, HotpotQA and Open-R1 Codeforces. Both arms train 40 waves from the same post-trained checkpoint.

What the agent may change

Prompts and labels, task generators, selection and deduplication.

What stays fixed

Base data, starting model, training recipe, reward rules, benchmarks and scorers.

Evaluation

PolyMath, MMLU-Pro, IFBench, LongBench V2 and LiveCodeBench V6, summarized as their geometric mean.

Research record

Research Agent · synthetic task design

The agent probed the starting model to find prompts at the edge of its ability, kept groups with both successes and failures, and froze 2,560 records, 512 per capability, at the 24-hour deadline.

Examples of generated tasks
CapabilityExample
MathAdd sixteen 13-digit numbers; the answer is recomputed from the prompt.
Multiple-choice reasoningChoose among arithmetic answers that share modular checks and their final six digits.
Instruction followingWrite with an exact paragraph count, a required first word and repeated-keyword constraints.
Long contextCombine sixteen target readings scattered across a 300-entry station log.
CodingImplement nine ordered string-normalization operations, checked by 24 tests.

Sources: release contract · operator excerpts · recorded measurements