Task Sample · Posttrain

Kev Decision Architecture

Kev is a Jev-inspired small neural model that makes decisions directly, without generating text. Improve its decision quality and confidence estimates on fixed training data.

Best score per runDecision probability score (0–1)
00.316
Model · harnessBestSubmissionsRuntime
◎ GPT-5.6 SolCodex · xhigh0.292334912.01 h
✳ Claude Opus 5Claude Code · max0.289552212.01 h

Absolute decision probability score, exp(−macro NLL); higher is better. Knowledge and other decision sources each receive half the NLL weight. The dashed 0.25186 reference is validation measured. These are the best submitted checkpoints, not necessarily the agents' later workspace state; each run used a 12-hour research budget.

Architecture illustration: shared context feeds three isolated question branches, each returning an illustrative probability distribution.
Source · Jev’s Architecture Unmasked
Archer Hume

The climb

Dots mark accepted Judge results, and each step line tracks that agent's best submitted score so far. The dashed line is the untouched reference measured during task validation. The 71 recorded submissions include one rejected Claude submission, shown as a cross; its next submission reproduced the best score. Claude's first Judge result arrives around 3.1 hours into the run.

0.2390.2550.2710.2860.302Baseline · 0.2518604812.1Elapsed time (h)Decision probability scoreGPT-5.6 Sol · submission 1 · 0.24868 · 0.265 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0.27539 · 0.574 h, Judge result recordedGPT-5.6 Sol · submission 3 · 0.28177 · 1.14 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0.28111 · 1.45 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.27361 · 1.78 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.27742 · 2.54 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.28177 · 2.61 h, Judge result recordedGPT-5.6 Sol · submission 8 · 0.28460 · 3.3 h, Judge result recordedGPT-5.6 Sol · submission 9 · 0.28694 · 3.49 h, Judge result recordedGPT-5.6 Sol · submission 10 · 0.28669 · 3.53 h, Judge result recordedGPT-5.6 Sol · submission 11 · 0.28705 · 3.56 h, Judge result recordedGPT-5.6 Sol · submission 12 · 0.28712 · 3.6 h, Judge result recordedGPT-5.6 Sol · submission 13 · 0.28717 · 3.63 h, Judge result recordedGPT-5.6 Sol · submission 14 · 0.28670 · 4.01 h, Judge result recordedGPT-5.6 Sol · submission 15 · 0.28691 · 4.38 h, Judge result recordedGPT-5.6 Sol · submission 16 · 0.28750 · 4.82 h, Judge result recordedGPT-5.6 Sol · submission 17 · 0.28777 · 5.23 h, Judge result recordedGPT-5.6 Sol · submission 18 · 0.28789 · 5.27 h, Judge result recordedGPT-5.6 Sol · submission 19 · 0.28790 · 5.31 h, Judge result recordedGPT-5.6 Sol · submission 20 · 0.28855 · 5.77 h, Judge result recordedGPT-5.6 Sol · submission 21 · 0.28879 · 5.84 h, Judge result recordedGPT-5.6 Sol · submission 22 · 0.28887 · 5.9 h, Judge result recordedGPT-5.6 Sol · submission 23 · 0.28863 · 6.03 h, Judge result recordedGPT-5.6 Sol · submission 24 · 0.28872 · 6.09 h, Judge result recordedGPT-5.6 Sol · submission 25 · 0.28887 · 6.21 h, Judge result recordedGPT-5.6 Sol · submission 26 · 0.28860 · 6.73 h, Judge result recordedGPT-5.6 Sol · submission 27 · 0.28915 · 6.82 h, Judge result recordedGPT-5.6 Sol · submission 28 · 0.28915 · 6.91 h, Judge result recordedGPT-5.6 Sol · submission 29 · 0.28906 · 7.03 h, Judge result recordedGPT-5.6 Sol · submission 30 · 0.28915 · 7.13 h, Judge result recordedGPT-5.6 Sol · submission 31 · 0.28893 · 7.41 h, Judge result recordedGPT-5.6 Sol · submission 32 · 0.28939 · 7.63 h, Judge result recordedGPT-5.6 Sol · submission 33 · 0.28900 · 7.83 h, Judge result recordedGPT-5.6 Sol · submission 34 · 0.28953 · 7.96 h, Judge result recordedGPT-5.6 Sol · submission 35 · 0.28959 · 8.06 h, Judge result recordedGPT-5.6 Sol · submission 36 · 0.28980 · 8.17 h, Judge result recordedGPT-5.6 Sol · submission 37 · 0.28993 · 8.28 h, Judge result recordedGPT-5.6 Sol · submission 38 · 0.28998 · 8.39 h, Judge result recordedGPT-5.6 Sol · submission 39 · 0.28757 · 8.84 h, Judge result recordedGPT-5.6 Sol · submission 40 · 0.29074 · 8.99 h, Judge result recordedGPT-5.6 Sol · submission 41 · 0.29133 · 9.11 h, Judge result recordedGPT-5.6 Sol · submission 42 · 0.29204 · 9.23 h, Judge result recordedGPT-5.6 Sol · submission 43 · 0.29219 · 9.35 h, Judge result recordedGPT-5.6 Sol · submission 44 · 0.29233 · 9.88 h, Judge result recordedGPT-5.6 Sol · submission 45 · 0.29225 · 10 h, Judge result recordedGPT-5.6 Sol · submission 46 · 0.29229 · 10.1 h, Judge result recordedGPT-5.6 Sol · submission 47 · 0.28352 · 10.5 h, Judge result recordedGPT-5.6 Sol · submission 48 · 0.28249 · 11 h, Judge result recordedGPT-5.6 Sol · submission 49 · 0.29233 · 11.2 h, Judge result recordedGPT-5.6 Sol · 0.29233Claude Opus 5 · submission 1 · 0.28264 · 3.08 h, Judge result recordedClaude Opus 5 · submission 2 · 0.28166 · 3.22 h, Judge result recordedClaude Opus 5 · submission 3 · 0.28391 · 3.58 h, Judge result recordedClaude Opus 5 · submission 4 · 0.28258 · 3.9 h, Judge result recordedClaude Opus 5 · submission 5 · 0.28391 · 4.12 h, Judge result recordedClaude Opus 5 · submission 6 · 0.28815 · 4.92 h, Judge result recordedClaude Opus 5 · submission 7 · 0.28815 · 5.03 h, Judge result recordedClaude Opus 5 · submission 8 · 0.28658 · 6.29 h, Judge result recordedClaude Opus 5 · submission 9 · 0.28815 · 7.12 h, Judge result recordedClaude Opus 5 · submission 10 · 0.28825 · 8.05 h, Judge result recordedClaude Opus 5 · submission 11 · 0.28478 · 8.31 h, Judge result recordedClaude Opus 5 · submission 12 · 0.28542 · 8.41 h, Judge result recordedClaude Opus 5 · submission 13 · 0.28825 · 8.59 h, Judge result recordedClaude Opus 5 · submission 14 · 0.28825 · 8.71 h, Judge result recordedClaude Opus 5 · submission 15 · 0.28825 · 9.33 h, Judge result recordedClaude Opus 5 · submission 16 · 0.28955 · 9.89 h, Judge result recordedClaude Opus 5 · submission 17 · 0.28955 · 10.2 h, Judge result recordedClaude Opus 5 · submission 18 · 0.28955 · 10.5 h, Judge result recordedClaude Opus 5 · submission 19 · 0.28955 · 10.8 h, Judge result recordedClaude Opus 5 · submission 20 · 0.28955 · 11.4 h, Judge result recordedClaude Opus 5 · submission 21 · 0.00000 · 11.6 h, Judge result recordedClaude Opus 5 · submission 22 · 0.28955 · 11.8 h, Judge result recordedClaude Opus 5 · 0.28955
The agent runs · 2 runs · submissionrejected submissionrunning bestbaseline · validation measuredTime since run start · points mark recorded Judge results

The task

Make a small model better at choosing among answers and expressing uncertainty. It reads shared context and answers yes/no, multiple-choice, and ordered-rating questions directly with probabilities, without generating an answer token by token. The challenge is to transfer from fixed training examples to new domains and harder decisions.

Environment

Reference baseline

The 0.25186 reference is the untouched Kev pointer-head and LoRA implementation, trained from the original Qwen base using the task's fixed data and reference recipe. The public maintainer guide records this measurement during task validation on 2026-09-21.

The task-owned distributed launcher changes how training is spread across two GPUs, not the model, data, global batch, or reference recipe. The instruction table's earlier unmeasured status predates validation. This is a documented validation result, not an upstream benchmark score or either agent's first submission; Judge evaluates only the candidate, without rerunning a baseline alongside it.

Research loop

  1. Form a hypothesis about a weakness in the model's representations, decision head, or training.
  2. Implement the change and train on the fixed corpus, starting independent experiments from the supplied base.
  3. Check probabilities, calibration, and transfer on the public development examples.
  4. Validate and submit the checkpoint, then use aggregate Judge feedback to retain or revise the idea.

What the agent may change

What stays fixed

Evaluation

Judge scores 1,958 questions: 500 MMLU, 500 MMLU-Pro, and 958 other transfer and control decisions. A further 110 evidence-removed requests provide confidence diagnostics only and do not contribute to the score. These are task-specific subsets, not full benchmark results.

For each source, Judge averages the negative log probability assigned to the correct answer. It then averages sources within the knowledge and other groups: macro NLL = 0.5 × knowledge NLL + 0.5 × other NLL. The displayed score is exp(−macro NLL), between zero and one; higher is better. It is not accuracy.

Feedback includes aggregate quality, calibration, timing, memory, and candidate validation results, not hidden examples or answers. Invalid candidates receive zero; infrastructure interruptions have no score. All 71 archived submissions have numeric scores. Claude's submission 21 received zero for a research note outside the editable directories; after removing it, submission 22 reproduced 0.28955. This rejection is not a measurement of the model's decision quality. The figure above is a schematic of the decision-model idea, not a measured result or the final architecture found by either agent.

Agent runs

Claude has 22 recorded submissions and Codex has 49. The Claude curve replaces the withdrawn run and follows the added requirement that every probability come from the Kev decision head; the Codex run predates that rule. Each figure and folded table preserves the complete recorded submission history. The paired summary excerpts describe reproduced submissions, not a certification of later workspace changes.

0.2410.2560.2710.2850.30.29233round 149Decision probability scoreGPT-5.6 Sol · Codex
GPT-5.6 Sol · 49 submissions · best 0.29233
GPT-5.6 Sol — every submission, in order (49)
#Score
10.24868
20.27539
30.28177
40.28111
50.27361
60.27742
70.28177
80.28460
90.28694
100.28669
110.28705
120.28712
130.28717
140.28670
150.28691
160.28750
170.28777
180.28789
190.28790
200.28855
210.28879
220.28887
230.28863
240.28872
250.28887
260.28860
270.28915
280.28915
290.28906
300.28915
310.28893
320.28939
330.28900
340.28953
350.28959
360.28980
370.28993
380.28998
390.28757
400.29074
410.29133
420.29204
430.29219
440.29233
450.29225
460.29229
470.28352
480.28249
490.29233
0.280.2830.2860.2880.2910.28955round 122Decision probability scoreClaude Opus 5 · Claude Code
Claude Opus 5 · 22 submissions · best 0.28955
Claude Opus 5 — every submission, in order (22)
#Score
10.28264
20.28166
30.28391
40.28258
50.28391
60.28815
70.28815
80.28658
90.28815
100.28825
110.28478
120.28542
130.28825
140.28825
150.28825
160.28955
170.28955
180.28955
190.28955
200.28955
210
220.28955
GPT-5.6 Sol — the agent’s own closing-summary excerpt
  • Reward: 0.292327
  • Macro NLL: 1.22988
  • Knowledge NLL: 1.68104
  • Other NLL: 0.77872
  • Judge evaluation time: 319.0 seconds, reduced from 354.2 seconds with identical predictions
  • Parameters: 518,415,232
  • Peak memory: 2.28 GB
  • Validation and isolation: passed

Additional rank-1 duration, learning-rate, seed and attention-only controls plus rank-2 capacity testing were rejected.

Claude Opus 5 — the agent’s own closing-summary excerpt

Recovered — 0.28955 restored, workspace scope clean.

The reason I was resubmitting at all: I found that --letter-style/--pointer-style did double duty, setting both the training and the serving rendering. Since the shipped artifact trains on desc and serves on key_desc, a fresh default run would not have reproduced it. I split them into --letter-style/--pointer-style (training) and --serve-letter-style/--serve-pointer-style (serving, defaulting to key_desc), then verified a fresh launcher run produces byte-identical adapters and head weights with all style fields written explicitly rather than relying on absent-key defaults. That artifact is what's now submitted.

Links