Task Sample · Posttrain

Learnability-Aware Long/Short CoT Adaptation

Adapt the supplied long and short mathematical reasoning examples so the same small language model learns more effectively under a fixed training recipe.

Best score per run5-task macro accuracy (pp) over AIME 2024, AMC 2023, OlympiadBench-EN, GSM8K test, MATH-500
051.3
Model · harnessBestSubmissionsRuntime
Claude Fable 5Claude Code47.5237.7 h
GPT-5.6 SolCodex · xhigh46.1277.81 h

Absolute five-benchmark macro accuracy for the candidate checkpoint; higher is better.

Teaser figure from Small-Model-Learnability-Gap (Small Models Struggle to Learn from Strong Reasoners)
Source project · Small-Model-Learnability-Gap · source ↗

The climb

All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.

3438424650Baseline · 46.073 pp02467.9Elapsed time (h)5-task macro accuracyClaude Fable 5 · submission 1 · 46.1 · 0.29 h, Judge result recordedClaude Fable 5 · submission 2 · 0 · 0.789 h, Judge result recordedClaude Fable 5 · submission 3 · 0 · 1.71 h, Judge result recordedClaude Fable 5 · submission 4 · 44.2 · 2.19 h, Judge result recordedClaude Fable 5 · submission 5 · 0 · 2.42 h, Judge result recordedClaude Fable 5 · submission 6 · 43 · 2.7 h, Judge result recordedClaude Fable 5 · submission 7 · 36.4 · 3.1 h, Judge result recordedClaude Fable 5 · submission 8 · 44.2 · 3.38 h, Judge result recordedClaude Fable 5 · submission 9 · 44.2 · 3.67 h, Judge result recordedClaude Fable 5 · submission 10 · 45.5 · 4.01 h, Judge result recordedClaude Fable 5 · submission 11 · 43.9 · 4.28 h, Judge result recordedClaude Fable 5 · submission 12 · 45.4 · 4.62 h, Judge result recordedClaude Fable 5 · submission 13 · 41.1 · 4.93 h, Judge result recordedClaude Fable 5 · submission 14 · 44.7 · 5.19 h, Judge result recordedClaude Fable 5 · submission 15 · 43.3 · 5.45 h, Judge result recordedClaude Fable 5 · submission 16 · 42.3 · 5.72 h, Judge result recordedClaude Fable 5 · submission 17 · 42.2 · 5.99 h, Judge result recordedClaude Fable 5 · submission 18 · 44.6 · 6.24 h, Judge result recordedClaude Fable 5 · submission 19 · 44 · 6.49 h, Judge result recordedClaude Fable 5 · submission 20 · 47.5 · 6.75 h, Judge result recordedClaude Fable 5 · submission 21 · 42.6 · 7.15 h, Judge result recordedClaude Fable 5 · submission 22 · 44 · 7.41 h, Judge result recordedClaude Fable 5 · submission 23 · 45.7 · 7.67 h, Judge result recordedClaude Fable 5 · 47.5GPT-5.6 Sol · submission 1 · 0 · 0.22 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0 · 0.239 h, Judge result recordedGPT-5.6 Sol · submission 3 · 0 · 0.246 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0 · 0.433 h, Judge result recordedGPT-5.6 Sol · submission 5 · 45.5 · 0.645 h, Judge result recordedGPT-5.6 Sol · submission 6 · 44.6 · 0.971 h, Judge result recordedGPT-5.6 Sol · submission 7 · 45.1 · 1.31 h, Judge result recordedGPT-5.6 Sol · submission 8 · 42 · 1.63 h, Judge result recordedGPT-5.6 Sol · submission 9 · 43.2 · 1.95 h, Judge result recordedGPT-5.6 Sol · submission 10 · 44.7 · 2.25 h, Judge result recordedGPT-5.6 Sol · submission 11 · 43.5 · 2.56 h, Judge result recordedGPT-5.6 Sol · submission 12 · 41.9 · 2.87 h, Judge result recordedGPT-5.6 Sol · submission 13 · 44.2 · 3.18 h, Judge result recordedGPT-5.6 Sol · submission 14 · 45 · 3.48 h, Judge result recordedGPT-5.6 Sol · submission 15 · 46.1 · 3.8 h, Judge result recordedGPT-5.6 Sol · submission 16 · 45.7 · 4.13 h, Judge result recordedGPT-5.6 Sol · submission 17 · 44.1 · 4.44 h, Judge result recordedGPT-5.6 Sol · submission 18 · 43.9 · 4.77 h, Judge result recordedGPT-5.6 Sol · submission 19 · 43.8 · 5.09 h, Judge result recordedGPT-5.6 Sol · submission 20 · 45.1 · 5.46 h, Judge result recordedGPT-5.6 Sol · submission 21 · 43.7 · 5.78 h, Judge result recordedGPT-5.6 Sol · submission 22 · 44.8 · 6.12 h, Judge result recordedGPT-5.6 Sol · submission 23 · 43.3 · 6.49 h, Judge result recordedGPT-5.6 Sol · submission 24 · 43.7 · 6.83 h, Judge result recordedGPT-5.6 Sol · submission 25 · 45.1 · 7.14 h, Judge result recordedGPT-5.6 Sol · submission 26 · 44.3 · 7.45 h, Judge result recordedGPT-5.6 Sol · submission 27 · 41.6 · 7.78 h, Judge result recordedGPT-5.6 Sol · 46.1
The agent runs · 2 runs · submissionrejected submissionrunning bestbaseline · Judge measuredTime since run start · points mark recorded Judge results

The task

Adapt the supplied long and short mathematical reasoning responses into training data that a fixed small language model can learn from more effectively.

Environment

Reference baseline

The official long/short mixture checkpoint is evaluated under the same Judge protocol as the candidate. The page reports each candidate's absolute macro accuracy while the Judge also returns its difference from that baseline.

Research loop

  1. Inspect the reasoning pools, reference checkpoint, and earlier aggregate feedback.
  2. Form one hypothesis about what the small model can learn and materialize the adapted data.
  3. Train the fixed recipe, submit the complete checkpoint, and compare benchmark-level aggregates before revising.

What the agent may change

What stays fixed

Evaluation

The Judge evaluates the reference and candidate checkpoints on the same five mathematical-reasoning benchmarks. The score is the candidate's unweighted macro accuracy in percentage points; higher is better. Only aggregate benchmark results and training provenance are exposed.

Agent runs

Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.

34.438.24245.749.547.5round 1235-task macro accuracyClaude Fable 5 · Claude Code
Claude Fable 5 · 23 submissions · best 47.5
Claude Fable 5 — every submission, in order (23)
#Score
146.1
20
30
444.2
50
643
736.4
844.2
944.2
1045.5
1143.9
1245.4
1341.1
1444.7
1543.3
1642.3
1742.2
1844.6
1944
2047.5
2142.6
2244
2345.7
40.842.343.845.446.946.1round 1275-task macro accuracyGPT-5.6 Sol · Codex
GPT-5.6 Sol · 27 submissions · best 46.1
GPT-5.6 Sol — every submission, in order (27)
#Score
10
20
30
40
545.5
644.6
745.1
842
943.2
1044.7
1143.5
1241.9
1344.2
1445
1546.1
1645.7
1744.1
1843.9
1943.8
2045.1
2143.7
2244.8
2343.3
2443.7
2545.1
2644.3
2741.6

Links