Task Sample · Systems

Fused Tied-Weight Linear Cross-Entropy on H100

Optimize a fused tied-weight linear cross-entropy implementation so it accelerates both the operator itself and a fixed language-model training step.

Best score per runReward index: 100 * sqrt(operator_ratio * sft_ratio), baseline = 100
01229
Model · harnessBestSubmissionsRuntime
GPT-5.6 SolCodex · high1138456.3 h
Claude Opus 5Claude Code · high1131226.26 h

Operator and training speed are measured against the official Liger implementation inside the same Judge run.

Teaser figure from Liger-Kernel
Source project · Liger-Kernel · source ↗

The climb

All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.

9541008106311171172Baseline · 100 reward index0246.3Elapsed time (h)Reward indexGPT-5.6 Sol · submission 2 · 1040 · 0.278 h, Judge result recordedGPT-5.6 Sol · submission 3 · 987 · 0.398 h, Judge result recordedGPT-5.6 Sol · submission 4 · 1018 · 0.546 h, Judge result recordedGPT-5.6 Sol · submission 5 · 994 · 0.759 h, Judge result recordedGPT-5.6 Sol · submission 6 · 1003 · 0.985 h, Judge result recordedGPT-5.6 Sol · submission 7 · 1022 · 1.17 h, Judge result recordedGPT-5.6 Sol · submission 8 · 1052 · 1.29 h, Judge result recordedGPT-5.6 Sol · submission 9 · 1023 · 1.4 h, Judge result recordedGPT-5.6 Sol · submission 10 · 1036 · 1.5 h, Judge result recordedGPT-5.6 Sol · submission 11 · 1074 · 1.58 h, Judge result recordedGPT-5.6 Sol · submission 12 · 1081 · 1.68 h, Judge result recordedGPT-5.6 Sol · submission 13 · 1054 · 1.78 h, Judge result recordedGPT-5.6 Sol · submission 14 · 1051 · 1.86 h, Judge result recordedGPT-5.6 Sol · submission 15 · 1073 · 1.96 h, Judge result recordedGPT-5.6 Sol · submission 16 · 1051 · 2.04 h, Judge result recordedGPT-5.6 Sol · submission 17 · 1050 · 2.34 h, Judge result recordedGPT-5.6 Sol · submission 18 · 1053 · 2.43 h, Judge result recordedGPT-5.6 Sol · submission 19 · 1014 · 2.54 h, Judge result recordedGPT-5.6 Sol · submission 20 · 1110 · 2.73 h, Judge result recordedGPT-5.6 Sol · submission 21 · 1039 · 2.87 h, Judge result recordedGPT-5.6 Sol · submission 22 · 1046 · 3.04 h, Judge result recordedGPT-5.6 Sol · submission 23 · 1094 · 3.15 h, Judge result recordedGPT-5.6 Sol · submission 24 · 1117 · 3.27 h, Judge result recordedGPT-5.6 Sol · submission 25 · 1097 · 3.38 h, Judge result recordedGPT-5.6 Sol · submission 26 · 1115 · 3.48 h, Judge result recordedGPT-5.6 Sol · submission 27 · 1119 · 3.61 h, Judge result recordedGPT-5.6 Sol · submission 28 · 1110 · 3.71 h, Judge result recordedGPT-5.6 Sol · submission 29 · 1096 · 3.82 h, Judge result recordedGPT-5.6 Sol · submission 30 · 1138 · 3.91 h, Judge result recordedGPT-5.6 Sol · submission 31 · 1088 · 4.09 h, Judge result recordedGPT-5.6 Sol · submission 32 · 1109 · 4.27 h, Judge result recordedGPT-5.6 Sol · submission 33 · 1089 · 4.38 h, Judge result recordedGPT-5.6 Sol · submission 34 · 1088 · 4.59 h, Judge result recordedGPT-5.6 Sol · submission 35 · 1094 · 4.68 h, Judge result recordedGPT-5.6 Sol · submission 36 · 1124 · 4.79 h, Judge result recordedGPT-5.6 Sol · submission 37 · 1085 · 4.89 h, Judge result recordedGPT-5.6 Sol · submission 38 · 1122 · 4.99 h, Judge result recordedGPT-5.6 Sol · submission 39 · 1092 · 5.07 h, Judge result recordedGPT-5.6 Sol · submission 40 · 1101 · 5.2 h, Judge result recordedGPT-5.6 Sol · submission 41 · 1120 · 5.54 h, Judge result recordedGPT-5.6 Sol · submission 42 · 1106 · 5.62 h, Judge result recordedGPT-5.6 Sol · submission 43 · 1097 · 5.83 h, Judge result recordedGPT-5.6 Sol · submission 44 · 1083 · 6.06 h, Judge result recordedGPT-5.6 Sol · submission 45 · 1104 · 6.3 h, Judge result recordedGPT-5.6 Sol · 1138Claude Opus 5 · submission 2 · 1033 · 0.179 h, Judge result recordedClaude Opus 5 · submission 3 · 1063 · 0.397 h, Judge result recordedClaude Opus 5 · submission 4 · 1031 · 0.507 h, Judge result recordedClaude Opus 5 · submission 6 · 1009 · 1.17 h, Judge result recordedClaude Opus 5 · submission 7 · 1069 · 1.28 h, Judge result recordedClaude Opus 5 · submission 8 · 1022 · 3.41 h, Judge result recordedClaude Opus 5 · submission 9 · 1064 · 3.51 h, Judge result recordedClaude Opus 5 · submission 10 · 1028 · 3.68 h, Judge result recordedClaude Opus 5 · submission 11 · 1016 · 3.88 h, Judge result recordedClaude Opus 5 · submission 12 · 1044 · 4.03 h, Judge result recordedClaude Opus 5 · submission 13 · 1029 · 4.13 h, Judge result recordedClaude Opus 5 · submission 14 · 1107 · 4.37 h, Judge result recordedClaude Opus 5 · submission 15 · 1128 · 4.5 h, Judge result recordedClaude Opus 5 · submission 16 · 1033 · 4.81 h, Judge result recordedClaude Opus 5 · submission 17 · 1037 · 4.92 h, Judge result recordedClaude Opus 5 · submission 18 · 1120 · 5.02 h, Judge result recordedClaude Opus 5 · submission 19 · 1131 · 5.18 h, Judge result recordedClaude Opus 5 · submission 20 · 1101 · 5.64 h, Judge result recordedClaude Opus 5 · submission 21 · 1050 · 5.93 h, Judge result recordedClaude Opus 5 · submission 22 · 1083 · 6.26 h, Judge result recordedClaude Opus 5 · 1131
The agent runs · 2 runs · submissionrunning bestbaseline · paired normalizedTime since run start · points mark recorded Judge results

The task

Replace the implementation behind a fused tied-weight linear cross-entropy callable so it runs faster both by itself and inside a fixed language-model training step without changing the result.

Environment

Reference baseline

Every submission runs the official Liger implementation and the candidate in a paired Judge evaluation. The reference therefore normalizes machine-level timing variation in the same run.

Research loop

  1. Inspect the current callable, official source, public diagnostics, and earlier aggregate feedback.
  2. Change one kernel, memory, or scheduling mechanism while preserving forward and backward behavior.
  3. Run public checks, submit the closed source file, and compare operator, training, memory, and numerical aggregates.

What the agent may change

What stays fixed

Evaluation

The Judge combines paired speed ratios from the standalone operator and the training step into one reward index; higher is better and the official implementation anchors the reference level. Correctness, gradient, optimizer, repeatability, and memory checks gate the score.

Agent runs

Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.

96010111063111411651138round 145Reward indexGPT-5.6 Sol · Codex
GPT-5.6 Sol · 45 submissions · best 1138
GPT-5.6 Sol — every submission, in order (45)
#Score
1
21040
3987
41018
5994
61003
71022
81052
91023
101036
111074
121081
131054
141051
151073
161051
171050
181053
191014
201110
211039
221046
231094
241117
251097
261115
271119
281110
291096
301138
311088
321109
331089
341088
351094
361124
371085
381122
391092
401101
411120
421106
431097
441083
451104
98710291070111111531131round 122Reward indexClaude Opus 5 · Claude Code
Claude Opus 5 · 22 submissions · best 1131
Claude Opus 5 — every submission, in order (22)
#Score
1
21033
31063
41031
5
61009
71069
81022
91064
101028
111016
121044
131029
141107
151128
161033
171037
181120
191131
201101
211050
221083
GPT-5.6 Sol — the agent’s own closing summary

Final source: candidate.py (12,787 bytes), SHA-256 60dc44d3138f719026c3a79524e2954bf23f39e11454d0c21ebb6ef7b4940a08, the sole workspace file. Best overall score remains 1138.2575.

Claude Opus 5 — the agent’s own closing summary

Final state: /workspace/candidate.py is the sole workspace entry and was judged valid. Best reward: 1130.91; operator performance: 57.8× the official Liger default; SFT throughput: 2.21×; memory stayed below baseline in every row; loss was bit-identical to the reference.

The remaining gap belongs to forward-GEMM scheduling: the permitted toolset did not expose a faster measured path.

Links