All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.
View by
The agent runs · 2 runs · submissionrunning bestbaseline · paired normalizedTime since run start · points mark recorded Judge results
Replace the implementation behind a fused tied-weight linear cross-entropy callable so it runs faster both by itself and inside a fixed language-model training step without changing the result.
Environment
The official Liger implementation, model, data, reference wrapper, and public diagnostics are provided.
The agent delivers one self-contained Python source file implementing the candidate callable.
Reference baseline
Every submission runs the official Liger implementation and the candidate in a paired Judge evaluation. The reference therefore normalizes machine-level timing variation in the same run.
Research loop
Inspect the current callable, official source, public diagnostics, and earlier aggregate feedback.
Change one kernel, memory, or scheduling mechanism while preserving forward and backward behavior.
Run public checks, submit the closed source file, and compare operator, training, memory, and numerical aggregates.
What the agent may change
Python orchestration inside the candidate callable.
Self-contained tensor operations and Triton kernels.
Token and vocabulary tiling.
Logit recomputation and gradient-accumulation strategy.
Launch topology, memory movement, and scheduling.
What stays fixed
The callable interface, model, weights, tokenizer, data, and optimizer step.
Forward loss, both gradients, numerical tolerances, and input immutability.
The paired workloads, timing rules, and memory envelope.
External fused-loss implementations and hidden evaluator state remain unavailable.
Evaluation
The Judge combines paired speed ratios from the standalone operator and the training step into one reward index; higher is better and the official implementation anchors the reference level. Correctness, gradient, optimizer, repeatability, and memory checks gate the score.
Agent runs
Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.
GPT-5.6 Sol · 45 submissions · best 1138GPT-5.6 Sol — every submission, in order (45)
#
Score
1
—
2
1040
3
987
4
1018
5
994
6
1003
7
1022
8
1052
9
1023
10
1036
11
1074
12
1081
13
1054
14
1051
15
1073
16
1051
17
1050
18
1053
19
1014
20
1110
21
1039
22
1046
23
1094
24
1117
25
1097
26
1115
27
1119
28
1110
29
1096
30
1138
31
1088
32
1109
33
1089
34
1088
35
1094
36
1124
37
1085
38
1122
39
1092
40
1101
41
1120
42
1106
43
1097
44
1083
45
1104
Claude Opus 5 · 22 submissions · best 1131Claude Opus 5 — every submission, in order (22)
#
Score
1
—
2
1033
3
1063
4
1031
5
—
6
1009
7
1069
8
1022
9
1064
10
1028
11
1016
12
1044
13
1029
14
1107
15
1128
16
1033
17
1037
18
1120
19
1131
20
1101
21
1050
22
1083
GPT-5.6 Sol — the agent’s own closing summary
Final source: candidate.py (12,787 bytes), SHA-256 60dc44d3138f719026c3a79524e2954bf23f39e11454d0c21ebb6ef7b4940a08, the sole workspace file. Best overall score remains 1138.2575.
Claude Opus 5 — the agent’s own closing summary
Final state:/workspace/candidate.py is the sole workspace entry and was judged valid. Best reward: 1130.91; operator performance: 57.8× the official Liger default; SFT throughput: 2.21×; memory stayed below baseline in every row; loss was bit-identical to the reference.
The remaining gap belongs to forward-GEMM scheduling: the permitted toolset did not expose a faster measured path.