All logged agent runs share one plot. Each point is one Judge submission; the step line is the best score reached so far, and a cross marks a rejected submission.
View by
The agent runs · 2 runs · submission×rejected submissionrunning bestbaseline · paired normalizedTime since run start · points mark recorded Judge results
Reduce the latency of a fixed vertical-plus-slash sparse-attention prefill operation without dropping attention work, changing causality, or weakening its output.
Environment
The source-equivalent fallback, public source, fixed interface, and public diagnostics are provided.
The agent delivers one candidate implementation plus brief unscored research notes.
Reference baseline
Every submission runs the frozen upstream fallback and the candidate on matching fresh inputs in the same Judge invocation. The score is therefore a paired speedup rather than an absolute latency.
Research loop
Inspect the fallback, sparse-attention contract, public diagnostics, and earlier aggregate feedback.
Test one execution-plan hypothesis and record the public evidence.
Finish the source, submit it, and compare latency, correctness, and resource aggregates before revising.
What the agent may change
Candidate dispatch and metadata handling.
Sparse-index conversion, canonicalization, and fusion.
Head grouping, batching, and intermediate layout.
Tiling and online-softmax accumulation.
Kernel decomposition, launch geometry, and scheduling.
Unscored research notes.
What stays fixed
The sparse-attention pattern, causal union, input and output semantics, and callable interface.
The model geometry, data types, reserved cases, and numerical agreement rules.
Input immutability, deterministic full-output behavior, and caller-stream execution.
The paired fallback, timing procedure, and offline Judge environment.
Evaluation
For every reserved case, the Judge compares fallback and candidate median latency and combines the ratios geometrically; higher is better. Scope, stream, determinism, full-write, mutation, and numerical checks must pass before the speed result is accepted.
Agent runs
Each figure is one run: a lollipop per submission, the running best as a step line, and a cross on the floor where a submission was rejected. Open the row below each figure for every Judge submission.
GPT-5.6 Sol · 61 submissions · best 1.54GPT-5.6 Sol — every submission, in order (61)
#
Score
1
1.08
2
0
3
1.14
4
1.14
5
1.19
6
1.2
7
1.24
8
1.31
9
1.33
10
1.33
11
1.34
12
1.36
13
1.35
14
1.36
15
1.36
16
1.36
17
1.39
18
1.39
19
1.39
20
1.4
21
1.39
22
1.38
23
1.46
24
1.46
25
1.46
26
1.47
27
1.45
28
1.47
29
1.46
30
0
31
1.47
32
1.46
33
1.46
34
1.46
35
1.46
36
1.47
37
1.47
38
1.47
39
1.48
40
1.47
41
1.48
42
1.47
43
1.48
44
1.48
45
1.47
46
1.48
47
1.48
48
1.48
49
1.47
50
1.47
51
1.47
52
1.53
53
1.54
54
1.54
55
1.53
56
1.53
57
1.54
58
1.53
59
1.54
60
1.54
61
1.54
Claude Opus 5 · 15 submissions · best 1.5Claude Opus 5 — every submission, in order (15)
#
Score
1
1.23
2
0
3
1.37
4
1.45
5
1.47
6
1.46
7
1.47
8
1.5
9
1.49
10
1.49
11
1.49
12
1.5
13
1.5
14
1.5
15
1.49
GPT-5.6 Sol — the agent’s own closing summary
Implemented and submitted the optimized operator. Best formal result: 1.5405066467× geometric-mean speedup (submission 60).
32K: 1.5770×
65K: 1.5369×
131K: 1.5084×
Correctness and policy checks passed
Peak HBM: 16,385.25 MiB
The retained plan uses certified translated-quad fusion, excludes boundary-crossing final quads, and removes redundant tail activity predicates through an exact causal-range certificate.
Claude Opus 5 — the agent’s own closing summary
Delivered:/workspace/submission.py at approximately 1.50× the frozen fallback, with best recorded reward 1.5041 (32K: 1.519×, 64K: 1.524×, 128K: 1.461×), plus a summary-first notes.md covering 23 hypotheses — 5 shipped and 18 rejected on measurement.
The retained gains came from causal-mask hoisting, a four-way merged path, closed-form index conversion where valid, a streaming converter elsewhere, and heaviest-first program ordering.