PR #1514 · SP8192 + Muon 0.97 + Legal Score-First TTT
The combined Muon 0.97 and causal token n-gram tilt variant improves the direct PR #1413 predecessor; the score delta is not attributed to either change alone.
Method before metrics.
On PR #1514: SP8192 with Muon 0.97 and legal score-first TTT; 3-seed sweep beats #1493 (p=0.020) Declared changes include: Score-first TTT only — every sliding-window chunk is scored under inferencemode() before any gradient update. No chunk is trained on before scoring. No SLOT, no pre-quant TTT on val data, no n-gram cache, no ETLB. [x] Score-first TTT ordering verified
- Hypothesis
- H-PG-1514 · The combined Muon 0.97 and causal token n-gram tilt variant improves the direct PR #1413 predecessor; the score delta is not attributed to either change alone.
- Derived from
- PR #1413 · SP8192 + QK-Gain 5 + Legal Score-First TTT
experiment:E-PG-1413 - Uses technique from
- PR #1394 · SP8192 + GPTQ Embeddings + Depth Recurrence + SDClip
experiment:E-PG-1394 - Compared with
- PR #1493 · SP8192 + 3-Layer Recurrence + Parallel Residuals + Legal TTT
experiment:E-PG-1493
The comparison contract.
A reproduction matches this protocol. A fork changes it and declares the deviation.
- model Family
- Parameter Golf language model
- dataset
- FineWeb validation
- benchmarks
- FineWeb validation BPB
- controlled Variables
- 16,000,000-byte artifact cap · 600-second training budget · 8×H100 SXM evaluation hardware · tokenizer-agnostic BPB
- independent Variable
- combined implementation changes declared in the source PR
- seeds
- 0 · 42 · 1234
- hardware
- 8×H100 SXM
- pr number
- 1,514
- source pr
- https://github.com/openai/parameter-golf/pull/1514
- source commit
- 59109e3036328d553f33db20ddf167a428b4150a
- source dataset commit
- f5c079314c4877fbb0af378c0abade5a8ca33d3a
- seed count
- 3
- metric direction
- lower_is_better
- artifact limit bytes
- 16,000,000
- training budget seconds
- 600
- declared changes
- Score-first TTT only — every sliding-window chunk is scored under inferencemode() before any gradient update. No chunk is trained on before scoring. · No SLOT, no pre-quant TTT on val data, no n-gram cache, no ETLB. · [x] Score-first TTT ordering verified
- declared relations
- [object Object] · [object Object] · [object Object]
What this experiment produced or evaluated.
- Produced
- SP8192 + Muon 0.97 + Legal Score-First TTT ↗
external-run:ER-PG-1514 - Uses
- benchmark:fineweb-validation-bpb
Change the evidence, not just the discussion.
Start locally, inspect the portable manifest, and sign in only if you choose to publish.
Attach a run to this protocol.
Record the measurement and its provenance. New evidence is unverified and pending review until CodeSOTA checks it.
Stable, portable, attributable.
- Stable ID
- E-PG-1514
- Visibility
- public
- Created
- 2026-04-09
- Updated
- 2026-08-20