Small English language models
Published benchmark results for compact base models around the 125M class. Parameter count is descriptive, not an eligibility cutoff.
- Best same-harness mean
- Hymba-125M · 49.35
- Models in LightEval ranking
- 11
- Models in rerun roster
- 13
- Primary protocol
- 0-shot · 5-task mean
LightEval small-model ranking
Ten source-reported rows from Hymba Table 6 plus the new CodeSOTA Pollock run. The source column keeps the executions explicit.
| 1 | Hymba-125M | 125M | Hymba Table 6 | 31.12 | 44.95 | 68.50 | 45.54 | 35.52 | 52.25 | 49.35 |
|---|---|---|---|---|---|---|---|---|---|---|
| 2 | SmolLM-135M | 135M | Hymba Table 6 | 30.23 | 43.99 | 69.60 | 42.30 | 33.60 | 52.70 | 48.44 |
| 3 | MobileLLM-125M | 125M | Hymba Table 6 | — | 35.51 | 65.30 | 38.90 | 39.50 | 53.10 | 46.46 |
| 4 | Mamba-130M | 130M | Hymba Table 6 | 27.41 | 33.01 | 63.33 | 33.86 | 30.40 | 51.54 | 42.43 |
| 5 | OPT-125M | 125M | Hymba Table 6 | 25.67 | 31.25 | 61.97 | 31.04 | 29.00 | 53.20 | 41.29 |
| 6 | LaMini-GPT-124M | 124M | Hymba Table 6 | 26.47 | 33.26 | 62.89 | 30.05 | 27.80 | 50.75 | 40.95 |
| 7 | GPT-Neo-125M | 125M* | Hymba Table 6 | 27.25 | 31.30 | 62.35 | 29.68 | 29.20 | 51.54 | 40.81 |
| 8 | GPT-2 small | 137M | Hymba Table 6 | 26.29 | 31.09 | 62.51 | 29.76 | 29.40 | 49.72 | 40.50 |
| 9 | Pollock 1.4 | 127.7M | CodeSOTA | 26.71 | 33.14 | 60.61 | 29.34 | 28.60 | 50.28 | 40.39 |
| 10 | Pythia-160M | 160M* | Hymba Table 6 | 26.68 | 31.92 | 61.64 | 29.55 | 27.80 | 49.49 | 40.08 |
| 11 | Cerebras-GPT-111M | 111M | Hymba Table 6 | 25.56 | 27.75 | 58.16 | 26.32 | 25.40 | 50.28 | 37.58 |
How to read it: the mean averages ARC, PIQA, HellaSwag, OpenBookQA, and WinoGrande; MMLU is displayed but excluded. Higher is better. Pollock is a full 0-shot BF16 CodeSOTA run using the pinned SmolLM2 LightEval recipe; its source link opens the result record. The other rows remain frozen exactly as reported in Hymba Table 6. * The study uses model-family labels: current Hugging Face artifacts report 150.4M parameters for GPT-Neo-125M and 212.7M for Pythia-160M.
Reported outside the reference run
Useful context, but not additional rows in Table 1. These values came from separate executions.
| Model | Current params | Source / protocol | MMLU | ARC c+e | PIQA | HellaSwag | OpenBookQA | WinoGrande | 5-task mean |
|---|---|---|---|---|---|---|---|---|---|
| SmolLM2-135M | 134.5M | SmolLM2 card · LightEval · 0-shot | 31.50 | 43.90 | 68.40 | 42.10 | 34.60 | 51.30 | 48.06 |
SmolLM2 remains here because its values are source-reported from its model card rather than the Hymba reference table. Its 48.06 mean is calculated from the five displayed model-card values. MMLU is the cloze/continuation task, not letter-choice MMLU.
Small-model diagnostics
Tests that reveal language-model quality at this scale without changing the primary five-task ranking.
| Model | MMLU continuation | BLiMP | SciQ | LAMBADA acc. | LAMBADA PPL ↓ | WikiText-2 BPB ↓ | Source |
|---|---|---|---|---|---|---|---|
| Pollock 1.4 | 25.52 | 78.07 | 65.60 | 27.83 | 53.54 | 1.007 | CodeSOTA · full BF16 run |
MMLU continuation scores answer text directly rather than selecting A–D. BLiMP tests grammatical minimal pairs; LAMBADA tests final-word prediction; WikiText-2 BPB is tokenizer-comparable and lower is better. These diagnostics are displayed separately because averaging unlike metrics would create an arbitrary composite.
Next measurements
Cells remain pending until the same pinned implementation has run for every comparison model. No placeholder score enters the table.
| Family | Benchmark | Published measurements | Pollock status | What it isolates |
|---|---|---|---|---|
| Language fit | Multi-domain BPB | BPB by domain ↓ | Pending | Tokenizer-comparable fit on prose, news, dialogue, scientific text, and code. |
| Grammar | BLiMP | Accuracy · category accuracy · mean margin ↑ | Aggregate complete | Adds phenomenon-level results and confidence, not only correct/incorrect pairs. |
| Grammar | BLiMP Supplement | Accuracy · category accuracy ↑ | Pending | Extends the minimal-pair grammar suite with broader lexical phenomena. |
| Context | LAMBADA OpenAI | Accuracy ↑ · target NLL ↓ · MRR ↑ · top-5 ↑ | Accuracy + PPL complete | Continuous target-token metrics retain signal below exact-match thresholds. |
| World knowledge | EWoK | Macro accuracy · category accuracy ↑ | Pending | Paired likelihood tests for elementary world knowledge; report distance from 50%. |
| Targeted syntax | SyntaxGym | Suite accuracy · surprisal effect ↑ | Pending | Tests whether expected syntactic surprisal effects appear in controlled suites. |
Reporting rule: publish raw benchmark columns and category breakdowns; do not create a single diagnostic average. Multi-domain BPB replaces cross-tokenizer word perplexity. EWoK receives confidence intervals because a small model may remain close to its 50% paired-choice baseline.
Unified CodeSOTA rerun roster
This is the model set that can answer the comparison cleanly. The exact total parameter count is shown where repository metadata differs from the model name.
| Model | Current total params | Architecture | Track | Harness status |
|---|---|---|---|---|
| Pollock 1.4 | 127.7M | Transformer | Core | Complete · BF16 |
| SmolLM2-135M | 134.5M | Transformer | Core | Run |
| MobileLLM-125M | ≈124.6M | Transformer | Core | Run |
| Mamba-130M | 129.1M | State-space | Core | Run |
| OPT-125M | ≈125M | Transformer | Core | Run |
| GPT-Neo-125M | 150.4M | Transformer | Core | Run |
| GPT-2 small | 137.0M | Transformer | Core | Run |
| Pythia-160M | 212.7M | Transformer | Core | Run |
| Cerebras-GPT-111M | 111M | Transformer | Core | Run |
| Pythia-70M | 95.6M | Transformer | Baseline | Run |
| DistilGPT2 | 82M | Transformer | Baseline | Run |
| Muse2-125M Base | 122.9M | Hybrid conv/attention | Additional | Adapter needed |
| TLM-100M | 99.0M | Tiered GPT-Neo | Additional | Custom class needed |
How the fresh ranking should be run
Published scores remain frozen above. New CodeSOTA results get their own source marker and never silently replace source-reported values.
ARC-Easy, ARC-Challenge, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Publish both ARC components and their arithmetic mean.
Unweighted mean of ARC c+e, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Higher is better.
MMLU continuation, SciQ, multi-domain BPB, BLiMP plus Supplement, LAMBADA target-token metrics, EWoK, and SyntaxGym. Show these separately; do not mix them into the main mean.
One pinned harness and dataset revision, 0-shot, full test splits, fixed checkpoint SHAs, matching precision, and saved raw outputs.