Small English language models
Published task results and a new 100-prompt English grammar diagnostic for small models, with Bielik v3 11B as a larger comparison. Parameter count is descriptive, not an eligibility cutoff.
- Best same-harness mean
- Hymba-125M · 49.35
- Models in LightEval ranking
- 11
- New grammar diagnostic
- 20 models · 2,000 attempts
- Primary protocol
- 0-shot · 5-task mean
English sentence-completion grammar
100 shared prefixes per model · 2,000 attempts · 20 models. Qwen judges grammar and completeness separately; both must pass. Every attempt counts.
Pollock: 78/100. Bielik v3 11B: 77/100. The one-point gap does not establish a quality difference. SmolLM, SmolLM2 and GPT-2 each score 90/100.
The original table uses the Bielik Instruct Q4_K_M checkpoint with raw completion and no chat template. The separate comparison below adds its checkpoint chat template with an assistant prefix. LaMini still receives bare prefixes. These are English completion diagnostics, not general chat-quality scores.
Bielik: raw completion or chat template
Same 11B Instruct Q4_K_M checkpoint, 100 prefixes and seeds, sampling settings, and grammar rubric. The raw result is preserved from the original run. The chat run uses the checkpoint template, a completion instruction and the same prefix prefilled as assistant text.
Raw completion: 77/100 passes. Grammar: 82/100. Complete: 84/100. 95% Wilson interval: 67.8–84.2%.
This measures instruction + template + assistant prefill together, not the effect of template delimiters alone. Single-seed automatic judgments; no human audit of the new chat run. Neither score measures factual correctness or Polish-language quality.
Scatter plots: 20 checkpoints, 21 configurations
Each point is one measured configuration. Bielik raw and chat use the same checkpoint and are shown separately. Polish-trained rows remain English transfer probes.

Select a model to inspect its 100 saved attempts and judge explanations.
| Model / inspect samples | Pass / 100 ↓ | Learned params | Grammar yes | Complete yes | 95% interval | Track |
|---|---|---|---|---|---|---|
| 90 | 134.52M | 94 | 94 | 82.6–94.5% | Small English model | |
| 90 | 134.52M | 95 | 92 | 82.6–94.5% | Small English model | |
| 90 | 124.44M | 96 | 91 | 82.6–94.5% | Small English model | |
| 88 | 129.14M | 97 | 90 | 80.2–93% | Small English model | |
| 85 | 125.20M | 94 | 87 | 76.7–90.7% | Small English model | |
| 84 | 98.99M | 92 | 90 | 75.6–89.9% | Small English model | |
| 83 | 125.24M | 91 | 87 | 74.5–89.1% | Small English model | |
| 78 | 127.67M | 88 | 84 | 68.9–85% | Small English model | |
| 77 | 162.32M | 89 | 81 | 67.8–84.2% | Small English model | |
| 77 | 11.17B | 82 | 84 | 67.8–84.2% | 11B instruct · Q4_K_M | |
| 69 | 81.91M | 85 | 73 | 59.4–77.2% | Small English model | |
| 60 | 70.43M | 70 | 74 | 50.2–69.1% | Small English model | |
| 27 | 124.44M | 58 | 31 | 19.3–36.4% | Small · instruction-tuned |
Intervals are 95% Wilson intervals over 100 prompts; they do not include judge error. Parameter counts are unique learned parameters, excluding buffers. Scores are not averaged with the five-task ranking below.
Ten prefix categories, one seeded attempt per prefix. Temperature 0.8, top-k 40; up to 160 subword tokens or 640 bytes for byte models. Small models use float32; Bielik uses Q4_K_M. Judge: Qwen3.8-27B W4A16, temperature zero, model identities hidden. No retries for bad generations.
A second LLM reviewed five random attempts per model: 92 agreements, 6 disagreements and 2 uncertain cases. This is not human gold validation. Original scores remain unchanged. Small gaps should not decide which model is better.
This adds a measured generation diagnostic. Use category results and failure examples alongside BPB, grammatical minimal pairs and downstream tasks. Multiple seeds, independent judging and rank consistency across model sizes remain to be measured.
Work through the benchmark ladder →MobileLLM: gated access. Cerebras: repository API returned 404. Muse2: custom implementation unavailable. Hymba-125M: no official checkpoint located. Backend samplers and precision differ; one of five Bielik repeatability samples changed with prompt caching disabled.
LightEval small-model ranking
Ten source-reported rows from Hymba Table 6 plus the new CodeSOTA Pollock run. The source column keeps the executions explicit.
| 1 | Hymba-125M | 125M | Hymba Table 6 | 31.12 | 44.95 | 68.50 | 45.54 | 35.52 | 52.25 | 49.35 |
|---|---|---|---|---|---|---|---|---|---|---|
| 2 | SmolLM-135M | 135M | Hymba Table 6 | 30.23 | 43.99 | 69.60 | 42.30 | 33.60 | 52.70 | 48.44 |
| 3 | MobileLLM-125M | 125M | Hymba Table 6 | — | 35.51 | 65.30 | 38.90 | 39.50 | 53.10 | 46.46 |
| 4 | Mamba-130M | 130M | Hymba Table 6 | 27.41 | 33.01 | 63.33 | 33.86 | 30.40 | 51.54 | 42.43 |
| 5 | OPT-125M | 125M | Hymba Table 6 | 25.67 | 31.25 | 61.97 | 31.04 | 29.00 | 53.20 | 41.29 |
| 6 | LaMini-GPT-124M | 124M | Hymba Table 6 | 26.47 | 33.26 | 62.89 | 30.05 | 27.80 | 50.75 | 40.95 |
| 7 | GPT-Neo-125M | 125M* | Hymba Table 6 | 27.25 | 31.30 | 62.35 | 29.68 | 29.20 | 51.54 | 40.81 |
| 8 | GPT-2 small | 137M | Hymba Table 6 | 26.29 | 31.09 | 62.51 | 29.76 | 29.40 | 49.72 | 40.50 |
| 9 | Pollock 1.4 | 127.7M | CodeSOTA | 26.71 | 33.14 | 60.61 | 29.34 | 28.60 | 50.28 | 40.39 |
| 10 | Pythia-160M | 160M* | Hymba Table 6 | 26.68 | 31.92 | 61.64 | 29.55 | 27.80 | 49.49 | 40.08 |
| 11 | Cerebras-GPT-111M | 111M | Hymba Table 6 | 25.56 | 27.75 | 58.16 | 26.32 | 25.40 | 50.28 | 37.58 |
How to read it: the mean averages ARC, PIQA, HellaSwag, OpenBookQA, and WinoGrande; MMLU is displayed but excluded. Higher is better. Pollock is a full 0-shot BF16 CodeSOTA run using the pinned SmolLM2 LightEval recipe; its source link opens the result record. The other rows remain frozen exactly as reported in Hymba Table 6. * The study uses model-family labels. Older artifact totals included buffers; the new grammar run counts unique learned weights: 125.2M for GPT-Neo and 162.3M for Pythia-160M. Source-reported Table 1 labels remain unchanged.
Reported outside the reference run
Useful context, but not additional rows in Table 1. These values came from separate executions.
| Model | Current params | Source / protocol | MMLU | ARC c+e | PIQA | HellaSwag | OpenBookQA | WinoGrande | 5-task mean |
|---|---|---|---|---|---|---|---|---|---|
| SmolLM2-135M | 134.5M | SmolLM2 card · LightEval · 0-shot | 31.50 | 43.90 | 68.40 | 42.10 | 34.60 | 51.30 | 48.06 |
SmolLM2 remains here because its values are source-reported from its model card rather than the Hymba reference table. Its 48.06 mean is calculated from the five displayed model-card values. MMLU is the cloze/continuation task, not letter-choice MMLU.
Small-model diagnostics
Tests that reveal language-model quality at this scale without changing the primary five-task ranking.
| Model | MMLU continuation | BLiMP | SciQ | LAMBADA acc. | LAMBADA PPL ↓ | WikiText-2 BPB ↓ | Source |
|---|---|---|---|---|---|---|---|
| Pollock 1.4 | 25.52 | 78.07 | 65.60 | 27.83 | 53.54 | 1.007 | CodeSOTA · full BF16 run |
MMLU continuation scores answer text directly rather than selecting A–D. BLiMP tests grammatical minimal pairs; LAMBADA tests final-word prediction; WikiText-2 BPB is tokenizer-comparable and lower is better. These diagnostics are displayed separately because averaging unlike metrics would create an arbitrary composite.
Next measurements
Cells remain pending until the same pinned implementation has run for every comparison model. No placeholder score enters the table.
| Family | Benchmark | Published measurements | Pollock status | What it isolates |
|---|---|---|---|---|
| Language fit | Multi-domain BPB | BPB by domain ↓ | Pending | Tokenizer-comparable fit on prose, news, dialogue, scientific text, and code. |
| Grammar | BLiMP | Accuracy · category accuracy · mean margin ↑ | Aggregate complete | Adds phenomenon-level results and confidence, not only correct/incorrect pairs. |
| Grammar | BLiMP Supplement | Accuracy · category accuracy ↑ | Pending | Extends the minimal-pair grammar suite with broader lexical phenomena. |
| Context | LAMBADA OpenAI | Accuracy ↑ · target NLL ↓ · MRR ↑ · top-5 ↑ | Accuracy + PPL complete | Continuous target-token metrics retain signal below exact-match thresholds. |
| World knowledge | EWoK | Macro accuracy · category accuracy ↑ | Pending | Paired likelihood tests for elementary world knowledge; report distance from 50%. |
| Targeted syntax | SyntaxGym | Suite accuracy · surprisal effect ↑ | Pending | Tests whether expected syntactic surprisal effects appear in controlled suites. |
Reporting rule: publish raw benchmark columns and category breakdowns; do not create a single diagnostic average. Multi-domain BPB replaces cross-tokenizer word perplexity. EWoK receives confidence intervals because a small model may remain close to its 50% paired-choice baseline.
Unified CodeSOTA rerun roster
Roster for the five-task rerun. Counts use learned weights where measured in the grammar run. Harness status below refers to task evaluation; grammar results are reported separately above.
| Model | Learned params | Architecture | Track | Harness status |
|---|---|---|---|---|
| Pollock 1.4 | 127.7M | Transformer | Core | Complete · BF16 |
| SmolLM2-135M | 134.5M | Transformer | Core | Run |
| MobileLLM-125M | ≈124.6M | Transformer | Core | Run |
| Mamba-130M | 129.1M | State-space | Core | Run |
| OPT-125M | ≈125M | Transformer | Core | Run |
| GPT-Neo-125M | 125.2M | Transformer | Core | Run |
| GPT-2 small | 124.4M | Transformer | Core | Run |
| Pythia-160M | 162.3M | Transformer | Core | Run |
| Cerebras-GPT-111M | 111M | Transformer | Core | Run |
| Pythia-70M | 70.4M | Transformer | Baseline | Run |
| DistilGPT2 | 82M | Transformer | Baseline | Run |
| Muse2-125M Base | 122.9M | Hybrid conv/attention | Additional | Adapter needed |
| TLM-100M | 99.0M | Tiered GPT-Neo | Additional | Custom class needed |
How the fresh ranking should be run
Published scores remain frozen above. New CodeSOTA results get their own source marker and never silently replace source-reported values.
ARC-Easy, ARC-Challenge, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Publish both ARC components and their arithmetic mean.
Unweighted mean of ARC c+e, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Higher is better.
MMLU continuation, SciQ, multi-domain BPB, BLiMP plus Supplement, LAMBADA target-token metrics, EWoK, and SyntaxGym. Show these separately; do not mix them into the main mean.
One pinned harness and dataset revision, 0-shot, full test splits, fixed checkpoint SHAs, matching precision, and saved raw outputs.