LLM/benchmark ladder/small English models

Small English language models

Published task results and a new 100-prompt English grammar diagnostic for small models, with Bielik v3 11B as a larger comparison. Parameter count is descriptive, not an eligibility cutoff.

Best same-harness mean
Hymba-125M · 49.35
Models in LightEval ranking
11
New grammar diagnostic
20 models · 2,000 attempts
Primary protocol
0-shot · 5-task mean
New measurement · 11 September 2026

English sentence-completion grammar

100 shared prefixes per model · 2,000 attempts · 20 models. Qwen judges grammar and completeness separately; both must pass. Every attempt counts.

Pollock: 78/100. Bielik v3 11B: 77/100. The one-point gap does not establish a quality difference. SmolLM, SmolLM2 and GPT-2 each score 90/100.

The original table uses the Bielik Instruct Q4_K_M checkpoint with raw completion and no chat template. The separate comparison below adds its checkpoint chat template with an assistant prefix. LaMini still receives bare prefixes. These are English completion diagnostics, not general chat-quality scores.

Bielik: raw completion or chat template

Same 11B Instruct Q4_K_M checkpoint, 100 prefixes and seeds, sampling settings, and grammar rubric. The raw result is preserved from the original run. The chat run uses the checkpoint template, a completion instruction and the same prefix prefilled as assistant text.

Raw completion: 77/100 passes. Grammar: 82/100. Complete: 84/100. 95% Wilson interval: 67.8–84.2%.

This measures instruction + template + assistant prefill together, not the effect of template delimiters alone. Single-seed automatic judgments; no human audit of the new chat run. Neither score measures factual correctness or Polish-language quality.

Scatter plots: 20 checkpoints, 21 configurations

Each point is one measured configuration. Bielik raw and chat use the same checkpoint and are shown separately. Polish-trained rows remain English transfer probes.

Two scatter plots: learned parameter count versus English completion passes, and grammar versus sentence completeness. Bielik raw and chat-template results are highlighted separately.

Select a model to inspect its 100 saved attempts and judge explanations.

Automatic English grammar and sentence completeness, 100 attempts per model. Higher pass count is better.
Model / inspect samplesPass / 100 ↓Learned paramsGrammar yesComplete yes95% intervalTrack
90134.52M949482.6–94.5%Small English model
90134.52M959282.6–94.5%Small English model
90124.44M969182.6–94.5%Small English model
88129.14M979080.2–93%Small English model
85125.20M948776.7–90.7%Small English model
8498.99M929075.6–89.9%Small English model
83125.24M918774.5–89.1%Small English model
78127.67M888468.9–85%Small English model
77162.32M898167.8–84.2%Small English model
7711.17B828467.8–84.2%11B instruct · Q4_K_M
6981.91M857359.4–77.2%Small English model
6070.43M707450.2–69.1%Small English model
27124.44M583119.3–36.4%Small · instruction-tuned

Intervals are 95% Wilson intervals over 100 prompts; they do not include judge error. Parameter counts are unique learned parameters, excluding buffers. Scores are not averaged with the five-task ranking below.

Protocol and provenance

Ten prefix categories, one seeded attempt per prefix. Temperature 0.8, top-k 40; up to 160 subword tokens or 640 bytes for byte models. Small models use float32; Bielik uses Q4_K_M. Judge: Qwen3.8-27B W4A16, temperature zero, model identities hidden. No retries for bad generations.

Judge audit: provisional evidence

A second LLM reviewed five random attempts per model: 92 agreements, 6 disagreements and 2 uncertain cases. This is not human gold validation. Original scores remain unchanged. Small gaps should not decide which model is better.

Place in the benchmark ladder

This adds a measured generation diagnostic. Use category results and failure examples alongside BPB, grammatical minimal pairs and downstream tasks. Multiple seeds, independent judging and rank consistency across model sizes remain to be measured.

Work through the benchmark ladder →
Coverage and limitations

MobileLLM: gated access. Cerebras: repository API returned 404. Muse2: custom implementation unavailable. Hymba-125M: no official checkpoint located. Backend samplers and precision differ; one of five Bielik repeatability samples changed with prompt caching disabled.

Table 1

LightEval small-model ranking

Ten source-reported rows from Hymba Table 6 plus the new CodeSOTA Pollock run. The source column keeps the executions explicit.

1Hymba-125M125MHymba Table 631.1244.9568.5045.5435.5252.2549.35
2SmolLM-135M135MHymba Table 630.2343.9969.6042.3033.6052.7048.44
3MobileLLM-125M125MHymba Table 635.5165.3038.9039.5053.1046.46
4Mamba-130M130MHymba Table 627.4133.0163.3333.8630.4051.5442.43
5OPT-125M125MHymba Table 625.6731.2561.9731.0429.0053.2041.29
6LaMini-GPT-124M124MHymba Table 626.4733.2662.8930.0527.8050.7540.95
7GPT-Neo-125M125M*Hymba Table 627.2531.3062.3529.6829.2051.5440.81
8GPT-2 small137MHymba Table 626.2931.0962.5129.7629.4049.7240.50
9Pollock 1.4127.7MCodeSOTA26.7133.1460.6129.3428.6050.2840.39
10Pythia-160M160M*Hymba Table 626.6831.9261.6429.5527.8049.4940.08
11Cerebras-GPT-111M111MHymba Table 625.5627.7558.1626.3225.4050.2837.58

How to read it: the mean averages ARC, PIQA, HellaSwag, OpenBookQA, and WinoGrande; MMLU is displayed but excluded. Higher is better. Pollock is a full 0-shot BF16 CodeSOTA run using the pinned SmolLM2 LightEval recipe; its source link opens the result record. The other rows remain frozen exactly as reported in Hymba Table 6. * The study uses model-family labels. Older artifact totals included buffers; the new grammar run counts unique learned weights: 125.2M for GPT-Neo and 162.3M for Pythia-160M. Source-reported Table 1 labels remain unchanged.

Table 2

Reported outside the reference run

Useful context, but not additional rows in Table 1. These values came from separate executions.

ModelCurrent paramsSource / protocolMMLUARC c+ePIQAHellaSwagOpenBookQAWinoGrande5-task mean
SmolLM2-135M134.5MSmolLM2 card · LightEval · 0-shot31.5043.9068.4042.1034.6051.3048.06

SmolLM2 remains here because its values are source-reported from its model card rather than the Hymba reference table. Its 48.06 mean is calculated from the five displayed model-card values. MMLU is the cloze/continuation task, not letter-choice MMLU.

Table 3

Small-model diagnostics

Tests that reveal language-model quality at this scale without changing the primary five-task ranking.

ModelMMLU continuationBLiMPSciQLAMBADA acc.LAMBADA PPL ↓WikiText-2 BPB ↓Source
Pollock 1.425.5278.0765.6027.8353.541.007CodeSOTA · full BF16 run

MMLU continuation scores answer text directly rather than selecting A–D. BLiMP tests grammatical minimal pairs; LAMBADA tests final-word prediction; WikiText-2 BPB is tokenizer-comparable and lower is better. These diagnostics are displayed separately because averaging unlike metrics would create an arbitrary composite.

Diagnostic protocol v2

Next measurements

Cells remain pending until the same pinned implementation has run for every comparison model. No placeholder score enters the table.

FamilyBenchmarkPublished measurementsPollock statusWhat it isolates
Language fitMulti-domain BPBBPB by domain ↓PendingTokenizer-comparable fit on prose, news, dialogue, scientific text, and code.
GrammarBLiMPAccuracy · category accuracy · mean margin ↑Aggregate completeAdds phenomenon-level results and confidence, not only correct/incorrect pairs.
GrammarBLiMP SupplementAccuracy · category accuracy ↑PendingExtends the minimal-pair grammar suite with broader lexical phenomena.
ContextLAMBADA OpenAIAccuracy ↑ · target NLL ↓ · MRR ↑ · top-5 ↑Accuracy + PPL completeContinuous target-token metrics retain signal below exact-match thresholds.
World knowledgeEWoKMacro accuracy · category accuracy ↑PendingPaired likelihood tests for elementary world knowledge; report distance from 50%.
Targeted syntaxSyntaxGymSuite accuracy · surprisal effect ↑PendingTests whether expected syntactic surprisal effects appear in controlled suites.

Reporting rule: publish raw benchmark columns and category breakdowns; do not create a single diagnostic average. Multi-domain BPB replaces cross-tokenizer word perplexity. EWoK receives confidence intervals because a small model may remain close to its 50% paired-choice baseline.

Table 4

Unified CodeSOTA rerun roster

Roster for the five-task rerun. Counts use learned weights where measured in the grammar run. Harness status below refers to task evaluation; grammar results are reported separately above.

ModelLearned paramsArchitectureTrackHarness status
Pollock 1.4127.7MTransformerCoreComplete · BF16
SmolLM2-135M134.5MTransformerCoreRun
MobileLLM-125M≈124.6MTransformerCoreRun
Mamba-130M129.1MState-spaceCoreRun
OPT-125M≈125MTransformerCoreRun
GPT-Neo-125M125.2MTransformerCoreRun
GPT-2 small124.4MTransformerCoreRun
Pythia-160M162.3MTransformerCoreRun
Cerebras-GPT-111M111MTransformerCoreRun
Pythia-70M70.4MTransformerBaselineRun
DistilGPT282MTransformerBaselineRun
Muse2-125M Base122.9MHybrid conv/attentionAdditionalAdapter needed
TLM-100M99.0MTiered GPT-NeoAdditionalCustom class needed
Protocol

How the fresh ranking should be run

Published scores remain frozen above. New CodeSOTA results get their own source marker and never silently replace source-reported values.

Primary tasks

ARC-Easy, ARC-Challenge, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Publish both ARC components and their arithmetic mean.

Primary rank

Unweighted mean of ARC c+e, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Higher is better.

Secondary tasks

MMLU continuation, SciQ, multi-domain BPB, BLiMP plus Supplement, LAMBADA target-token metrics, EWoK, and SyntaxGym. Show these separately; do not mix them into the main mean.

Controls

One pinned harness and dataset revision, 0-shot, full test splits, fixed checkpoint SHAs, matching precision, and saved raw outputs.