LLM/benchmark ladder/small English models

Small English language models

Published benchmark results for compact base models around the 125M class. Parameter count is descriptive, not an eligibility cutoff.

Best same-harness mean
Hymba-125M · 49.35
Models in LightEval ranking
11
Models in rerun roster
13
Primary protocol
0-shot · 5-task mean
Table 1

LightEval small-model ranking

Ten source-reported rows from Hymba Table 6 plus the new CodeSOTA Pollock run. The source column keeps the executions explicit.

1Hymba-125M125MHymba Table 631.1244.9568.5045.5435.5252.2549.35
2SmolLM-135M135MHymba Table 630.2343.9969.6042.3033.6052.7048.44
3MobileLLM-125M125MHymba Table 635.5165.3038.9039.5053.1046.46
4Mamba-130M130MHymba Table 627.4133.0163.3333.8630.4051.5442.43
5OPT-125M125MHymba Table 625.6731.2561.9731.0429.0053.2041.29
6LaMini-GPT-124M124MHymba Table 626.4733.2662.8930.0527.8050.7540.95
7GPT-Neo-125M125M*Hymba Table 627.2531.3062.3529.6829.2051.5440.81
8GPT-2 small137MHymba Table 626.2931.0962.5129.7629.4049.7240.50
9Pollock 1.4127.7MCodeSOTA26.7133.1460.6129.3428.6050.2840.39
10Pythia-160M160M*Hymba Table 626.6831.9261.6429.5527.8049.4940.08
11Cerebras-GPT-111M111MHymba Table 625.5627.7558.1626.3225.4050.2837.58

How to read it: the mean averages ARC, PIQA, HellaSwag, OpenBookQA, and WinoGrande; MMLU is displayed but excluded. Higher is better. Pollock is a full 0-shot BF16 CodeSOTA run using the pinned SmolLM2 LightEval recipe; its source link opens the result record. The other rows remain frozen exactly as reported in Hymba Table 6. * The study uses model-family labels: current Hugging Face artifacts report 150.4M parameters for GPT-Neo-125M and 212.7M for Pythia-160M.

Table 2

Reported outside the reference run

Useful context, but not additional rows in Table 1. These values came from separate executions.

ModelCurrent paramsSource / protocolMMLUARC c+ePIQAHellaSwagOpenBookQAWinoGrande5-task mean
SmolLM2-135M134.5MSmolLM2 card · LightEval · 0-shot31.5043.9068.4042.1034.6051.3048.06

SmolLM2 remains here because its values are source-reported from its model card rather than the Hymba reference table. Its 48.06 mean is calculated from the five displayed model-card values. MMLU is the cloze/continuation task, not letter-choice MMLU.

Table 3

Small-model diagnostics

Tests that reveal language-model quality at this scale without changing the primary five-task ranking.

ModelMMLU continuationBLiMPSciQLAMBADA acc.LAMBADA PPL ↓WikiText-2 BPB ↓Source
Pollock 1.425.5278.0765.6027.8353.541.007CodeSOTA · full BF16 run

MMLU continuation scores answer text directly rather than selecting A–D. BLiMP tests grammatical minimal pairs; LAMBADA tests final-word prediction; WikiText-2 BPB is tokenizer-comparable and lower is better. These diagnostics are displayed separately because averaging unlike metrics would create an arbitrary composite.

Diagnostic protocol v2

Next measurements

Cells remain pending until the same pinned implementation has run for every comparison model. No placeholder score enters the table.

FamilyBenchmarkPublished measurementsPollock statusWhat it isolates
Language fitMulti-domain BPBBPB by domain ↓PendingTokenizer-comparable fit on prose, news, dialogue, scientific text, and code.
GrammarBLiMPAccuracy · category accuracy · mean margin ↑Aggregate completeAdds phenomenon-level results and confidence, not only correct/incorrect pairs.
GrammarBLiMP SupplementAccuracy · category accuracy ↑PendingExtends the minimal-pair grammar suite with broader lexical phenomena.
ContextLAMBADA OpenAIAccuracy ↑ · target NLL ↓ · MRR ↑ · top-5 ↑Accuracy + PPL completeContinuous target-token metrics retain signal below exact-match thresholds.
World knowledgeEWoKMacro accuracy · category accuracy ↑PendingPaired likelihood tests for elementary world knowledge; report distance from 50%.
Targeted syntaxSyntaxGymSuite accuracy · surprisal effect ↑PendingTests whether expected syntactic surprisal effects appear in controlled suites.

Reporting rule: publish raw benchmark columns and category breakdowns; do not create a single diagnostic average. Multi-domain BPB replaces cross-tokenizer word perplexity. EWoK receives confidence intervals because a small model may remain close to its 50% paired-choice baseline.

Table 4

Unified CodeSOTA rerun roster

This is the model set that can answer the comparison cleanly. The exact total parameter count is shown where repository metadata differs from the model name.

ModelCurrent total paramsArchitectureTrackHarness status
Pollock 1.4127.7MTransformerCoreComplete · BF16
SmolLM2-135M134.5MTransformerCoreRun
MobileLLM-125M≈124.6MTransformerCoreRun
Mamba-130M129.1MState-spaceCoreRun
OPT-125M≈125MTransformerCoreRun
GPT-Neo-125M150.4MTransformerCoreRun
GPT-2 small137.0MTransformerCoreRun
Pythia-160M212.7MTransformerCoreRun
Cerebras-GPT-111M111MTransformerCoreRun
Pythia-70M95.6MTransformerBaselineRun
DistilGPT282MTransformerBaselineRun
Muse2-125M Base122.9MHybrid conv/attentionAdditionalAdapter needed
TLM-100M99.0MTiered GPT-NeoAdditionalCustom class needed
Protocol

How the fresh ranking should be run

Published scores remain frozen above. New CodeSOTA results get their own source marker and never silently replace source-reported values.

Primary tasks

ARC-Easy, ARC-Challenge, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Publish both ARC components and their arithmetic mean.

Primary rank

Unweighted mean of ARC c+e, PIQA, HellaSwag, OpenBookQA, and WinoGrande. Higher is better.

Secondary tasks

MMLU continuation, SciQ, multi-domain BPB, BLiMP plus Supplement, LAMBADA target-token metrics, EWoK, and SyntaxGym. Show these separately; do not mix them into the main mean.

Controls

One pinned harness and dataset revision, 0-shot, full test splits, fixed checkpoint SHAs, matching precision, and saved raw outputs.