Vision & documents
OCR · parsing · VLMs
OCR →Document parsing →Polish OCR →Independent benchmark evidence across models, tasks, and modalities. Every published score is dated, source-tiered, and linked back to where it came from.
A leaderboard row is not a fact until it can be inspected.
One protocol-specific leader per benchmark. A model is printed only when the row is verified and its source URL is inspectable.
Leaders are computed only within the named metric and protocol. Incomplete or malformed evidence is withheld rather than silently promoted.
OCR · parsing · VLMs
OCR →Document parsing →Polish OCR →Reasoning · knowledge · search
LLMs →Polish LLM →Embeddings →Generation · repair · tool use
Code generation →Agent benchmarks →RL environments →ASR · TTS · audio
Speech-to-text →Text-to-speech →Audio →Medical · inspection · robotics
Medical AI →Industrial →Robotics →Measured results, operational claims, and source quality stay separate so a clean-looking table cannot hide a weak comparison.
OCRBench v2 · English private split
CodeSOTA preserves metric direction, the exact benchmark variant, source type, and snapshot used. A vendor claim can be useful, but it is never presented as an independent reproduction.
Benchmark health tells you whether a table still separates models. The decision workspace then adds hardware, openness, language, and evidence constraints.
31.4-point spread in the current matched frontier snapshot.
A strong signal only when harness and scaffold details travel with the score.
0.8-point spread across the matched three-model snapshot.
No spread in the matched snapshot; use harder successors for frontier choices.
These are observations from one matched CodeSOTA snapshot, not permanent benchmark properties. Inspect the research question →
Filter the registry before comparing scores. A slightly lower benchmark result may be the correct choice when it is open, local, cheaper, or proven on your language.
I need a local OCR model for Polish invoices that fits in 24 GB VRAM→This is the useful part of the older CodeSOTA homepage: public experiments, negative results, and lineage—not just another table of scores.
Follow a decision back through its hypotheses, protocol, evidence, corrections, and negative results. The graph keeps research memory inspectable instead of flattening it into a marketing number.
Explore the research graph →Source-aware experiments, PR transitions, and bits-per-byte evidence.
Open the lineage →Original model workCheckpoint, configuration, training evidence, and honest evaluation limits.
Read the report →Evaluation infrastructureRanked by discriminative power, saturation risk, and available evidence.
Browse RL environments →Agents should not scrape leaderboards. Call a stable, source-aware endpoint and retain the snapshot id so the number in a report can be reproduced later.
GET /api/sotacurl "https://codesota.com/api/sota/ocr?tier=sota"
{
"task": "ocr",
"tier": "sota",
"pick": {
"model_name": "ovis2-5-9b",
"score": 63.4,
"score_metric": "overall-en-private",
"higher_is_better": true
},
"snapshot_id": "2026-04-27"
}Use your own words. CodeSOTA maps the request to benchmark rules, shows the interpretation back, and falls back to a weekly digest when the request is broad.
Broad request detected. Defaulting to a weekly digest.
Ranking position, source links, corrections, and verification outcomes are not for sale. Paid plans cover API volume and work done around the public registry.
For anyone checking a number before they cite it.
For teams that need a stable model-routing surface in production.
For a model choice that public leaderboards cannot answer.
Paid work is labelled wherever it appears. A customer cannot buy a better rank, suppress an unfavourable result, or remove a source trail.
Compare every API plan →✓ 9,102 results tracked across 163 models
+ 14 public experiments retain their protocol and outcome
± Superseded claims remain visible in the provenance trail