Open SOTA registryRegistry snapshot 2026-04-27

Find the best AI model
for the job.

Independent benchmark evidence across models, tasks, and modalities. Every published score is dated, source-tiered, and linked back to where it came from.

A leaderboard row is not a fact until it can be inspected.
Try
9,102Results
163Models tracked
371Benchmarks indexed
121Tasks
9Capability areas
70Published SOTA
§ 01 · Inspectable frontier

What can the evidence support right now?

One protocol-specific leader per benchmark. A model is printed only when the row is verified and its source URL is inspectable.

General document OCROCRBenchqwen3-5-397b-a17bscore931.0✓ Inspect source
Multilingual document OCROCRBench v2 · Englishovis2-5-9boverall-en-private63.4✓ Inspect source
Document parsingOmniDocBenchGLM-OCRcomposite94.62✓ Inspect source
PDF conversion qualityolmOCR-Benchinfinity-parser2-propass-rate87.6%✓ Inspect source
Terminal coding agentsTerminal-Bench 2Codex / GPT-5.5accuracy82.0%✓ Inspect source
Software engineering agentsSWE-bench VerifiedClaude Opus 4.7resolve-rate87.6%✓ Inspect source

Leaders are computed only within the named metric and protocol. Incomplete or malformed evidence is withheld rather than silently promoted.

§ 02 · Choose by task

Start from the job, not the model.

§ 03 · Evidence

A leaderboard is only as good as the trail behind it.

Measured results, operational claims, and source quality stay separate so a clean-looking table cannot hide a weak comparison.

Every public number carries its audit trail.

CodeSOTA preserves metric direction, the exact benchmark variant, source type, and snapshot used. A vendor claim can be useful, but it is never presented as an independent reproduction.

  • Verified — inspectable evidence meets the publication floor
  • Vendor-reported — primary claim, not independently reproduced
  • Reproduced — an independent run and protocol are available
  • Withheld — the available row does not meet the evidence floor
Read the methodology →
§ 04 · Decision quality

A top score is the start of a decision, not the end.

Benchmark health tells you whether a table still separates models. The decision workspace then adds hardware, openness, language, and evidence constraints.

Matched-snapshot signal
Separating

ARC-AGI-1

31.4-point spread in the current matched frontier snapshot.

Protocol-sensitive

SWE-bench Verified

A strong signal only when harness and scaffold details travel with the score.

Low separation

ARC-Challenge

0.8-point spread across the matched three-model snapshot.

Saturated snapshot

GSM8K

No spread in the matched snapshot; use harder successors for frontier choices.

These are observations from one matched CodeSOTA snapshot, not permanent benchmark properties. Inspect the research question →

§ 05 · Original research

The registry remembers how a result became a claim.

This is the useful part of the older CodeSOTA homepage: public experiments, negative results, and lineage—not just another table of scores.

Public research graph

Question → experiment → evidence → claim → finding

Follow a decision back through its hypotheses, protocol, evidence, corrections, and negative results. The graph keeps research memory inspectable instead of flattening it into a marketing number.

Explore the research graph →
Questions
7
Experiments
14
Claims
18
Findings
6
§ 06 · API

The same registry, callable by software.

Agents should not scrape leaderboards. Call a stable, source-aware endpoint and retain the snapshot id so the number in a report can be reproduced later.

  • Public, CORS-open JSON with task-first rankings
  • Benchmark, metric direction, result date, and snapshot id
  • Short task aliases for OCR, code, ASR, TTS, and more
Read the API docs →
GET /api/sotacurl "https://codesota.com/api/sota/ocr?tier=sota"

{
  "task": "ocr",
  "tier": "sota",
  "pick": {
    "model_name": "ovis2-5-9b",
    "score": 63.4,
    "score_metric": "overall-en-private",
    "higher_is_better": true
  },
  "snapshot_id": "2026-04-27"
}
§ 07 · Alerts

Tell CodeSOTA what to watch.

Use your own words. CodeSOTA maps the request to benchmark rules, shows the interpretation back, and falls back to a weekly digest when the request is broad.

Free weekly digestDouble opt-inUnsubscribe readyNo search text sent to analytics
§ 08 · Access

The rankings stay free. Paid access funds delivery and private evaluation.

Ranking position, source links, corrections, and verification outcomes are not for sale. Paid plans cover API volume and work done around the public registry.

Public registryOpen evidence
$0forever

For anyone checking a number before they cite it.

  • Current rankings and task pages
  • Source tier, metric direction, and snapshot date
  • Every public source URL and correction trail
  • CORS-open /api/sota endpoint with no key
Browse the registry
Custom benchmarkPrivate evaluation
$2,000one-time

For a model choice that public leaderboards cannot answer.

  • Private hold-out benchmark on your data
  • Candidate-model evaluation and recommendation
  • Written decision report
  • Reusable harness and rubric you keep
Scope an evaluation

Paid work is labelled wherever it appears. A customer cannot buy a better rank, suppress an unfavourable result, or remove a source trail.

Compare every API plan →
§ 09 · Registry loop

Maintained in public—and open to correction.

Full activity log →
  • 9,102 results tracked across 163 models

  • + 14 public experiments retain their protocol and outcome

  • ± Superseded claims remain visible in the provenance trail