Codesota · Computer Vision · Image Classification · ImageNet-1KTasks/Computer Vision/Image Classification
Image Classification · benchmark dataset · 2012 · EN

ImageNet Large Scale Visual Recognition Challenge 2012.

1.28M training images, 50K validation images across 1,000 object classes. The standard benchmark for image classification since 2012.

Saturated benchmark

Benchmark near ceiling or stagnant — no meaningful SOTA movement in 2+ years

Paper Download datasetSubmit a result
§ 01 · Leaderboard

Best published scores.

47 results indexed across 3 metrics. Shaded row marks current SOTA; ties broken by submission date.


Primary
top-1-accuracy · higher is better
All metrics
accuracy, pass@1, top-1-accuracy
accuracy
26 rows
#ModelOrgSubmittedPaper / codeaccuracy
01BEiT-L+Jun 2021BEiT: BERT Pre-Training of Image Transformers · code89.50
02AIMv2 ViT-3B/14 448pxNov 2024Multimodal Autoregressive Pre-training of Large Vision E… · code89.50
03ALIGNFeb 2021Scaling Up Visual and Vision-Language Representation Lea… · code88.64
04Vision Transformer (ViT-H/14)Oct 2020An Image is Worth 16x16 Words: Transformers for Image Re… · code88.55
05DINOv3 (7B)Aug 2025DINOv3 · code88.40
06MAE (ViT-H, 448)Nov 2021Masked Autoencoders Are Scalable Vision Learners · code87.80
07ConvNeXt (XL)Jan 2022A ConvNet for the 2020s · code87.80
08BiT-LDec 2019Big Transfer (BiT): General Visual Representation Learni… · code87.54
09DINOv2 (ViT-g/14)Apr 2023DINOv2: Learning Robust Visual Features without Supervis… · code86.50
10V-JEPA 2 ViT-g (1B, 384px)Jun 2025V-JEPA 2: Self-Supervised Video Models Enable Understand… · code85.10
11SigLIP 2 (g/16)Feb 2025SigLIP 2: Multilingual Vision-Language Encoders with Imp… · code85
12ResNet-152OpenMicrosoftDec 2015Deep Residual Learning for Image Recognition · code80.62
13DINO (ViT-B/8)Apr 2021Emerging Properties in Self-Supervised Vision Transforme… · code80.10
14YOLO26x-clsJan 2026pwc-dump · code79.90
15YOLO26l-clsJan 2026pwc-dump · code79
16YOLOv8x-clsJan 2023pwc-dump · code79
17YOLO26m-clsJan 2026pwc-dump · code78.10
18YOLOv8m-clsJan 2023pwc-dump · code76.80
19YOLOv8l-clsJan 2023pwc-dump · code76.80
20CLIPFeb 2021Learning Transferable Visual Models From Natural Languag… · code76.20
21YOLO26s-clsJan 2026pwc-dump · code76
22AltCLIPNov 2022AltCLIP: Altering the Language Encoder in CLIP for Exten… · code74.50
23YOLOv8s-clsJan 2023pwc-dump · code73.80
24YOLO26n-clsJan 2026pwc-dump · code71.40
25YOLOv8n-clsJan 2023pwc-dump · code69
26CN-CLIPNov 2022Chinese CLIP: Contrastive Vision-Language Pretraining in… · code59.60
pass@1
1 row
#ModelOrgSubmittedPaper / codepass@1
01pMF-H + FD-lossOpenN/Apaper0.720
top-1-accuracy· primary
20 rows
#ModelOrgSubmittedPaper / codetop-1-accuracy
01CoCa (finetuned)OpenGoogleDec 2025google-research91
02ViT-G/14OpenGoogleDec 2025google-research90.45
03SoViT-400m/14OpenGoogle DeepMindApr 2026neurips-202390.30
04AIMv2-3BOpenAppleApr 2026arxiv-paper89.50
05ConvNeXt V2 HugeOpenMetaDec 2025meta-research88.90
06ViT-H/14OpenGoogleDec 2025google-research88.55
07InternViT-6B (InternVL)OpenOpenGVLabApr 2026cvpr-202488.23
08Swin Transformer LargeOpenMicrosoftDec 2025microsoft-research87.30
09EfficientNetV2-LOpenGoogleDec 2025google-research85.70
10MambaVision-L2OpenNVIDIAApr 2026cvpr-202585.30
11DeiT-B DistilledOpenMetaDec 2025meta-research85.20
12EfficientNet-B7OpenGoogleDec 2025google-research84.40
13DeiT-BOpenMetaDec 2025meta-research83.10
14ConvNeXt V2 TinyOpenMetaDec 2025meta-research83
15ViT-L/16OpenGoogleDec 2025google-research82.70
16ViT-B/16OpenGoogleDec 2025google-research81.20
17ResNet-50 (A3 training)OpenTimmDec 2025timm-research80.40
18ResNet-152OpenMicrosoftDec 2025microsoft-research78.60
19EfficientNet-B0OpenGoogleDec 2025google-research77.10
20ResNet-50OpenMicrosoftDec 2025pytorch-vision76.15
Fig 2 · Rows sorted by score within each metric. Shaded row marks SOTA. Dates reflect model or paper release where available, otherwise the date Codesota accessed the source.
§ 03 · Progress

1 steps
of state of the art.

Each row below marks a model that broke the previous record on top-1-accuracy. Intermediate submissions are kept in the leaderboard above; only SOTA-setting entries are re-listed here.

Higher scores win. Each subsequent entry improved upon the previous best.

SOTA line · top-1-accuracy
  1. Dec 18, 2025CoCa (finetuned)Google91
Fig 3 · SOTA-setting models only. 1 entries span Dec 2025 Dec 2025.
§ 04 · Literature

16 papers
tied to this benchmark.

Every paper below corresponds to at least one row in the leaderboard above. Click through for the arXiv preprint and, when available, the reference implementation.

§ 06 · Contribute

Have a score that beats
this table?

Submit a checkpoint and a reproduction script. We will run it, publish the score, and — if it takes the top — annotate the step on the progress chart with your name.

Submit a result Read submission guide
What a submission needs
  • 01A public checkpoint or API endpoint
  • 02A reproduction script with frozen commit + seed
  • 03Declared evaluation environment (Python, deps)
  • 04One row per metric declared by this dataset
  • 05A contact so we can follow up on discrepancies
ImageNet-1K — Image Classification | CodeSOTA