Action recognition benchmark with 101 action categories
3 results indexed across 1 metric. Shaded row marks current SOTA; ties broken by submission date.
| # | Model | Org | Submitted | Paper / code | accuracy |
|---|---|---|---|---|---|
| 01 | VideoMAE ViT-B | — | Mar 2022 | VideoMAE: Masked Autoencoders are Data-Efficient Learner… · code | 96.10 |
| 02 | DINOv3 (7B) | — | Aug 2025 | DINOv3 · code | 93.50 |
| 03 | DINOv2 (ViT-g/14) | — | Apr 2023 | DINOv2: Learning Robust Visual Features without Supervis… · code | 91.20 |
Every paper below corresponds to at least one row in the leaderboard above. Click through for the arXiv preprint and, when available, the reference implementation.
Submit a checkpoint and a reproduction script. We will run it, publish the score, and — if it takes the top — annotate the step on the progress chart with your name.