Fine-grained temporal action understanding with objects
5 results indexed across 1 metric. Shaded row marks current SOTA; ties broken by submission date.
| # | Model | Org | Submitted | Paper / code | accuracy |
|---|---|---|---|---|---|
| 01 | V-JEPA 2 ViT-g (1B, 384px) | — | Jun 2025 | V-JEPA 2: Self-Supervised Video Models Enable Understand… · code | 77.30 |
| 02 | VideoMAE ViT-L | — | Mar 2022 | VideoMAE: Masked Autoencoders are Data-Efficient Learner… · code | 75.40 |
| 03 | DINOv3 (7B) | — | Aug 2025 | DINOv3 · code | 70.80 |
| 04 | VideoPrism-g | — | Feb 2024 | VideoPrism: A Foundational Visual Encoder for Video Unde… · code | 68.50 |
| 05 | DINOv2 (ViT-g/14) | — | Apr 2023 | DINOv2: Learning Robust Visual Features without Supervis… · code | 38.30 |
Every paper below corresponds to at least one row in the leaderboard above. Click through for the arXiv preprint and, when available, the reference implementation.
Submit a checkpoint and a reproduction script. We will run it, publish the score, and — if it takes the top — annotate the step on the progress chart with your name.