Codesota · Benchmark · DEMON BenchHome/Leaderboards/Multimodal Media/Any-to-Any Omni Models/DEMON Bench
Unknown

DEMON Bench.

Evaluates any-to-any multimodal models across diverse modality combinations

Paper ↗Leaderboard ↓
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

Multi Image Reasoning

Multi Image Reasoning is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Multi Image Reasoningverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Multi-Image Reasoning subtask score.
verified53.652024Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Multi-Image Reasoning subtask score.
verified50.282024Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Multi-Image Reasoning subtask score.
verified48.682024Source ↗Looks wrong?
04InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified48.552024Source ↗Looks wrong?
05LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified44.032024Source ↗Looks wrong?
06Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified43.852024Source ↗Looks wrong?
07MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified43.52024Source ↗Looks wrong?
08mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified42.52024Source ↗Looks wrong?
09OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified41.632024Source ↗Looks wrong?
10LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified41.532024Source ↗Looks wrong?
11BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Multi-Image Reasoning subtask score.
verified39.652024Source ↗Looks wrong?

Grounded Qa

Grounded Qa is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Grounded Qaverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Grounded QA subtask score.
verified52.932024Source ↗Looks wrong?
02Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Grounded QA subtask score.
verified512024Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Grounded QA subtask score.
verified48.62024Source ↗Looks wrong?
04InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified47.42024Source ↗Looks wrong?
05LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified44.82024Source ↗Looks wrong?
06Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified41.672024Source ↗Looks wrong?
07BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified39.232024Source ↗Looks wrong?
08LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified36.22024Source ↗Looks wrong?
09mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified33.272024Source ↗Looks wrong?
10OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified322024Source ↗Looks wrong?
11MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Grounded QA subtask score.
verified30.272024Source ↗Looks wrong?

Knowledge Images Qa

Knowledge Images Qa is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Knowledge Images Qaverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Knowledge Images QA subtask score.
verified49.332024Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Knowledge Images QA subtask score.
verified44.932024Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Knowledge Images QA subtask score.
verified44.932024Source ↗Looks wrong?
04InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified44.42024Source ↗Looks wrong?
05BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified33.532024Source ↗Looks wrong?
06mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified32.472024Source ↗Looks wrong?
07LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified322024Source ↗Looks wrong?
08OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified30.62024Source ↗Looks wrong?
09LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified28.332024Source ↗Looks wrong?
10Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified27.732024Source ↗Looks wrong?
11MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Knowledge Images QA subtask score.
verified26.42024Source ↗Looks wrong?

Multimodal Dialogue

Multimodal Dialogue is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Multimodal Dialogueverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Multimodal Dialogue subtask score.
verified42.72024Source ↗Looks wrong?
02Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Multimodal Dialogue subtask score.
verified38.142024Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Multimodal Dialogue subtask score.
verified37.52024Source ↗Looks wrong?
04InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified33.582024Source ↗Looks wrong?
05BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified26.122024Source ↗Looks wrong?
06OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified16.882024Source ↗Looks wrong?
07Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified15.372024Source ↗Looks wrong?
08LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified14.222024Source ↗Looks wrong?
09MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified13.692024Source ↗Looks wrong?
10mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified12.672024Source ↗Looks wrong?
11LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Multimodal Dialogue subtask score.
verified7.792024Source ↗Looks wrong?

Accuracy

Accuracy is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Average accuracy across all 7 DEMON subtasks.
verified39.282024Source ↗Looks wrong?
02Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Average accuracy across all 7 DEMON subtasks.
verified37.222024Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Average accuracy across all 7 DEMON subtasks.
verified36.372024Source ↗Looks wrong?
04InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified332024Source ↗Looks wrong?
05BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified26.922024Source ↗Looks wrong?
06LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified26.32024Source ↗Looks wrong?
07OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified25.832024Source ↗Looks wrong?
08Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified24.512024Source ↗Looks wrong?
09mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified23.132024Source ↗Looks wrong?
10MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified22.212024Source ↗Looks wrong?
11LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Average accuracy across all 7 DEMON subtasks.
verified21.242024Source ↗Looks wrong?

Relation Cloze

Relation Cloze is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Relation Clozeverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Relation Cloze subtask score.
verified27.152024Source ↗Looks wrong?
02Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Relation Cloze subtask score.
verified22.952024Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Relation Cloze subtask score.
verified22.152024Source ↗Looks wrong?
04OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified21.652024Source ↗Looks wrong?
05InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified21.22024Source ↗Looks wrong?
06LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified182024Source ↗Looks wrong?
07BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified17.942024Source ↗Looks wrong?
08MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified16.62024Source ↗Looks wrong?
09mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified16.252024Source ↗Looks wrong?
10Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified162024Source ↗Looks wrong?
11LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Relation Cloze subtask score.
verified15.852024Source ↗Looks wrong?

Visual Inference

Visual Inference is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Visual Inferenceverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Visual Inference subtask score.
verified27.152024Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Visual Inference subtask score.
verified25.92024Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Visual Inference subtask score.
verified25.52024Source ↗Looks wrong?
04OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified13.852024Source ↗Looks wrong?
05LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified13.512024Source ↗Looks wrong?
06InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified11.492024Source ↗Looks wrong?
07Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified11.392024Source ↗Looks wrong?
08BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified10.672024Source ↗Looks wrong?
09LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified8.272024Source ↗Looks wrong?
10MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified7.952024Source ↗Looks wrong?
11mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Visual Inference subtask score.
verified5.402024Source ↗Looks wrong?

Storytelling

Storytelling is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Storytellingverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Cheetah with VPG-C and Vicuna-13B backbone. Best Cheetah variant. Zhejiang University / NUS. Storytelling subtask score.
verified26.592024Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Cheetah with VPG-C and Vicuna-7B backbone. Main result from the paper (ICLR 2024 Spotlight). Zhejiang University / NUS. Storytelling subtask score.
verified25.22024Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Cheetah with VPG-C and LLaMA2-7B backbone. From Table 3 ablations. Zhejiang University / NUS. Storytelling subtask score.
verified24.762024Source ↗Looks wrong?
04InstructBLIP
InstructBLIP with Vicuna-7B. Salesforce, NeurIPS 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified24.412024Source ↗Looks wrong?
05OpenFlamingo
OpenFlamingo-9B (open-source Flamingo). University of Washington, 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified24.222024Source ↗Looks wrong?
06BLIP-2
BLIP-2 (Bootstrapping Language-Image Pre-training 2) with FlanT5-XXL. Salesforce, ICML 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified21.312024Source ↗Looks wrong?
07mPLUG-Owl
mPLUG-Owl (Modularized Multimodal Large Language Model). Alibaba DAMO, 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified19.332024Source ↗Looks wrong?
08LLaMA-Adapter V2
LLaMA-Adapter V2 (7B). Shanghai AI Lab, 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified17.572024Source ↗Looks wrong?
09MiniGPT-4
MiniGPT-4 (aligning large language model with advanced large vision model). KAUST, 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified17.072024Source ↗Looks wrong?
10Otter
Otter (multi-modal model built on OpenFlamingo). Nanyang Technological University, 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified15.572024Source ↗Looks wrong?
11LLaVA
LLaVA (Large Language and Vision Assistant, original v1). UW-Madison, NeurIPS 2023. Baseline on DEMON-Core. Storytelling subtask score.
verified10.72024Source ↗Looks wrong?
§ 04 · Submit a result

Add to the leaderboard.

← Back to Any-to-Any Omni Models