Evaluates any-to-any multimodal models across diverse modality combinations
Multi Image Reasoning is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 53.65 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 50.28 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 48.68 | 2024 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 48.55 | 2024 | Source ↗ | Looks wrong? |
| 05 | LLaMA-Adapter V2 | verified | 44.03 | 2024 | Source ↗ | Looks wrong? |
| 06 | Otter | verified | 43.85 | 2024 | Source ↗ | Looks wrong? |
| 07 | MiniGPT-4 | verified | 43.5 | 2024 | Source ↗ | Looks wrong? |
| 08 | mPLUG-Owl | verified | 42.5 | 2024 | Source ↗ | Looks wrong? |
| 09 | OpenFlamingo | verified | 41.63 | 2024 | Source ↗ | Looks wrong? |
| 10 | LLaVA | verified | 41.53 | 2024 | Source ↗ | Looks wrong? |
| 11 | BLIP-2 | verified | 39.65 | 2024 | Source ↗ | Looks wrong? |
Grounded Qa is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 52.93 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (LLaMA2-7B) | verified | 51 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 48.6 | 2024 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 47.4 | 2024 | Source ↗ | Looks wrong? |
| 05 | LLaMA-Adapter V2 | verified | 44.8 | 2024 | Source ↗ | Looks wrong? |
| 06 | Otter | verified | 41.67 | 2024 | Source ↗ | Looks wrong? |
| 07 | BLIP-2 | verified | 39.23 | 2024 | Source ↗ | Looks wrong? |
| 08 | LLaVA | verified | 36.2 | 2024 | Source ↗ | Looks wrong? |
| 09 | mPLUG-Owl | verified | 33.27 | 2024 | Source ↗ | Looks wrong? |
| 10 | OpenFlamingo | verified | 32 | 2024 | Source ↗ | Looks wrong? |
| 11 | MiniGPT-4 | verified | 30.27 | 2024 | Source ↗ | Looks wrong? |
Knowledge Images Qa is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 49.33 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 44.93 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 44.93 | 2024 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 44.4 | 2024 | Source ↗ | Looks wrong? |
| 05 | BLIP-2 | verified | 33.53 | 2024 | Source ↗ | Looks wrong? |
| 06 | mPLUG-Owl | verified | 32.47 | 2024 | Source ↗ | Looks wrong? |
| 07 | LLaMA-Adapter V2 | verified | 32 | 2024 | Source ↗ | Looks wrong? |
| 08 | OpenFlamingo | verified | 30.6 | 2024 | Source ↗ | Looks wrong? |
| 09 | LLaVA | verified | 28.33 | 2024 | Source ↗ | Looks wrong? |
| 10 | Otter | verified | 27.73 | 2024 | Source ↗ | Looks wrong? |
| 11 | MiniGPT-4 | verified | 26.4 | 2024 | Source ↗ | Looks wrong? |
Multimodal Dialogue is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (LLaMA2-7B) | verified | 42.7 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-13B) | verified | 38.14 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 37.5 | 2024 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 33.58 | 2024 | Source ↗ | Looks wrong? |
| 05 | BLIP-2 | verified | 26.12 | 2024 | Source ↗ | Looks wrong? |
| 06 | OpenFlamingo | verified | 16.88 | 2024 | Source ↗ | Looks wrong? |
| 07 | Otter | verified | 15.37 | 2024 | Source ↗ | Looks wrong? |
| 08 | LLaMA-Adapter V2 | verified | 14.22 | 2024 | Source ↗ | Looks wrong? |
| 09 | MiniGPT-4 | verified | 13.69 | 2024 | Source ↗ | Looks wrong? |
| 10 | mPLUG-Owl | verified | 12.67 | 2024 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 7.79 | 2024 | Source ↗ | Looks wrong? |
Accuracy is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 39.28 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (LLaMA2-7B) | verified | 37.22 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 36.37 | 2024 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 33 | 2024 | Source ↗ | Looks wrong? |
| 05 | BLIP-2 | verified | 26.92 | 2024 | Source ↗ | Looks wrong? |
| 06 | LLaMA-Adapter V2 | verified | 26.3 | 2024 | Source ↗ | Looks wrong? |
| 07 | OpenFlamingo | verified | 25.83 | 2024 | Source ↗ | Looks wrong? |
| 08 | Otter | verified | 24.51 | 2024 | Source ↗ | Looks wrong? |
| 09 | mPLUG-Owl | verified | 23.13 | 2024 | Source ↗ | Looks wrong? |
| 10 | MiniGPT-4 | verified | 22.21 | 2024 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 21.24 | 2024 | Source ↗ | Looks wrong? |
Relation Cloze is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 27.15 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (LLaMA2-7B) | verified | 22.95 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 22.15 | 2024 | Source ↗ | Looks wrong? |
| 04 | OpenFlamingo | verified | 21.65 | 2024 | Source ↗ | Looks wrong? |
| 05 | InstructBLIP | verified | 21.2 | 2024 | Source ↗ | Looks wrong? |
| 06 | LLaMA-Adapter V2 | verified | 18 | 2024 | Source ↗ | Looks wrong? |
| 07 | BLIP-2 | verified | 17.94 | 2024 | Source ↗ | Looks wrong? |
| 08 | MiniGPT-4 | verified | 16.6 | 2024 | Source ↗ | Looks wrong? |
| 09 | mPLUG-Owl | verified | 16.25 | 2024 | Source ↗ | Looks wrong? |
| 10 | Otter | verified | 16 | 2024 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 15.85 | 2024 | Source ↗ | Looks wrong? |
Visual Inference is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 27.15 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 25.9 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 25.5 | 2024 | Source ↗ | Looks wrong? |
| 04 | OpenFlamingo | verified | 13.85 | 2024 | Source ↗ | Looks wrong? |
| 05 | LLaMA-Adapter V2 | verified | 13.51 | 2024 | Source ↗ | Looks wrong? |
| 06 | InstructBLIP | verified | 11.49 | 2024 | Source ↗ | Looks wrong? |
| 07 | Otter | verified | 11.39 | 2024 | Source ↗ | Looks wrong? |
| 08 | BLIP-2 | verified | 10.67 | 2024 | Source ↗ | Looks wrong? |
| 09 | LLaVA | verified | 8.27 | 2024 | Source ↗ | Looks wrong? |
| 10 | MiniGPT-4 | verified | 7.95 | 2024 | Source ↗ | Looks wrong? |
| 11 | mPLUG-Owl | verified | 5.40 | 2024 | Source ↗ | Looks wrong? |
Storytelling is the reported evaluation metric for DEMON Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 26.59 | 2024 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 25.2 | 2024 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 24.76 | 2024 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 24.41 | 2024 | Source ↗ | Looks wrong? |
| 05 | OpenFlamingo | verified | 24.22 | 2024 | Source ↗ | Looks wrong? |
| 06 | BLIP-2 | verified | 21.31 | 2024 | Source ↗ | Looks wrong? |
| 07 | mPLUG-Owl | verified | 19.33 | 2024 | Source ↗ | Looks wrong? |
| 08 | LLaMA-Adapter V2 | verified | 17.57 | 2024 | Source ↗ | Looks wrong? |
| 09 | MiniGPT-4 | verified | 17.07 | 2024 | Source ↗ | Looks wrong? |
| 10 | Otter | verified | 15.57 | 2024 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 10.7 | 2024 | Source ↗ | Looks wrong? |