Codesota · Models · LLaVA8 results · 1 benchmarks
Model card

LLaVA.

Image CaptioningImage to TextApache 2.0

Strong open-source VLM. LLaVA-1.6 significantly improved.

§ 01 · Card

Model card,
inline.

Rendered server-side from the upstream README on Hugging Face — same content as the source repo, with editorial typography. The full card, sample weights, and revision history live on HF.


Source
liuhaotian/llava-v1.6-34b
License
apache-2.0
Pipeline
image-text-to-text

<br> <br>

LLaVA Model Card

Model details

Model type: LLaVA is an open-source chatbot trained by fine-tuning LLM on multimodal instruction-following data. It is an auto-regressive language model, based on the transformer architecture. Base LLM: NousResearch/Nous-Hermes-2-Yi-34B

Model date: LLaVA-v1.6-34B was trained in December 2023.

Paper or resources for more information: https://llava-vl.github.io/

License

NousResearch/Nous-Hermes-2-Yi-34B license.

Where to send questions or comments about the model: https://github.com/haotian-liu/LLaVA/issues

Intended use

Primary intended uses: The primary use of LLaVA is research on large multimodal models and chatbots.

Primary intended users: The primary intended users of the model are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.

Training dataset

  • 558K filtered image-text pairs from LAION/CC/SBU, captioned by BLIP.
  • 158K GPT-generated multimodal instruction-following data.
  • 500K academic-task-oriented VQA data mixture.
  • 50K GPT-4V data mixture.
  • 40K ShareGPT data.

Evaluation dataset

A collection of 12 benchmarks, including 5 academic VQA benchmarks and 7 recent benchmarks specifically proposed for instruction-following LMMs.

Card content reproduced from huggingface.co/liuhaotian/llava-v1.6-34b under the upstream license. Rendering trims fenced HTML, raw widgets and tables for safety; tap the link for the untouched original.
§ 02 · Benchmarks

Every benchmark LLaVA has a recorded score for.

#BenchmarkArea · TaskMetricValueRankDateSource
01demon-bench—multi-image-reasoning41.5%#10/112026-04-20source ↗
02demon-bench—grounded-qa36.2%#8/112026-04-20source ↗
03demon-bench—knowledge-images-qa28.3%#9/112026-04-20source ↗
04demon-bench—accuracy21.2%#11/112026-04-20source ↗
05demon-bench—relation-cloze15.8%#11/112026-04-20source ↗
06demon-bench—storytelling10.7%#11/112026-04-20source ↗
07demon-bench—visual-inference8.3%#9/112026-04-20source ↗
08demon-bench—multimodal-dialogue7.8%#11/112026-04-20source ↗
Rank column shows this model’s position vs all other models scored on the same benchmark + metric (competitors after the slash). #1 in red means current SOTA. Sorted by rank, then newest result.
§ 06 · Sources & freshness

Where these numbers come from.

codesota-api
8
results
8 of 8 rows marked verified.