Open-weight AI training-data modality disclosures, 2024-2025

Ten major open-weight model families by disclosed modality, 2024-2025, category matrix.

What the Data Shows

Training-data modality disclosures are compared for 10 major open-weight model families released in 2024-2025. The matrix lists each model family with organization country, input modalities, disclosed training-data modalities and release-term bucket. Four of the ten families disclose multimodal input or training modalities, while six are text-only or text-first in the selected evidence. The matrix does not compare training-data volume because the official cards disclose modality detail unevenly across families.

Data table

Meta Llama 4 — Country U.S.
Meta Llama 4 — Inputs Text+image
Meta Llama 4 — Training Text+vision
Meta Llama 4 — Terms Custom
Google Gemma 3 — Country U.S.
Google Gemma 3 — Inputs Text+image
Google Gemma 3 — Training Text+multi
Google Gemma 3 — Terms Custom
Microsoft Phi-4 multimodal — Country U.S.
Microsoft Phi-4 multimodal — Inputs Text+img+aud
Microsoft Phi-4 multimodal — Training Text+speech
Microsoft Phi-4 multimodal — Terms MIT
Mistral Small 3.1 — Country France
Mistral Small 3.1 — Inputs Text+image
Mistral Small 3.1 — Training Image-text
Mistral Small 3.1 — Terms Apache 2.0
IBM Granite 3.3 — Country U.S.
IBM Granite 3.3 — Inputs Text
IBM Granite 3.3 — Training Text+code
IBM Granite 3.3 — Terms Apache 2.0
AllenAI OLMo 2 — Country U.S.
AllenAI OLMo 2 — Inputs Text
AllenAI OLMo 2 — Training Text
AllenAI OLMo 2 — Terms Apache 2.0

Sources

Source 1: Meta

Publisher: Meta

Dataset/table: Llama 4 license/model card

Units: licence and modality disclosure

Date checked: 2026-05-24

Evidence: Supports Llama 4 as a Meta open-weight family under the Llama 4 Community License.

Source 2: Google DeepMind / Google

Publisher: Google DeepMind / Google

Dataset/table: Gemma 3 model card

Units: licence and modality disclosure

Date checked: 2026-05-24

Evidence: Supports Gemma 3 as a Google open-weight family governed by Gemma Terms of Use.

Source 3: Microsoft

Publisher: Microsoft

Dataset/table: Phi-4-multimodal-instruct model card

Units: licence and modality disclosure

Date checked: 2026-05-24

Evidence: Supports Phi-4 multimodal as a Microsoft open-weight model family under the MIT license.

Source 4: IBM

Publisher: IBM

Dataset/table: Granite-3.3-8B-Instruct model card

Units: licence and model-purpose disclosure

Date checked: 2026-05-24

Evidence: Supports IBM Granite 3.3 as an Apache 2.0 language-model release.

Source 5: Allen Institute for AI

Publisher: Allen Institute for AI

Dataset/table: OLMo 2 model card

Units: licence and model-purpose disclosure

Date checked: 2026-05-24

Evidence: Supports OLMo 2 as an Apache 2.0 open model family.

Source 6: Mistral AI

Publisher: Mistral AI

Dataset/table: Mistral Small 3.1 model card

Units: licence and modality disclosure

Date checked: 2026-05-24

Evidence: Supports Mistral Small 3.1 as an Apache 2.0 multimodal model family.

Source 7: Qwen / Alibaba Cloud

Publisher: Qwen / Alibaba Cloud

Dataset/table: Qwen3 model card

Units: licence and model-purpose disclosure

Date checked: 2026-05-24

Evidence: Supports Qwen3 as an Alibaba/Qwen open-weight family under Apache 2.0.

Source 8: DeepSeek AI

Publisher: DeepSeek AI

Dataset/table: DeepSeek-R1 model card

Units: licence and model-purpose disclosure

Date checked: 2026-05-24

Evidence: Supports DeepSeek-R1 as a MIT-licensed open-weight reasoning model family.

Source 9: Technology Innovation Institute

Publisher: Technology Innovation Institute

Dataset/table: Falcon3 model card

Units: licence, training-data and model-purpose disclosure

Date checked: 2026-05-24

Evidence: Supports Falcon 3 as a TII open-weight family under the TII Falcon-LLM License.

Source 10: Cohere Labs

Publisher: Cohere Labs

Dataset/table: Aya Expanse 8B model card

Units: licence and modality disclosure

Date checked: 2026-05-24

Evidence: Supports Aya Expanse 8B as a Cohere Labs open-weight research release under a noncommercial licence.

Data notes & caveats

The matrix records modality categories that the source cards disclose. It does not compare training-token totals or data volumes because those values are not consistently available across rows.

Methodology

Rows were selected from the same 10-family major open-weight scope. Modality wording is normalized into reader-facing categories while preserving the official-card evidence trail in source cards.

Source consistency notes

Rows use mixed 2024-2025 releases and official model-card evidence. The stat is source-labelled because disclosure categories are harmonized across uneven model-card fields.

Licensing

Branded MarketStats graphics are free to use*

Keep the visible MarketStats attribution, source line and canonical page link intact. For clean unbranded assets, commercial use or higher-volume reuse, choose one of the licensing tiers below.

View tiers and terms
Request unbranded access

This specific graphic needs a source-rights review before an unbranded download. Branded use remains available.

* Free branded use is subject to the Terms and requires attribution/canonical-link preservation. MarketStats licenses the graphic presentation and MarketStats-created editorial layer, not the underlying third-party source data. Source attribution and source-provider obligations still apply.