GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท llama-4-maverick

Llama 4 Maverick

v400B

Meta

Modelopen-sourceself-hosteddata-sovereigntyon-premises
89
Strong
About This Model

Meta's flagship open-weight model (released April 5, 2025): a mixture-of-experts with 400B total / 17B active parameters (128 experts), natively multimodal (text + image), with a 1M-token context window. Meta's last open-weight release: the company has since pivoted to closed models with Muse Spark (April 2026), and newer open models from DeepSeek, Qwen, and Moonshot have surpassed it on most benchmarks.

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

Excellent performance for an open-source model, approaching frontier proprietary models. Performance and latency depend heavily on deployment infrastructure.

task accuracy code

Standard coding benchmarks

Evidence
HumanEval โ€” 86.7% on HumanEval
highVerified: 2026-07-09
task accuracy reasoning

PhD-level reasoning benchmarks

Evidence
GPQA Diamond โ€” 65.2% on PhD-level questions
highVerified: 2026-07-09
task accuracy general

Comprehensive knowledge testing

Evidence
MMLU-Pro โ€” 76.8% on graduate knowledge
highVerified: 2026-07-09
output consistency

Community evaluation

Evidence
Community Testing โ€” Good consistency reported by community
mediumVerified: 2026-07-09
latency p50

Community deployment reports

Evidence
Self-hosted deployments โ€” Latency varies by infrastructure (typically 2-5s)
lowVerified: 2026-07-09
latency p95

Community reports

Evidence
Community reports โ€” Varies significantly
lowVerified: 2026-07-09
context window

Official specification

Evidence
Meta AI Blog - The Llama 4 herd โ€” 1M-token context window for Maverick (Instruct)
Hugging Face - Welcome Llama 4 Maverick & Scout โ€” 1M context confirmed; hosted providers often expose less
highVerified: 2026-07-09
uptime

Deployment model analysis

Evidence
Model Architecture โ€” Uptime controlled by deployment infrastructure
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Security is customer-controlled with self-hosting. Excellent for data sovereignty but requires in-house security expertise.

prompt injection resistance

Community security testing

Evidence
Community Security Testing โ€” Moderate resistance, requires additional guardrails
mediumVerified: 2026-07-09
jailbreak resistance

Safety evaluation

Evidence
Meta Safety Card โ€” Basic safety training, additional tuning recommended
mediumVerified: 2026-07-09
data leakage prevention

Deployment model analysis

Evidence
Self-hosted deployment โ€” Complete data control with on-premises deployment
highVerified: 2026-07-09
output safety

Safety testing

Evidence
Meta Safety Evaluation โ€” Safety training included, additional guardrails recommended
mediumVerified: 2026-07-09
api security

Architecture analysis

Evidence
Deployment model โ€” Security entirely under customer control
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Exceptional privacy - best-in-class. Self-hosting provides complete data control, enabling any compliance framework.

data residency

Deployment model analysis

Evidence
Self-hosted deployment โ€” Complete control over data location
highVerified: 2026-07-09
training data optout

Deployment architecture

Evidence
Self-hosted model โ€” No data sent to Meta for training
highVerified: 2026-07-09
data retention

Architecture analysis

Evidence
Self-hosted deployment โ€” Complete control over data retention
highVerified: 2026-07-09
pii handling

Deployment model analysis

Evidence
On-premises deployment โ€” Customer implements PII handling
highVerified: 2026-07-09
compliance certifications

Deployment model analysis

Evidence
Customer infrastructure โ€” Compliance depends on customer infrastructure (enables HIPAA, etc.)
highVerified: 2026-07-09
zero data retention

Architecture analysis

Evidence
Self-hosted model โ€” No external data transmission
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Excellent transparency as open-source model. Full access to weights and detailed documentation.

explainability

Capability evaluation

Evidence
Model capabilities โ€” Good reasoning explanations
mediumVerified: 2026-07-09
hallucination rate

Community evaluation

Evidence
Community testing โ€” Moderate hallucination rate, similar to other models
mediumVerified: 2026-07-09
bias fairness

Bias benchmark evaluation

Evidence
Meta Responsible AI โ€” Bias testing and mitigation included
mediumVerified: 2026-07-09
uncertainty quantification

Qualitative assessment

Evidence
Model behavior โ€” Reasonable uncertainty expression
mediumVerified: 2026-07-09
model card quality

Documentation review

Evidence
Llama 4 Model Card โ€” Comprehensive open-source model card
highVerified: 2026-07-09
training data transparency

Research paper review

Evidence
Llama 4 Paper โ€” Detailed training data description in paper
highVerified: 2026-07-09
guardrails

Safety mechanism review

Evidence
Meta Safety Tools โ€” Basic guardrails, customer can add more
mediumVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Strong operational maturity with massive open-source ecosystem. Requires in-house ML ops expertise.

api design quality

Tooling review

Evidence
Deployment tools โ€” Standard inference libraries available
highVerified: 2026-07-09
sdk quality

SDK ecosystem review

Evidence
Hugging Face Transformers โ€” Excellent community SDK support
highVerified: 2026-07-09
versioning policy

Release policy review

Evidence
Meta Release Process โ€” Clear versioning with model checkpoints
highVerified: 2026-07-09
monitoring observability

Tooling analysis

Evidence
Customer implementation โ€” Customer must implement monitoring
mediumVerified: 2026-07-09
support quality

Support channel assessment

Evidence
Community Support โ€” Active community, Meta eng engagement
highVerified: 2026-07-09
ecosystem maturity

Ecosystem analysis

Evidence
Open Source Ecosystem โ€” Massive ecosystem (Hugging Face, vLLM, etc.)
highVerified: 2026-07-09
license terms

License review

Evidence
Llama 4 License โ€” Permissive license for commercial use
highVerified: 2026-07-09
Strengths
  • +Best privacy and data sovereignty - complete on-premises control
  • +Open-source with permissive commercial license
  • +No recurring API costs - one-time infrastructure investment
  • +Customizable and fine-tunable for specific domains
  • +Excellent transparency with full model access
  • +No vendor lock-in or rate limits
  • +Best for highly regulated industries and government
Limitations
  • !Requires significant ML ops expertise and infrastructure
  • !Performance and latency depend on hardware investment
  • !Behind frontier proprietary models and newer open models (DeepSeek V4, Qwen3.5, Kimi K2.6) on benchmarks
  • !No managed service or enterprise support from Meta
  • !Requires customer implementation of safety guardrails
  • !High upfront hardware costs (8x A100/H100 GPUs minimum)
  • !Legacy status: Meta's last open-weight release; Meta pivoted to closed models (Muse Spark, April 2026), so no successor open weights are expected
Metadata
pricing
input: $0 (self-hosted)
output: $0 (self-hosted)
notes: Free open weights, infrastructure costs only (8x H100 GPUs ~$200K+) for self-hosting. Also served by third-party APIs at low per-token rates as of July 2026.
last verified: 2026-07-09
context window: 1000000
languages
0: English
1: Spanish
2: French
3: German
4: Italian
5: Portuguese
6: Japanese
7: Korean
8: Chinese
9: Arabic
10: Hindi
11: 100+ languages
modalities
0: text
1: image (input)
api endpoint: Self-hosted or third-party hosts (Together, etc.)
open source: true
architecture: Mixture-of-experts (128 routed experts + shared expert), natively multimodal via early fusion
parameters: 400B total / 17B active

Use Case Ratings

code generation

Strong coding for open-source model. Excellent for on-premises code assistance with data sovereignty.

customer support

Good for self-hosted customer support requiring data privacy.

content creation

Excellent for content creation with strong multilingual capabilities.

data analysis

Good for data analysis with complete data control for sensitive datasets.

research assistant

Excellent for research requiring data sovereignty. Transparent open-source nature aids reproducibility.

legal compliance

Excellent for legal work requiring on-premises deployment. Complete data control enables any compliance framework.

healthcare

Outstanding for healthcare with on-premises HIPAA compliance. Best data sovereignty of any option.

financial analysis

Strong for financial services requiring data residency and air-gapped deployment.

education

Excellent for education with strong multilingual capabilities and customizability.

creative writing

Good for creative writing with ability to fine-tune for specific styles.