← Model indexMODEL PROFILE / llama-4-scout

Meta / open

Llama 4 Scout

A natively multimodal open-weight MoE model designed for single-H100 efficiency and exceptionally long context.

Context10M
ModalityText · Image
AccessWeights · Partners
ReasoningGeneral

Where it fits.

Distinctive for controlled deployments that combine very long inputs with image-text understanding. Scout is a natively multimodal MoE with 109B total, 17B active parameters and 16 experts. Meta highlights a 10M context and single-H100 deployment with Int4 quantization.

02

Core capabilities

  • Native text-image multimodality with early fusion
  • 10M maximum context for large corpora and codebases
  • 109B total / 17B active MoE architecture
  • Open weights available through Meta and partners
03

Best-fit workflows

  • Ultra-long-context analysis
  • Private multimodal deployment
  • Customization and distillation
04 / OPERATIONS

From evaluation
to production.

Deployment notes

Use supported partner stacks or a tested self-host runtime. Int4 can fit on one H100 according to Meta, but production memory also includes KV cache and concurrency; benchmark target sequence lengths, not only empty-model fit.

Watch before choosing

Practical throughput, memory use and retrieval quality at 10M context depend heavily on serving and quantization. The published 10M capacity is not a guarantee of uniform recall at every position. Long-context systems still need retrieval tests, positional probes, memory planning and careful image-count validation.

05

Family & versions

Specifications and positioning checked against provider documentation.

Open official source