← Defici Newsai-news

Meta Releases Llama 4 Scout Open Weights as Open-Source AI Race Heats Up

By Defici Editorial · 19 Jul 2026

Meta has made Llama 4 Scout available for download under its community licence, releasing the full model weights for what it describes as its most capable openly-available language model. Scout uses a sparse mixture-of-experts architecture with 17 billion active parameters and 109 billion total parameters, meaning it achieves strong benchmark performance while requiring only fraction of the compute of a dense model the same total size.

The release of open weights is significant for the enterprise and research community. Scout can be fine-tuned and self-hosted on hardware that a single organisation can own and control, removing the dependency on cloud API providers that many security-sensitive organisations have been reluctant to accept. The model runs on 8xH100 with full 128k context active, or on 4xH100 with reduced context using quantisation.

On standard benchmarks, Scout scores competitively with models in the commercial API market: 87.2% on MMLU, 71.3% on HumanEval, and 62.1% on GPQA — numbers that place it clearly ahead of the previous open-source leader Llama 3 70B, and within striking distance of some commercial mid-tier offerings.

Meta argues that open models are essential for AI safety research because they allow the broader research community to study model internals, conduct red-teaming, and develop mechanistic interpretability techniques that are impossible with black-box API access. The company published a detailed model card alongside the weights covering training data composition, safety evaluations, and known failure modes.

The practical deployment picture is not uniformly rosy. Scout's mixture-of-experts architecture requires specialised inference infrastructure to achieve the latency advantages it theoretically offers — naive deployments may actually run slower than dense models of the same active-parameter count because of expert-routing overhead. Meta has published reference serving configurations using its custom inference stack, but integration with standard frameworks like vLLM and TGI still has rough edges that the open-source community is actively smoothing.

Several cloud providers have announced they will make Scout available as a hosted inference endpoint within weeks, which will benefit organisations that want open-model flexibility without the infrastructure burden of self-hosting.

The release also intensifies the competitive dynamic between Meta and the commercial AI labs. With Scout available at zero API cost for self-hosters, the price pressure on commercial mid-tier models is likely to increase, potentially forcing further price cuts from providers already competing on cost.

ShareXWhatsAppLinkedIn

Get Defici News in your inbox