← Defici Newsai-news

Meta Releases Llama 4 Scout 17B for On-Device Inference, Targeting Mobile and Edge Deployments

By Defici Editorial · 26 Jul 2026

Llama Goes to the Edge

Meta released Llama 4 Scout under an open-weight licence this week, a 17B parameter model optimised for on-device inference on hardware ranging from Apple M-series chips to NVIDIA RTX 4080 consumer GPUs. The model fits within 10 GB of VRAM in 4-bit quantised form, a key threshold for current mobile AI accelerator chips.

Scout is positioned as a complement to Llama 4's flagship 70B and 405B variants, not a replacement. Its design philosophy prioritises instruction-following precision and function-calling reliability over raw reasoning depth, making it suitable for agentic tool use, document summarisation, and real-time translation at the edge.

Benchmark Position

On MT-Bench and IFEval, Scout scores comparably to GPT-4o-mini while achieving 40% lower latency on an Apple M4 Pro chip running MLX. Meta's internal evaluation shows Scout outperforming Mistral Nemo 12B and Phi-4 14B across code generation and structured data extraction tasks, the two use cases most requested in developer community surveys.

Licensing and Deployment

The model is released under the Llama 4 Community License, which permits commercial use with attribution for deployments under 700 million monthly active users — a threshold only the largest platforms approach. Weights are available via Hugging Face, Meta AI's website, and AWS Bedrock's model catalogue.

Hardware partners including Qualcomm, MediaTek, and Samsung have already announced optimised NPU runtimes targeting sub-200ms first-token latency on next-generation smartphone chipsets expected in H2 2026.

ShareXWhatsAppLinkedIn

Get Defici News in your inbox