Llama 4 Scout 109B
Multimodal 109B MoE model with 17B active parameters for powerful workstations and servers.
DeviceTerra has not lab-tested this exact artifact on every hardware combination. RAM and VRAM floors are LocalLens planning estimates, not publisher guarantees.
Will it run?
LocalLens uses 96 GB as a conservative planning floor for this profile and its listed quantization. This is not measured speed. Processor, backend support, context length, memory bandwidth and other applications still affect the result.
Check your complete system →What is it useful for?
Suggested Ollama artifact
Confirm the tag on the official page before downloading, then run:
ollama run llama4:scoutStart with a representative non-sensitive task. Record the exact runtime, quantization, context, load time, generation speed, memory use and answer quality.
Evidence and limits
Published context: Very long context; practical limit is hardware-dependent. The source establishes specifications; it does not guarantee that maximum context fits at the LocalLens minimum RAM floor.