2026-06-23 BREAKTHROUGHS☀ AM
Meta releases 405 billion parameter Llama 3.1 with open weights
📰 THE BRIEF
Meta published the full weights for Llama 3.1 405B. The model matches or exceeds GPT-4 on standard benchmarks and runs locally on consumer GPUs or inexpensive cloud instances. No per-token API fees apply.
💡 WHY IT MATTERS
Open weights shift the cost structure from usage fees to hardware and electricity. Users can now run large models without sending data to third-party servers. This changes deployment decisions from cloud subscription to infrastructure planning.
👥 WHO'S DOING IT
Hugging Face hosts the model weights and reports thousands of daily downloads. Independent developers have deployed quantized versions achieving 30 tokens per second on single RTX 4090 cards.
⚡ TRY IT
- Visit huggingface.co/meta-llama/Meta-Llama-3.1-405B and accept the license.
- Install Ollama from ollama.com and run 'ollama run llama3.1:405b'.
- Query the local endpoint and observe inference speed without API costs.