Dev Tools · 1h ago
Open-source engine runs Gemma 4 26B on Macs with 2GB RAM
TurboFieldfare, a Swift/Metal inference engine, enables 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using ~2GB RAM by streaming experts from SSD. It achieves 5-6 tok/s on 8GB M2 MacBook Air and 31-35 tok/s on M5 MacBook Pro. The open-source project includes an OpenAI-compatible local server with streaming and tool calls.
Meridian48 take
This is a clever optimization for running large models on memory-constrained devices, but real-world usability depends on SSD speed and the model's performance at 4-bit quantization.
Read the full reporting
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac →
Hacker News
open-source-aimac-inference