Dev Tools · 7h ago
Gemma 4 E2B Runs on Single TPU v6e Chip, QAT Variants Fail
Google's Gemma 4 E2B model serves 213 tokens per second on one TPU v6e chip with 32 GB HBM, scaling to 2,200 output tok/s across concurrent streams. The plain model handles function calling and vision tasks, but all QAT variants fail to load due to unimplemented quantization paths and loader bugs. The issues are filed upstream as tpu-inference #3225.
Meridian48 take
The QAT failures highlight the gap between model releases and production-ready inference, especially for architectures with KV-sharing.
gemma-4tpu-inference