TUESDAY, JULY 21, 2026 48° E  /  GLOBAL TECH · SUMMARISED SUBSCRIBE
AI, business, devices, policy — global tech, summarised every 30 minutes.
Dev Tools · 7h ago

Gemma 4 E2B Runs on Single TPU v6e Chip, QAT Variants Fail

By Meridian48 News Desk · Summarised from DEV Community ·

Google's Gemma 4 E2B model serves 213 tokens per second on one TPU v6e chip with 32 GB HBM, scaling to 2,200 output tok/s across concurrent streams. The plain model handles function calling and vision tasks, but all QAT variants fail to load due to unimplemented quantization paths and loader bugs. The issues are filed upstream as tpu-inference #3225.

Meridian48 take
The QAT failures highlight the gap between model releases and production-ready inference, especially for architectures with KV-sharing.
Read the full reporting
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive →
DEV Community
gemma-4tpu-inference
More dev tools briefs
Go deeper on dev tools
AllAIStartupsBusinessDevicesPolicySecurityDev ToolsPakistan