WEDNESDAY, JULY 22, 2026 48° E  /  GLOBAL TECH · SUMMARISED SUBSCRIBE
AI, business, devices, policy — global tech, summarised every 30 minutes.
Dev Tools · 1h ago

ExecuTorch MLX Delegate Speeds Qwen3 4.52x on M1 Max

By Meridian48 News Desk · Summarised from DEV Community ·

ExecuTorch's experimental MLX delegate, released in May 2026, enables PyTorch models to run on Apple Silicon GPUs. Testing Qwen3-0.6B on an M1 Max showed decode throughput of 188.9 tokens/s with INT4 quantization, 4.52x faster than PyTorch MPS BF16. However, INT4 altered output in two of three prompts, highlighting a trade-off between speed and accuracy.

Meridian48 take
The speed gains are impressive, but the output changes with INT4 quantization mean developers must weigh performance against reliability for their use case.
Read the full reporting
Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max →
DEV Community
executorchmlx-delegate
More dev tools briefs
Go deeper on dev tools
AllAIStartupsBusinessDevicesPolicySecurityDev ToolsPakistan