Dev Tools · 1h ago
ExecuTorch MLX Delegate Speeds Qwen3 4.52x on M1 Max
ExecuTorch's experimental MLX delegate, released in May 2026, enables PyTorch models to run on Apple Silicon GPUs. Testing Qwen3-0.6B on an M1 Max showed decode throughput of 188.9 tokens/s with INT4 quantization, 4.52x faster than PyTorch MPS BF16. However, INT4 altered output in two of three prompts, highlighting a trade-off between speed and accuracy.
Meridian48 take
The speed gains are impressive, but the output changes with INT4 quantization mean developers must weigh performance against reliability for their use case.
Read the full reporting
Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max →
DEV Community
executorchmlx-delegate