AI · 2h ago
LongCat-Video-Avatar 1.5 cuts inference to 8 steps, adds INT8 support
Meituan's open-source talking-avatar model v1.5 reduces generation to 8 sampling steps via DMD2 distillation, swaps audio encoder to Whisper-Large-v3, and adds INT8 quantization for consumer GPUs. The 13.6B-parameter model claims parity or better results versus HeyGen and Kling Avatar 2.0 in internal evaluations. Setup requires PyTorch 2.6.0, CUDA 12.4, and FlashAttention-2.
Meridian48 take
The performance claims are vendor-reported and lack independent verification, so treat the benchmarks as promising but unproven.
open-source-aitalking-avatar