Qwen 3.8 Flash Next is now supported in DwarfStar with fast inference speeds of 50-70 tokens/second and over 1400 tokens/second for prefill on 64GB Mac systems using Metal optimization.
Antirez added support for a quantized DeepSeek V4.1 Flash variant to DwarfStar, claiming it runs with SSD streaming on an M5-Max MacBook Pro with 128GB RAM. The commit was merged recently, so stability is unverified.