← Back to live feed1 story
The local inference engine Lily supports the Qwen3.6-35B-A3B model to enable hybrid compute on Mac devices. Benchmarks conducted on an M5 Max MacBook Pro show that Lily delivers 1.35x higher decode throughput and 1.23x higher prefill throughput than MLX-LM across 10 prompt lengths and 10 decode contexts.
Perplexity released the source code for the engine to ensure on-device processing does not bottleneck tasks within its Computer agent system. The software maintains current output quality levels while optimizing for Apple silicon to reduce the latency of local model execution.
Sign in to suggest edits