Nanbeige4.2-3B and LFM2.5-2.6B tied for the highest average intelligence score of 63 on the iPhone 17 Pro and Galaxy S26 Ultra. The evaluation, conducted by Artificial Analysis in partnership with Liquid AI, limited models to a 16K context window to mirror the memory constraints of handheld devices. LFM2.5-2.6B demonstrated higher efficiency by processing a 1,024 token prompt in 8.0 seconds using 2.3 GB of memory, while Nanbeige4.2-3B required 21.4 seconds and 4.0 GB.

The partnership introduces Pipette, an open source benchmarking suite featuring a public dataset of 10,000 results across 35 model classes and four device types. Performance on the iPhone 17 Pro varied widely, with end to end generation times spanning 30x and peak memory usage spanning 19x among tested models. The tool includes macOS, Windows, iOS, and Android clients, allowing developers to measure quality, speed, and latency directly on target hardware.

Sign in to suggest edits

Key sources

  1. SOURCE@artificialanlys“LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory”x.com
  2. SOURCE@liquidai“A public dataset containing more than 10k results, including 1000+ tested performance configurations across models, quantization levels, runtimes, and devices”x.com
  3. SUPPORT@artificialanlys“When the context limit is raised to 64K (represented by dots in the image), Ling 3.0 Tiny takes the top spot at 66”x.com
  4. SUPPORT@artificialanlys“Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set and runs out of its 16K window on 29% of generations”x.com
  5. SUPPORT@artificialanlys“End-to-End Generation Time, measured as the total time to generate 256 output tokens after a 1,024 token input, spans 30x across the models we tested on an iPhone 17 Pro”x.com
  6. SUPPORT@artificialanlys“Peak memory at 4K context spans roughly 19x on the iPhone 17 Pro”x.com
  7. SUPPORT@artificialanlys“Falcon-H1R-7B takes MATH-500 at 97%”x.com
Markdown