Thứ Hai, 24 thg 8, 2026

Trong đánh giá mới nhất từ Artificial Analysis phối hợp cùng Liquid AI, hai mô hình Nanbeige4.2-3B và LFM2.5-2.6B cùng đạt điểm thông minh trung bình cao nhất là 63 trên iPhone 17 Pro và Galaxy S26 Ultra. Để mô phỏng giới hạn bộ nhớ của thiết bị cầm tay, nhóm nghiên cứu đã giới hạn context window ở mức 16K token. Kết quả cho thấy LFM2.5-2.6B hoạt động hiệu quả hơn khi xử lý prompt 1.024 token chỉ mất 8 giây với 2,3 GB RAM, trong khi Nanbeige4.2-3B cần tới 21,4 giây và 4,0 GB.

Hợp tác này cũng giới thiệu Pipette, một bộ công cụ benchmarking mã nguồn mở đi kèm bộ dữ liệu công khai gồm 10.000 kết quả trên 35 loại mô hình và 4 dòng thiết bị khác nhau. Hiệu suất trên iPhone 17 Pro có sự chênh lệch rất lớn, với thời gian tạo nội dung end-to-end và mức chiếm dụng bộ nhớ tối đa dao động lần lượt từ 30 lần và 19 lần giữa các mô hình được thử nghiệm.

Đăng nhập để góp ý, chỉnh sửa
Nguồn chính
SOURCE@artificialanlys16:14 24 thg 8LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory
SOURCE@liquidai15:12 24 thg 8A public dataset containing more than 10k results, including 1000+ tested performance configurations across models, quantization levels, runtimes, and devices
SUPPORT@artificialanlys16:14 24 thg 8When the context limit is raised to 64K (represented by dots in the image), Ling 3.0 Tiny takes the top spot at 66
SUPPORT@artificialanlys16:14 24 thg 8Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set and runs out of its 16K window on 29% of generations
SUPPORT@artificialanlys16:14 24 thg 8End-to-End Generation Time, measured as the total time to generate 256 output tokens after a 1,024 token input, spans 30x across the models we tested on an iPhone 17 Pro
SUPPORT@artificialanlys16:14 24 thg 8Peak memory at 4K context spans roughly 19x on the iPhone 17 Pro
SUPPORT@artificialanlys16:14 24 thg 8Falcon-H1R-7B takes MATH-500 at 97%