Trong đánh giá mới nhất từ Artificial Analysis phối hợp cùng Liquid AI, hai mô hình Nanbeige4.2-3B và LFM2.5-2.6B cùng đạt điểm thông minh trung bình cao nhất là 63 trên iPhone 17 Pro và Galaxy S26 Ultra. Để mô phỏng giới hạn bộ nhớ của thiết bị cầm tay, nhóm nghiên cứu đã giới hạn context window ở mức 16K token. Kết quả cho thấy LFM2.5-2.6B hoạt động hiệu quả hơn khi xử lý prompt 1.024 token chỉ mất 8 giây với 2,3 GB RAM, trong khi Nanbeige4.2-3B cần tới 21,4 giây và 4,0 GB.

Hợp tác này cũng giới thiệu Pipette, một bộ công cụ benchmarking mã nguồn mở đi kèm bộ dữ liệu công khai gồm 10.000 kết quả trên 35 loại mô hình và 4 dòng thiết bị khác nhau. Hiệu suất trên iPhone 17 Pro có sự chênh lệch rất lớn, với thời gian tạo nội dung end-to-end và mức chiếm dụng bộ nhớ tối đa dao động lần lượt từ 30 lần và 19 lần giữa các mô hình được thử nghiệm.

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SOURCE@artificialanlys“LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory”x.com
  2. SOURCE@liquidai“A public dataset containing more than 10k results, including 1000+ tested performance configurations across models, quantization levels, runtimes, and devices”x.com
  3. SUPPORT@artificialanlys“When the context limit is raised to 64K (represented by dots in the image), Ling 3.0 Tiny takes the top spot at 66”x.com
  4. SUPPORT@artificialanlys“Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set and runs out of its 16K window on 29% of generations”x.com
  5. SUPPORT@artificialanlys“End-to-End Generation Time, measured as the total time to generate 256 output tokens after a 1,024 token input, spans 30x across the models we tested on an iPhone 17 Pro”x.com
  6. SUPPORT@artificialanlys“Peak memory at 4K context spans roughly 19x on the iPhone 17 Pro”x.com
  7. SUPPORT@artificialanlys“Falcon-H1R-7B takes MATH-500 at 97%”x.com
Bản Markdown