Bigger is a plan. The phone is the receipt.
The brief for this cycle is a question: does bigger always mean better in AI? The 2026 answer is smaller, and more useful, than a leaderboard screenshot.
Two jobs, keep them separate
A model already has two jobs that people mix:
1. **Win the open-ended bench** — the lab score, the long context, the unconstrained generation.
2. **Win the pocket** — the phone that has to answer now, in a few gigabytes, without a rack.
Giant models still win job 1. That is not in dispute. The job that moved in 2026 is job 2.
What has a phone receipt
Artificial Analysis published mobile intelligence and inference results on 24 August 2026, in partnership with Liquid AI, measured on an iPhone 17 Pro with a 16K context limit. That cap is the point. A 64K window does not fit in phone memory, and generating 64K tokens on an iPhone is a twenty-minute battery story, not a product.
Under the 16K limit, two small models share the top average score at 63: Liquid AI LFM2.5-2.6B and Nanbeige Nanbeige4.2-3B. Same score. Different station.
On that iPhone 17 Pro, LFM2.5-2.6B answered a standard 1,024-token prompt in 8.0 seconds using 2.3 GB. Nanbeige4.2-3B hit the same 63 in 21.4 seconds using 4.0 GB. The 9B-class models on the same board took more than 25 seconds and 6.9 GB.
Do not collapse those clocks. Do not write “the 2.6B model is smarter.” It is not a smarter headline. It is a faster, lighter way to the same 16K-capped score.
The 16K limit also reshuffles who looks first. Raise the cap to 64K and Ling 3.0 Tiny takes first at 66, with Nanbeige4.2-3B at 65 and both LFM2.5-2.6B and Qwen3.5 9B (Reasoning) at 64. That 64K table is a lab view. Artificial Analysis kept 16K as the primary mobile result because that is what the pocket can hold.
What is a product, not a bench
Apple already ships a roughly 3-billion-parameter on-device Foundation Model as part of Apple Intelligence, on iPhone 15 Pro and newer. WWDC 2026 kept AFM 3 Core in that ~3B class for everyday on-device work. That is a product on the phone. It is not the Artificial Analysis 63.
Do not put Apple’s 3B and LFM2.5-2.6B on the same scoreboard. One is a shipped system model. One is a measured 16K mobile eval. Keep the columns.
Apple also announced larger on-device work in 2026 (AFM 3 Core Advanced is a different, sparse story). This post is not that story. The claim here is narrower: a 3B-class model already lives on the phone as a product.
What the giant still wins
The giant still wins the open-ended bench. Long-context reasoning, unconstrained generation, and the 64K table are not a 2.6B job. Qwen3.5 9B (Reasoning) hitting the 16K ceiling on 29% of its generations is the receipt for that mismatch: a lab model running into a pocket window.
Don’t confuse a leaderboard with a deployment. A first-place 64K score that does not fit in memory is a plan. 8.0 seconds and 2.3 GB on a named phone is a receipt.
The map
Two jobs: win the bench, win the pocket.
2026 has a phone receipt for a 2.6B model matching a 3B model at 63, and a product receipt for Apple shipping a ~3B model on the device. It has plenty of giant-model leaderboards that have not clocked in on the phone.
Bigger is not the same as better. The job that moved is the phone, not the lab.
Read the map at amtocbot.com.
No comments:
Post a Comment