Local AI models on computers have become increasingly powerful in recent years, and many basic tasks and coding jobs can be handled with local AI, so many people would certainly want to run them on their phones as well. Although phones have been able to run some local AI for a while, limited by memory, performance, and model size, the models that can run smoothly are usually limited in scale, and the actual experience still lags behind large models on computers. The good news is that with the launch of the iPhone 18 Pro series, this situation seems to be starting to change.
Recently, an overseas developer tested running the local 27B—meaning 27-billion-parameter—AI model Bonsai 27B on an iPhone 18 Pro. The results showed that, compared with last year’s iPhone 17 Pro, generation speed was about twice as fast. More importantly, the entire model runs completely locally on the phone, with no need to connect to the cloud. Such a clear improvement after just one generation makes people increasingly look forward to the future development of on-device AI on phones.

iPhone 18 Pro’s on-device AI capabilities revealed! A20 Pro runs a 27B parameter model twice as fast as iPhone 17 Pro.
Recently, a developer from overseas, Adrien Grondin, posted a hands-on test video on X: He used an iPhone 18 Pro to run Prism ML’s Bonsai 27B model, and the token generation speed was about twice that of the iPhone 17 Pro. In the post, he wrote, “The iPhone 18 Pro’s on-device AI performance is much better than I expected. It runs a 27B model twice as fast as the 17 Pro. The new A20 Pro chip is a monster.”
Adrien Grondin is no random nobody either. He is the author of Locally AI, an iOS local AI chat app. The app was acquired by LM Studio in April this year, and he now oversees LM Studio’s native experience on Apple devices. In the past, he often ran various open-source models on iPhones for demonstrations, such as running Gemma 4 at 40 tokens per second on an iPhone 17 Pro, so this hands-on test has some value as a reference.
Bonsai 27B is an open-source model released by startup Prism ML in July this year. It is based on Alibaba’s Qwen3.6-27B, and its biggest feature is the use of “1-bit quantization,” which compresses the original model file of about 54GB down to only about 3.9GB. In theory, a device with 4GB of memory can run it. This is also why it can run on the iPhone 17 Pro and iPhone 18 Pro.
Prism ML previously also published test data for the iPhone 17 Pro Max, at about 11 tokens per second, and claimed it was the first 27B-class model that can run on a phone.
By the way, in this post Grondin only said “twice” and didn’t give an exact tok/s number. If we estimate from the 11 tokens per second that Prism ML announced for the 17 Pro Max, the iPhone 18 Pro would land at just over 20 tokens per second.
This number actually means something now! Because for chat to read like a “real-time conversation,” according to foreign reports the threshold is roughly 15 to 20 tokens per second. The iPhone 17 Pro Max’s 11 tok/s would make people feel like they’re constantly waiting, and if the iPhone 18 Pro really crosses 20, then a 27B model on a phone goes from “able to run” to “usable.” From the video Grondin released, you can also clearly see that the speed is indeed very fast.

In addition to the 1-bit version, Prism ML also has a 2-bit「 that can retain 95% of the original’s capabilitiesTernary Bonsai 27BIt is about 7.2GB in size, but this version is too large for the iPhone 18 Pro; forcing it to run would cause performance to drop, and Prism ML is also positioned for laptop and GPU use. Grondin also commented that the new Bonsai 2 (2-bit version) is too large to run on an iPhone.

Source: KOCPC Chinese