For the first time in Apple Silicon history, the A20 Pro sports a dual-16-core Neural Engine that’s designed to tackle AI workloads like no other SoC before it, and in short, it’s an upgrade that we thought we didn’t need, but it’s absolutely paramount if you wish to run on-device AI models on a smartphone. In the latest demonstration, an iPhone 18 Pro is shown to run a 27B parameter model at double the speed of the iPhone 17 Pro.
The dual-16-core Neural Engine is certainly faster in token generation, but the consequences of limited memory configurations mean that 2-bit quantized AI models are too big to run on an iPhone 18 Pro
Seeing as how both the iPhone 18 Pro and iPhone 18 Pro Max ship with 12GB 96-bit LPDDR5X RAM that’s significantly faster than the configuration in the iPhone 17 Pro and iPhone 17 Pro Max, this upgrade and the inclusion of the dual-16-core Neural Engine push on-device AI performance to the next level. Adrien Grondin demonstrates these gains by running Bonsai 27B on an iPhone 18 Pro, and you can clearly spot how incredibly fast the token generation speed is.
Of course, while it’s a major step for iPhones when it comes to running denser 27B AI models without an internet connection, there are some trade-offs that users will experience. Firstly, in the X post, it’s mentioned that with Bonsai 2 released, what’s the need to keep running Bonsai? The answer is disappointing, but that’s the harsh reality of running AI models on smartphones; Bonsai 2 with 2-bit quantization is far too big to fit locally on an iPhone 18 Pro, leading to performance degradation.
With Bonsai being a 1-bit quantized AI model, it can effortlessly run on devices with 4GB RAM, and on handsets packing 8GB or even 12GB of memory, it’ll be off to the races. On the iPhone 18 Pro, which isn’t just equipped with an A20 Pro but exceptionally faster memory, a unified memory bandwidth of 115.2GB/s, and a dual-16-core Neural Engine whose peak throughput is faster than the SoC’s 7-core GPU in workloads designed for the NPU, Bonsai 27B will show no signs of slowing down.
Hopefully, when Apple transitions to bigger memory configurations in future releases, we’ll see denser models being supported.
News Source: Adrien Grondin
Follow Wccftech on Google to get more of our news coverage in your feeds.