Particle.news
Download on the App Store

PrismML Says Its 27B ‘Bonsai’ Model Can Run on iPhone as Apple Evaluates Technology

Validation could let iPhones run much larger AI locally, reshaping privacy, latency and cloud cost trade-offs.

Overview

  • PrismML publicly released Bonsai 27B on Tuesday, July 14, saying it compressed Alibaba’s 27-billion-parameter Qwen model from about 54 GB to under 4 GB so the full model can fit within the memory limits of recent iPhones.
  • PrismML CEO Babak Hassibi told reporters that Apple and other companies are actively evaluating the models on devices but that talks are early and no licensing, acquisition, or deployment has been announced.
  • The company achieves the shrinkage through extreme quantization that stores weights as 1-bit or ternary values, which cuts memory dramatically and raises clocked token throughput but typically costs a few percentage points in factual recall and other benchmarks.
  • Analysts and observers say independent, real-world tests are essential, including long-prompt behavior, battery and power measurements during multitasking, KV-cache and activation limits on phones, and reliability under millions of queries before any wide rollout.
  • If the claims hold, the shift could move some compute from datacenters to devices and change chip and cloud demand, but competitors such as Google and Qualcomm and Apple’s own on-device work mean outcomes could include licensing, acquisition, or in-house alternatives.