Khosla-backed startup runs largest AI model ever on iPhone
PrismML, a Khosla-backed AI startup, says it compressed Alibaba's 27-billion-parameter Qwen 3.6 model from roughly 54GB to under 4GB and ran it fully on an iPhone 17 Pro — the largest AI model ever on Apple's phone, per The Information. That would let Vinod Khosla-backed teams and Apple run far more intelligence on-device, without the cloud.
The report, covered by MacRumors and AppleInsider, lands as Apple races to expand on-device Apple Intelligence. If the claims hold, iPhone owners could eventually get server-grade reasoning and agent-style tasks without sending data to remote servers.
Key Takeaways
- PrismML says it ran all 27 billion parameters of Qwen 3.6 simultaneously on an iPhone 17 Pro after shrinking the model from roughly 54GB to about 4GB.
- Apple has held meetings with PrismML about using the compression tech to run much larger models directly on iPhones, per The Information.
- PrismML's dense approach contrasts with Apple's sparse AFM 3 Core Advanced model, which activates only 1 billion to 4 billion of its 20 billion parameters at once.
- More on-device power could cut Apple's reliance on Private Cloud Compute while improving privacy and lowering infrastructure costs.
- No partnership is confirmed — AppleInsider notes there is no guarantee the companies will work together or when users would see benefits.
What Did the Khosla-Backed Startup Actually Demonstrate?
According to The Information, PrismML used mathematical compression to shrink Alibaba's open-source Qwen 3.6 large language model. The result, AppleInsider reports, is a system that normally demands server-scale storage running inside a smartphone footprint.
PrismML says the on-phone model can handle complex conversations, reasoning, and fully autonomous agents — tasks that typically require cloud-backed models. The startup also claims its technique does not sacrifice performance, a common trade-off when models are compressed.
Unlike Apple's current on-device stack, every parameter in PrismML's version stays active at the same time. That full activation is central to why the company calls the demo the largest AI model ever run entirely on an iPhone.
Why Is Apple Interested in PrismML's Technology?
Apple has long favored processing sensitive AI workloads on the device rather than in the cloud. MacRumors notes that running bigger models locally would let more Apple Intelligence features skip Apple's Private Cloud Compute servers.
That shift could reduce Apple's operating costs and strengthen privacy promises — both strategic priorities as rivals pour billions into data-center AI. The Information reports Apple has already met with PrismML to discuss how the startup's methods might fit future iPhones.
AppleInsider emphasizes that interest does not mean a deal is imminent. Representatives may simply be scouting outside expertise after in-house compression efforts faced performance hurdles.
How Does PrismML Compare to Apple's Current On-Device AI?
Apple's AFM 3 Core Advanced model, which powers iOS 27 enhancements like more expressive Siri voices and improved systemwide dictation on iPhone 17 Pro and iPhone Air, contains 20 billion parameters, MacRumors reports.
However, Apple uses a sparse architecture. Only about 1 billion to 4 billion parameters are active during any given task. PrismML's compressed Qwen 3.6 keeps all 27 billion parameters live at once — a denser design that the startup argues enables heavier workloads on the same hardware.
On paper, that makes PrismML's demonstration larger in total active capacity than Apple's current on-device flagship, even though both run on the latest Pro-tier iPhone silicon.
When Could iPhone Users See Smarter On-Device AI?
For now, the news is about meetings and claims, not product launches. AppleInsider cautions that even if Apple and PrismML partner, it is unknown when Siri or Apple Intelligence customers would gain access to the compressed models.
Still, the timing matters for anyone tracking how big tech buys innovation. A Vinod Khosla-backed firm showing server-scale AI on a phone puts fresh pressure on Apple to close the gap between cloud assistants and local intelligence — a storyline worth watching in our Celebrity Breaking News coverage as Silicon Valley power players place new bets.