Apple is evaluating technology from startup PrismML that could allow large AI models to run directly on iPhones by significantly reducing memory usage. According to CNBC, CEO Babak Hassibi confirmed that Apple is testing the startup’s compression methods, which aim to shrink massive models—such as Alibaba’s 27-billion-parameter Qwen—to run on mobile hardware with a footprint as small as 4 GB.
How PrismML Technology Shrinks AI Models for Mobile
The core of the technology involves compressing large language models (LLMs) to function within the constrained memory environments of smartphones. PrismML recently demonstrated this by reducing a 54 GB version of Alibaba’s open-source Qwen model down to less than 4 GB. This level of compression allows the model to operate on an iPhone 15 or newer devices.
By moving the computational heavy lifting from the cloud to the device, this shift could address two primary bottlenecks in consumer AI: latency and user privacy. When AI processing occurs locally, requests do not need to be transmitted to a remote server, potentially making voice assistants like Siri faster and more secure.
Did you know?
PrismML’s compression process aims to reduce memory consumption by up to 15 times, potentially allowing sophisticated AI to run without constant internet connectivity.
The Current Status of Apple’s Evaluation
While the prospect of on-device AI is significant, the collaboration remains in the early stages. Babak Hassibi, CEO of PrismML, stated to CNBC that Apple is currently evaluating the speed, energy efficiency, and overall performance of the models on their hardware. He noted that the conversations are exploratory, and it is not yet clear what the final outcome of these tests will be.

Industry analysts, as reported by CNBC, maintain that even with successful compression, the broader AI ecosystem will likely continue to rely on massive data centers. While on-device processing eases the burden on individual requests, the training of models and the management of large-scale AI services will still require substantial chip and server infrastructure.
Market Impact and Future Hardware Trends
The news of Apple’s interest in PrismML comes as the company continues to refine its strategy. Following the report, Apple shares saw a decline of 0.9 percent on Tuesday. The shift toward local AI processing aligns with broader industry trends where manufacturers are racing to optimize hardware for local execution over cloud-reliant architectures.
Pro Tip:
Keep an eye on future iOS updates and hardware release notes. If Apple integrates this type of compression, you may notice faster response times for local Siri queries and improved functionality when your device is offline.
Frequently Asked Questions
Can current iPhones run large AI models?
Currently, iPhones require significant hardware optimization to handle complex AI. PrismML’s technology aims to bridge this gap by compressing large models so they fit within the memory constraints of devices like the iPhone 15.
Why is running AI on the device instead of the cloud important?
Running AI locally improves privacy, as user data does not need to leave the device. It also reduces latency, as the device does not have to wait for a response from a remote cloud server.
Will this replace the need for data centers?
No. According to analysts cited by CNBC, while on-device AI reduces some demand, data centers remain essential for the intensive computing required to train AI models and manage large-scale operations.
What are your thoughts on the future of on-device AI? Are you prioritizing privacy or raw processing power in your next smartphone upgrade? Join the discussion in the comments below or subscribe to our newsletter for the latest updates on mobile technology trends.
Keep reading