Treatmybrand


a Kainjoo SA Venture
Ch. du Vernay 14a
1196 Gland
+41.21.561.34.96
[email protected]

Support


Monday to Friday
8AM to 8PM
[email protected]
Back

Apple’s Latest AI Architecture Breaks On-Device Memory Barrier with 20B Parameter Model

On-device AI models have traditionally been limited by DRAM capacity, restricting their size and capability compared to server-side models. Apple’s new AFM 3 Core Advanced model, announced at WWDC26, overcomes this by storing its entire 20-billion-parameter weight set in NAND flash memory instead of DRAM. This approach, developed with Google, includes a novel routing system that loads only the necessary model components into DRAM per prompt, significantly reducing memory demands.

Unlike conventional Mixture of Experts models that route per token, Apple’s system routes once per prompt to select experts, balancing performance and memory bandwidth. Active parameters scale dynamically from 1 billion to 4 billion depending on task complexity, optimizing resource use. However, Apple has yet to release detailed performance metrics or clarify when tasks offload from device to cloud, creating some uncertainty for enterprise deployment and compliance.

This architectural shift presents a new option for regulated industries needing powerful local AI agents without cloud dependency, while simpler tasks remain on-device and complex ones can offload to AFM 3 Cloud Pro running on Google Cloud. Enterprises must now consider this private/cloud boundary as a strategic decision, weighing the benefits and cloud reliance. Further technical details and benchmarks are expected in a summer report.

Venturebeat
Venturebeat