Qualcomm unveiled its next flagship mobile chipsets this week, the Snapdragon 8 Elite Gen 6 and a higher-end Elite Extreme Gen 6 variant, both built on a 2-nanometer process and both designed around a mobile processor's most demanding new job: running increasingly large AI models directly on the device rather than routing every request to the cloud.

Both chips use Qualcomm's custom Oryon CPU architecture with two prime cores clocked at 5GHz, a first for a Qualcomm flagship, paired with six performance cores running at 4GHz. Qualcomm says the standard Elite Gen 6 delivers a 10% CPU performance gain and 37% better power efficiency compared with the prior generation, while GPU performance rises 35% with a 40% efficiency improvement. The higher-end Elite Extreme Gen 6 pushes further, with a 13% CPU gain and a 44% year-over-year GPU improvement, and adds support for the VVC (H.266) video codec for more efficient video encoding and playback.

The real story is the NPU

The more significant changes are in the chip's neural processing unit, the dedicated hardware block responsible for running AI models locally rather than in the cloud. Qualcomm's Hexagon NPU in the new generation includes what the company calls an Element Accelerator, purpose-built for transformer workloads, the architecture underlying essentially every modern large language model.

Qualcomm says the standard Elite Gen 6's NPU delivers a 14% performance gain and 20% better AI performance-per-watt over the previous generation. The Elite Extreme Gen 6 goes further still, with a 35% NPU performance improvement and 33% better AI efficiency, along with 50% more shared memory than the standard variant. That expanded memory pool is what allows the Extreme version to support mixture-of-experts AI models with more than 30 billion parameters running entirely on the device, a scale of model that would have required a cloud data center to run just a few years ago.

Qualcomm also introduced Adreno Neural Fusion, a feature that applies AI-driven rendering techniques to graphics workloads and can cut power consumption by up to 40% in supported games, an approach similar in spirit to AI upscaling technologies that have become standard in PC gaming graphics cards.

Why on-device AI capacity matters now

The push toward running larger models locally reflects a shift phone makers and chipmakers have been signaling for roughly two years: as AI features move from novelty chatbots to background agents that handle scheduling, photo editing, real-time translation and proactive suggestions, sending every request to a cloud server introduces latency, cost and privacy tradeoffs that on-device processing avoids. A phone that can run a capable model locally can respond faster, keep working without a network connection, and keep more of a user's personal data, calendar entries, messages, photos, from ever leaving the device.

The tradeoff has always been capability: cloud-hosted models can be far larger and more capable than anything that fits in a phone's power and thermal budget. The jump to supporting 30-billion-parameter mixture-of-experts models on a flagship phone chip narrows that gap meaningfully, though it remains well short of the largest frontier models that continue to run exclusively in data centers.

First devices

Motorola has confirmed its upcoming Signature 27 flagship will use the Snapdragon 8 Elite Extreme Gen 6, and additional device announcements from other phone makers are expected in the weeks following Qualcomm's unveiling, following the industry's typical pattern of staggered flagship launches through the final months of the year.