Overview

Accelerate agentic AI on Arm CPUs

AI Summary

Arm SME2 accelerates matrix-heavy AI workloads directly on the CPU, helping deliver responsive, private on-device experiences. With expanded SME2 capability in the latest Arm C2 CPU cluster, developers can accelerate speech, search, generative AI, and other workloads that power emerging agentic experiences.

Benefits

Why SME2 matters?

Up to 1.7x faster agentic AI

Accelerate voice, personal memory, and reasoning workloads that form the building blocks of agentic AI on mobile.*

Responsive on-device AI

Run lightweight AI models on the CPU to support low-latency, on-device experiences across agentic workflows.

Broad framework support

Access SME2 acceleration through Arm KleidiAI integrations across more than 12 AI frameworks and 100 third-party applications.

Proven across mobile AI use cases

Support real-world experiences including live translation, speech recognition, image recognition, background blur, and object detection.

*Up to 1.7x faster execution with CSS for Mobile 2 versus the previous-generation reference platform across tested agentic AI workloads.

Features

Built for modern AI workloads

Matrix acceleration for lightweight AI

Accelerate matrix-heavy operations used by specialized on-device models for voice, memory retrieval, reasoning, and other AI tasks.

Low-precision model support

Support compact, quantized AI models, including 2-bit workloads, to reduce memory demands and improve execution efficiency.

Concurrent AI execution

Run transcription, translation, voice generation, and other AI workloads concurrently on the CPU for responsive multi-stage experiences.

Scalable by design

Delivers flexible performance from entry-tier to flagship mobile devices, ensuring consistent developer and user experiences across devices.

See SME2 documentation
Success story
SME2: Advancing on-device AI on Android

See how Arm and Google use heterogeneous computing to advance on-device AI on Android. With SME2 accelerating AI directly on the CPU, developers can deliver faster, more responsive, and scalable AI experiences across the Android ecosystem.

 
"LiteRT and Google AI Edge are built to take advantage of the best available acceleration on-device, including selecting Arm SME2 by default on supported platforms. In Google AI Edge Gallery, SME2 helps accelerate demanding generative AI workloads directly on the CPU, enabling responsive experiences such as on-device translation with models like Gemma. Our collaboration with Arm is helping developers unlock more of this performance through optimized software and hardware working together.”
Na Li, On-device ML Developer Tools Lead, Google
Use cases

SME2 in action

SME2 powers intelligent, low-latency, private on-device AI workloads, enabling use cases such as agentic calling, personalized workout coaching, immersive NPC interactions, and neural imaging.

 
 
 
 
Developer

Built for developers

Thanks to native SME2 support across leading AI frameworks and runtime libraries—including PyTorch, ONNX Runtime, XNNPACK, and llama.cpp—developers can access SME2 benefits without changing a single line of code. SME2-enhanced performance is also portable across Arm-based platforms, from iOS and iPadOS to macOS and, soon, Android.

 

Explore the new Arm Developer Launchpad for SME2 to understand SME2 acceleration and use cases, supported hardware, step-by-step tutorials, and hands-on learning paths.

Start building with SME2Read developer blog
RELATED PRODUCTS

Explore the Arm CSS for Mobile 2 platform

Arm CSS for Mobile 2

Arm CSS for Mobile 2

Mali G2-Ultra NX is part of CSS for Mobile 2, an integrated compute platform that combines AI orchestration, neural graphics, System IP, and software to enable next-generation mobile AI experiences.

Arm C2 CPU Cluster

Arm C2 CPU cluster

The high-performance Arm CPU cluster with SME2 for AI orchestration and responsive application performance.

Stay connected

Subscribe to stay up to date on the latest news, case studies, and insights.

Newsletter signup

Frequently asked questions: SME2

What is Arm SME2?

SME2 (Scalable Matrix Extension 2) is an advanced set of CPU instructions in the Armv9.3-A architecture designed to accelerate AI and ML workloads, particularly matrix-heavy tasks like LLMs and computer vision. It integrates seamlessly with popular AI frameworks via Arm KleidiAI, delivering higher performance and efficiency without code changes.

How does SME2 improve AI performance on devices?

By executing matrix operations directly on the CPU, SME2 enables up to 6X faster inference for large language models and 3X improvements in vision and audio processing—without requiring separate NPUs or cloud resources.

Which devices will support SME2?

Available now on iPhone 17 (A19), Apple M series devices, and flagship Android phones.

How does SME2 benefit developers?

SME2 integrates automatically with frameworks like Pytorch, ONNX Runtime, and XNNPACK, so developers can accelerate AI workloads without rewriting code. Developers can explore Arm AI on mobile resources for toolchains, SDKs, and training to get started quickly.

Can SME2 help with generative AI applications?

Absolutely. SME2 accelerates generative AI tasks, such as real-time translation, photo/video enhancement, audio generation, and motion analysis, directly on-device. This enables faster, more private, and more energy-efficient user experiences. Developers can learn how to implement these capabilities with Arm AI on mobile resources.