Skip to main content

Blog

Why Arm CPU? Your guide to Arm-based server infrastructure for AI, cloud and enterprise workloads

For a long time, Arm in the datacentre was treated as a hyperscale experiment. The assumption was simple: x86 owned enterprise servers, while Arm belonged either in mobile devices or in custom cloud platforms built by companies with the engineering depth to make their own silicon choices.

That assumption is now out of date.

Arm says around 50% of the compute shipped to top cloud service providers is now Arm-based, and more than 1.25 billion Neoverse cores have shipped into datacentres. The largest cloud providers have already put Arm-based processors into production for first-party workloads including ads, search, databases, e-commerce and application platforms. AWS Graviton, Google Axion, Microsoft Cobalt and Alibaba Yitian all point in the same direction: Arm is no longer a side bet. It is part of the mainstream cloud compute base.

The buyer question has changed with it. The point is not whether Arm servers can run serious workloads. They already do. It’s more useful to look at where Arm changes the economics of AI-era infrastructure, and how enterprise, cloud and neocloud buyers should evaluate it against the x86 platforms they know.

ARM-CPU

 

The infrastructure question has changed

Server buying used to be framed around socket performance, processor familiarity and software compatibility. Those still matter, but they are no longer enough. The binding constraints in modern datacentres are power, cooling, rack space, memory bandwidth, I/O and the ability to keep increasingly expensive accelerator infrastructure fed with useful work.

That is why Arm has become more relevant. Its core architectural advantage is efficiency: more useful compute per watt, more predictable performance under dense deployment, and a design model that can be tuned for the workload rather than locked into one generic server profile. At rack level, where power and cooling limits are hard constraints, that advantage matters more than a single headline benchmark.

The shift is especially important for AI infrastructure. GPUs and specialist accelerators attract most of the attention, but they do not run alone. Every AI platform needs CPU capacity for orchestration, control-plane work, data preparation, retrieval, search, tokenisation, networking, storage, security, logging, web serving, databases and application logic. As AI systems become more agentic, they will only generate more of this surrounding work.

In other words: AI does not make the CPU irrelevant. Instead, it raises the standard for what the CPU layer has to deliver.

 

What the hyperscalers have already proven

The clearest proof point for Arm is hyperscale adoption.

Cloud providers are ruthless about infrastructure economics. They buy at huge scale, optimise for power and rack density, and measure performance in production rather than in abstract. If Arm could not deliver consistent value across real cloud workloads, it would not have become such a large part of the compute base for top cloud service providers.

The workloads now associated with Arm-based cloud infrastructure are broad: ads, search, big data services, cloud databases, e-commerce platforms and first-party application infrastructure. These workloads are high-volume, latency-sensitive and commercially critical.

That matters for enterprise buyers because hyperscaler adoption reduces risk. It proves that Arm has the software ecosystem, operating system support, developer tooling and operational maturity required for serious infrastructure. It also sets expectations for the rest of the market. Once hyperscalers demonstrate a better cost-performance envelope, enterprise buyers and neocloud operators begin asking whether they can capture the same advantage in their own environments.

What Arm AGI CPU brings to the market

The next phase is about taking Arm beyond bespoke hyperscaler silicon and making it a more standard server infrastructure option.

Arm’s AGI CPU platform is designed for cloud and AI workloads on the Neoverse V-Series. The headline configuration brings up to 136 Neoverse V3 cores, 2MB of dedicated L2 cache per core, Armv9.2 ISA support and SVE2 with machine-learning acceleration instructions for bf16 and int8 workloads. The product family includes 136-core, 128-core and 64-core variants, with configurable TDP ranges that allow buyers to tune core count, frequency and power for different deployment profiles. See table below.

Memory and I/O are just as central to the design as core count. All three variants support 12 channels of DDR5 (up to 8800 MT/s), 96 lanes of PCIe Gen 6, and CXL 3.0 Type 3 for memory expansion – addressing the reality that AI-era workloads are often bottlenecked by data movement and memory access, not just CPU arithmetic.

Specs Arm AGI CPU 136C (max core count)Arm AGI CPU 128C (TCO optimised)Arm AGI CPU 64C (max mem/core)
Processing cores136 Neoverse V3

2x 128 SVE

2MB/core L2

128 Neoverse V3

2x 128 SVE

2MB/core L2

64 Neoverse V3

2x 128 SVE

2MB/core L2

CPU architectureArmv9.2

bfloat16 and INT8 AI instructions

Armv9.2

bfloat16 and INT8 AI instructions

Armv9.2

bfloat16 and INT8 AI instructions

System-level cache128MB128MB128MB
Max Frequency3.5GHz3.5GHz3.7GHz
Base TDP*

 

*Represents a preset TDP value within the configurable TDP range

 

300W300W300W
RDIMM memory12x DDR5

Up to 8800 MT/s

12x DDR5

Up to 8800 MT/s

12x DDR5

Up to 8800 MT/s

PCIe/IO96x lanes PCIe Gen6

CXL 3.0 Type 3

96x lanes PCIe Gen6

CXL 3.0 Type 3

96x lanes PCIe Gen6

CXL 3.0 Type 3

PCIe control lanes6x 1 Gen46x 1 Gen46x 1 Gen4
2-Socket supportYesYesYes

 

(Reference: Arm newsroom)

Why AI makes the CPU more important

The popular story of AI infrastructure is that GPUs won and CPUs became background plumbing. That is too simplistic.

Accelerators handle the most compute-intensive parts of training and inference, but the surrounding system has to prepare, route and feed the work. Agentic AI makes the surrounding layer busier because workloads become more fragmented. A single user request may trigger retrieval, search, tool calls, database reads, policy checks, model calls, post-processing and application actions. Much of that work sits on general-purpose compute.

This is why Arm’s AGI CPU positioning talks about both cloud workloads and AI CPU head nodes. The CPU layer has to provide high single-thread performance, predictable memory access, strong I/O bandwidth to accelerators, and enough efficiency to operate economically at scale. It also has to support familiar enterprise workloads: databases, video, Java, analytics, web serving and AI inference support tasks.

The clearest use cases are not necessarily the largest GPU jobs. They are the workloads that surround and enable AI systems. Think vector databases, search, KV-cache handling, tokenisation, data pre- and post-processing, media processing, application serving, control planes, logging, security services and analytical databases. These are high-volume workloads where performance-per-watt, memory bandwidth and rack density can materially change cost.

The economics are rack-level

A common mistake in processor comparisons is to stop at socket-level benchmarks. That is not how infrastructure budgets are constrained.

Datacentres run into power envelopes, cooling limits, rack density limits and floor-space constraints. Once those limits are fixed, the question becomes how much useful throughput a buyer can extract from each rack over the life of the deployment. That is where Arm’s efficiency case is strongest.

In Arm’s own rack-level comparison, an Arm AGI CPU rack delivers more than 2x the performance per rack versus a comparable x86 setup, at the same 36kW power envelope: 30x 1U Arm-based servers and 8,160 CPU cores, versus 17x 2U x86 servers and 4,352 CPU cores. Arm’s own materials note this is a projected comparison, but the message is clear: at a fixed power budget, rack design and core density can matter more than any single chip-to-chip benchmark

Software compatibility is no longer the blocker it once was

The historic objection to Arm servers – software compatibility – has weakened significantly.

The hyperscaler story proves much of the point: cloud-native applications, Linux operating systems, developer tools and commercial software stacks now run across Arm environments at scale. Arm’s own positioning points to a mature ecosystem across Linux and operating systems, cloud-native platforms, SaaS and enterprise AI/ML software.

That does not mean migration is automatic. Buyers still need to check operating system support, ISV certification, container images, runtime dependencies, observability agents, security tooling and performance characteristics under their own workload. But the question has shifted from “is Arm supported?” to “which of our workloads are already easy to move, and which need more planning?”

That is a healthier procurement question. It lets buyers sequence migration sensibly: start with high-volume cloud-native workloads, containerised services, stateless application tiers, analytics workloads or AI support services where compatibility is strongest and the efficiency upside is easiest to measure. Leave harder bare-metal or deeply integrated legacy workloads until the business case is clearer.

Where this leaves buyers

None of this makes x86 obsolete, and nobody serious is arguing it should be ripped out overnight. What’s changed is the burden of proof: Arm has already done the hard part at hyperscale, the software runs, the ecosystem is mature, the efficiency case holds up in production. The question now isn’t “does Arm work?” It’s “where does it work best for us, and what does that do to our power, rack and cost envelope?”

Start with the workloads where the case is easiest to prove, cloud-native, containerised, AI-support, and model it against your own environment rather than a benchmark slide.

If you’re weighing up whether Arm-based infrastructure makes sense for your environment, get in touch and we’ll help you work through the case properly.

Scroll back up to page top
Follow us
Contact us