AI is quietly moving out of the data center and into your pocket, and Arm Holdings just placed a large bet on that migration. On September 7 and 8, the chip-IP giant announced a new collaboration with Samsung plus a refresh of its mobile compute platform, both built around one thesis: the next generation of AI runs on the device, not in the cloud.
What Actually Got Announced
Arm and Samsung are developing a 2nm on-device AI accelerator SoC. Arm is supplying the AI accelerator architecture and core design IP, while Samsung's System LSI division handles full SoC integration and its foundry manufactures the chip on the SF2 2nm process, which uses Gate-All-Around (GAA) transistors. The stated goal is to run richer AI locally and reduce dependence on remote inference. Details on Arm's side are on the Arm newsroom.
That partnership arrived the same week Arm unveiled CSS for Mobile 2, its next mobile compute platform. The headline CPU element is a new C2 cluster that pairs the C2-Ultra performance core and C2-Pro efficiency cores with two SME2 vector-multiply units. Arm says doubling the SME2 capability delivers a 70% speedup on the latest Small Language Models, and that C2-Ultra reaches 1.7x the AI performance of the previous C1-Ultra while using up to 38% less power at the same workload. The accompanying Mali G2-Ultra NX GPU adds dedicated neural accelerators and claims up to 4x performance per watt for neural graphics.
Two independent sources cover the partnership. Arm's own newsroom is the primary source, and 24/7 Wall St. breaks down the silicon economics alongside data center incumbents.
Why On-Device Is The Real Story
Read the two announcements together and they describe a single engineering bet, not two separate products. The Samsung SoC and the C2 cluster both exist to pull inference onto the phone.
The table below lays out the two platforms and the concrete numbers behind each.
| Element | Platform | What it is | Key figure |
|---|---|---|---|
| CPU cluster | CSS for Mobile 2 (Arm C2) | C2-Ultra + C2-Pro cores, two SME2 vector units | 1.7x AI performance vs C1-Ultra at 38% less power |
| GPU | Mali G2-Ultra NX | Dedicated neural accelerators + ray tracing | 4x performance per watt for neural graphics |
| Accelerator | Arm-Samsung SoC | 2nm SF2 GAA, Arm IP, Samsung foundry | On-device inference, reduced cloud dependence |
| Market | Arm Neoverse | Data center cores in hyperscaler CPUs | 1.5 billion cores shipped, roughly 50% of hyperscaler share |
The data-science angle is economics. Every token a model generates in the cloud costs GPU time, memory bandwidth, and network egress. A chat that returns 500 tokens to a phone might burn 20 to 50 milliseconds of frontier-GPU time and cost fractions of a cent, but scale that across billions of devices running constantly and the marginal cost compounds fast. Move that same model onto the device and the cost curve inverts: the per-inference cloud bill collapses to near zero, and the latency drops to the local memory cycle time. The trade is a heavier upfront silicon cost and a smaller model that fits in on-device RAM.
Small Language Models are the obvious fit, which is why Arm doubled its SME2 units rather than just boosting clock speed. SME2 is a vector extension tuned for the matrix operations LLM inference spends most of its time doing. Doubling the unit count is a direct throughput play on token generation, and the 70% speedup figure reflects exactly that.
The engineering angle is the same story from the other side. GAA transistors at 2nm switch faster while leaking less current than the FinFET design that dominated for over a decade. Lower leakage matters enormously for a phone, where a model running continuously would otherwise drain the battery and throttle the SoC. Arm's own numbers show the pressure: 1.7x more AI work at 38% less power. That is a power-budget problem solved with architecture rather than a bigger battery.
The Data Center Is Not Going Away
The on-device bet is real, but it would be a mistake to read it as Arm displacing the data center. Arm's own center-of-gravity points the other way.
CEO Rene Haas said data center royalty revenue more than doubled year over year, that Arm Neoverse shipments have surpassed 1.5 billion cores, and that Arm is targeting a $15 billion silicon business against a data center TAM exceeding $100 billion by 2030. The AGI CPU has already drawn more than $2 billion in demand across fiscal 2027 and 2028. Arm silicon already sits inside Nvidia's Vera CPU, Google's Axion, Microsoft's Cobalt, and Amazon's Graviton 5, which puts it at roughly 50% of CPU compute share among the top hyperscalers. A full breakdown of the data center economics, including how the on-device deal fits against the incumbent, is on 24/7 Wall St..
Meanwhile the incumbent Arm must coexist with is still expanding. Nvidia's fiscal Q2 2027 data center revenue hit $89.02 billion, up 105.8% year over year, and Jensen Huang said demand is growing 100% a year even though Nvidia can only supply about 70% of it. Revenue opportunity per gigawatt is climbing from roughly $18 billion on Hopper to $40 billion on Vera Rubin.
So Arm is playing both ends. On-device inference eats a slice of the workload that used to require a round trip to the cloud, but the total volume of AI compute keeps growing because the models getting better every generation. More inference per device does not shrink the data center; it just shifts where a portion of that inference happens.
The Engineer's Read
The Samsung deal quietly does something structural beyond the chip itself. Until now, TSMC has been the unavoidable toll booth for leading-edge AI silicon, fabricating Nvidia's Rubin, Arm's AGI partners, and Qualcomm's custom accelerators. Samsung manufacturing an Arm-designed SoC on its own SF2 node is a real path toward partial disintermediation of TSMC for a class of products. The bull case is that whoever wins the accelerator war still needs wafers, and the risk is that Samsung's 2nm ramp is guided to dilute its own gross margin by 3 to 4 percentage points in Q3.
From a builder's perspective, the most useful number is the SME2 doubling. A 70% speedup on small language models at fixed power is the kind of gain that lets a 3B parameter model run acceptably on a phone instead of degrading into a slow cloud call. That is a product-defining improvement, not a benchmark curiosity.
Arm is priced for the data center at a P/E of roughly 298, after a 130% year-to-date run. That valuation assumes the $15 billion silicon business materializes while an Arm license dispute with Qualcomm heads toward a Q4 2026 trial. The on-device partnership with Samsung is real, additive, and genuinely efficient, but it does not close that legal risk or explain the multiple.
For the people who actually ship models, the takeaway is simpler: on-device AI is moving from a nice-to-have to a first-class platform, and the silicon to support it just got meaningfully faster. You can read the earlier look at running large local models on a workstation at Ryzen AI Halo local AI to see the other end of the same spectrum, far from the phone and close to the rack.
Outlook
Expect the next 12 months to be about whether on-device inference can move past demo models into the models people actually use. The architecture is arriving on time. The question is whether the smaller model fits the workload well enough to make the cloud round trip optional rather than default, and whether Samsung's 2nm can actually ship at volume while Arm's data center business keeps compounding. Both bets are the same bet, made twice.