AMD Advancing AI 2026 day one: all about the inference
I'm writing this from Moscone West in San Francisco, at the end of day one of AMD Advancing AI 2026.
The headlines arrived early: EPYC "Venice," the first x86 server chip in volume production on TSMC's 2nm process; the Instinct MI450 series; and Helios, a rack-scale system that puts 72 GPUs and 31TB of HBM4 memory in a single double-wide rack. Between them, Meta and OpenAI have committed something like 12 gigawatts of AMD compute. Wow.
You'll be able to read the announcement coverage everywhere by the time this posts, so I won't rehash it. What I found more interesting, flipping through the program and trying to decide which of the 100+ sessions and 30+ workshops to sit in, was how often the same few ideas kept coming up. Four of them, by my count. It's possible I'm pattern-matching, one attendee with one schedule, but here goes.
Inference has taken over the agenda
These events used to be mostly about training: bigger clusters, bigger models, bigger benchmark numbers. This year the balance has tipped. Two of the headline luminary speakers are Simon Mo of the vLLM project and Ying Sheng, co-creator of SGLang, the serving engines you'll find underneath a large share of production LLM deployments. The workshop track is thick with sessions on inference tuning, serving optimization, and AMD's catalog of pre-optimized models (AIMs). Goldman Sachs Research reckons token consumption will grow 24-fold by 2030, to 120 quadrillion tokens a month.
Even AMD's own hands-on workshop on hybrid multi-agent systems sells itself on achieving "optimum token expenditure." When a chip vendor's workshop copy starts worrying about your token bill, either something real has shifted or the marketing team knows its audience very well. My money is on both. Every coding agent and agentic workflow a team adopts turns compute from a capital line item into a meter that runs whether anyone's watching or not, and I don't think most organizations have worked out who watches the meter yet.
Open is a strategy with delivery dates
AMD isn't being subtle about its pitch this week: open source ROCm software, the open UALink interconnect inside Helios, the OpenClaw agent framework, and a speaker roster (George Hotz, Chris Lattner, the PyTorch and Linux Foundations) that reads like a maintainers' summit. The argument, as I read it, is that if you build on open standards and open software, you can change your mind later without rewriting everything.
Parts of that story are still a roadmap, though. The first Helios systems run UALink over Ethernet until dedicated switching silicon arrives next year, and ROCm's training story trails its inference story, where it's now a first-class citizen in PyTorch, vLLM and SGLang. None of this makes the open bet wrong. As a company that builds on open source, we'd like it to pay off. But "open" is probably best treated as a direction of travel with milestones attached.
AI factories are a facilities and finance problem too
The infrastructure track covers network transport for AI factories (how Multipath Reliable Connection gets past the limits of RoCEv2), rack-scale deployment patterns ("scale-up first, scale-out when it matters"), power and density constraints, and the true total cost of ownership of AI infrastructure. Research shared by theCUBE from its studio on the expo floor found 64% of organizations surveyed pointing to data and infrastructure bottlenecks as their main obstacle to AI deployment, not model availability. That matches what we see with customers: the models are mostly good enough now, and the pain has moved down the stack.
Near-term Helios supply is reportedly spoken for by the hyperscalers, and analysts put broader availability well into 2027. So for most enterprises, I suspect it's more about making the GPUs you already have, or can plausibly get, earn their keep in production, with the sort of controls a CFO and a security team will both sign off on. Though if your capex budget says otherwise, I'd be curious how you swung it.
Is inference moving back toward the data?
Alongside all the gigawatt talk there's a second thread in the program, easier to miss. Sovereign AI has its own sessions and its own silicon now (the MI430X is aimed at sovereign and FP64 workloads). AMD's "Agent Computer" concept, which shows up across two workshops, routes work between a Ryzen AI PC on your desk and Instinct GPUs in the data center depending on cost, latency and sensitivity. And there's a workshop called "Vibe Coding with Local Models," which is either a sign of where developer culture is heading or evidence that AMD's events team has a sense of humor. Possibly both.
The drivers behind this thread are ones we hear daily from our own customers: data residency, regulatory pressure in healthcare, finance, defense and telco, and a growing awareness that sending every request to a frontier API is a privacy decision as well as a pricing decision, made implicitly, thousands of times a day.
Where Spectro Cloud fits (and where to find us)
Those last two themes are why we're here in force this week. AMD is a partner: PaletteAI supports AMD-powered AI infrastructure end to end, from AMD Instinct GPUs and the GPU Operator through the ROCm runtime to AMD-optimized models from the AIMs catalog. AMD Ventures also took part in our $100M+ Series D, announced last week.
And this week we launched PaletteAI Inference Launchpad, a turnkey, locally managed inference stack built on one operating principle: local by default, frontier when needed. It runs validated open-weight models on your own GPUs, weighs each request for sensitivity, task type, policy, quota and capacity, then routes it to the best-fit local or frontier model, metering every token as it goes. Developers keep the tools they already use, and the platform team gets to see what those tools consume, with some controls to do something about it. Get the routing right and, depending on workload mix, token costs can drop by up to 70%. It's validated on AMD Instinct GPUs as well as NVIDIA, which felt like the right week to mention.
If you're at the show, come say hello. We're at demo station 515A in the ISV Pavilion, next to the Expo Theater stage, with live demos through Thursday. Our CTO and co-founder Saad Malik gave a 15-minute Expo Theater talk on Wednesday ("Bring AI inference home. Lose the token tax."), and the team is hosting demos and slower conversations in our hospitality suite in the Living Room at the InterContinental, a couple of minutes' walk from Moscone. You might also see some bright yellow tees and some vibrant signage outside the venue...

If any of the themes above landed close to home, two suggestions. The Inference Launchpad announcement is a short read on what local-first inference with real governance looks like. And our AI TCO calculator will give you a first pass at your own numbers. The gigawatt deals are a hyperscaler business for now. The token bill is much closer to home, and that one you can do something about before 2027.
Find the Spectro Cloud team at AMD Advancing AI 2026, July 22–23, at demo station 515A, or visit spectrocloud.com/amd to book a meeting. We’ll also be at Ai4 in Las Vegas in August.

.jpg)