GLOBAL RESEARCH ARCHIVE
AMD: Advancing AI 2026 Keynote
Research evidence excerpt
AMD: Advancing AI 2026 Keynote
July 24, 2026
Customer System Co-Design Becomes the Real Differentiator
• AMD emphasized that large AI customers are moving from buying chips to co-designing
full systems across compute, power, cooling, networking, and software.
• Meta reinforced this point, saying AI infrastructure is no longer just a GPU game and
that CPUs and GPUs need to be treated as “conjoined things.”
• Meta’s MI450 comments were especially important because it described a move from
MI300 experimentation to production-scale deployment and deeper engineering co-
design.
Inference Segmenting Into Different Workloads
• AMD indicated that inference is not one market: some workloads need maximum
throughput, some need balanced responsiveness, and some require ultra-low latency.
• This supports a broader portfolio argument, with AMD positioning different compute
architectures for different inference use cases rather than a one-size-fits-all solution.
Cerebras Highlights Ultra-Low Latency Inference
• Cerebras was the clearest proof point for disaggregated inference, combining AMD
CPUs, Helios, and the Cerebras Wafer Scale Engine.
• The partnership is aimed at customers that need both high throughput and very low
latency, with Cerebras claiming 5x throughput while maintaining speed.
OpenAI Reinforces the Need for More Compute
• OpenAI framed models as evolving from chatbots to reasoners, agents, and eventually
“interns,” which keeps pushing demand for scaled compute.
• OpenAI also said exploding token budgets are showing up internally across engineering
and broader enterprise workflows, not just in research workloads.
Enterprise AI Moves From Use Cases to Workflow Rebuilds
• AT&T said AI is moving beyond point use cases into rebuilding entire workflows across
The English excerpt is extracted automatically from the cited source page and may contain layout or recognition errors. It is never batch translated.
Open report viewer