ReportGem ReportGem 中文

REAL-TIME GLOBAL RESEARCH

US Semiconductors: UBS Chip Chats: Ultra-fast and Disaggregated Inferencing

Published: 2026-06-29Institution: UBS EquitiesPages: 12Original language: EnglishEvidence page: 1

Research evidence excerpt

US Semiconductors: UBS Chip Chats: Ultra-fast and Disaggregated Inferencing

Global Research

29 June 2026ab

US Semiconductors Equities

AmericasUBS Chip Chats: Ultra-fast and Disaggregated

Inferencing Semiconductors

Timothy Arcuri

Analyst

timothy.arcuri@ubs.com

Summary +1-415-352 5676

This past week, we hosted an expert call with a former META/INTC hardware engineer to Natalia Winkler, CFA

discuss emerging system architectures for AI inference. We covered ultra-fast Analyst

inferencing and the trade-offs of SRAM architectures, memory hierarchy, and latency natalia.winkler@ubs.com

constraints across different systems. We also explored the disaggregated inferencing +1-415-352 4626

approach that players like NVDA and AWS have been pursuing, as well as the challenges Alex Kivali

associated with this new heterogeneous model. The decode stage of inference is unique Analyst

in that, unlike training or prefill, it is structurally memory-bound rather than compute- alex.kivali@ubs.com

bound, which creates opportunities for non-HBM architectures to address memory +1-212-713 3945

constraints. Gianmarco Vella

Associate Analyst

SRAM-based Architectures: Optimizing for Low-Latency Inference gianmarco.vella@ubs.com

SRAM-based architectures tightly couple compute and offer a structural advantage for +1-415-352 4555

inference decode workloads, where performance is dominated by memory bandwidth Aaryan Wadhwa

and latency rather than compute/FLOPs. By colocating high-speed SRAM directly with Associate Analyst

compute, these systems eliminate reliance on external HBM for critical paths, materially aaryan.wadhwa@ubs.com

reducing data movement and control latency. The result is longer uninterrupted +1-212-821 6481

execution sequences and meaningfully improved per-user interactivity.

The English excerpt is extracted automatically from the cited source page and may contain layout or recognition errors. It is never batch translated.

Open report viewer