普通外文研报
US Semiconductors "UBS Chip Chats: Ultra-fast and Disaggregated..."
研报英文原文证据摘录
US Semiconductors "UBS Chip Chats: Ultra-fast and Disaggregated..."
Global Research
29 June 2026ab
US Semiconductors Equities
AmericasUBS Chip Chats: Ultra-fast and Disaggregated
Inferencing Semiconductors
Timothy Arcuri
Analyst
timothy.arcuri@ubs.com
Summary +1-415-352 5676
This past week, we hosted an expert call with a former META/INTC hardware engineer to Natalia Winkler, CFA
discuss emerging system architectures for AI inference. We covered ultra-fast Analyst
inferencing and the trade-offs of SRAM architectures, memory hierarchy, and latency natalia.winkler@ubs.com
constraints across different systems. We also explored the disaggregated inferencing +1-415-352 4626
approach that players like NVDA and AWS have been pursuing, as well as the challenges Alex Kivali
associated with this new heterogeneous model. The decode stage of inference is unique Analyst
in that, unlike training or prefill, it is structurally memory-bound rather than compute- alex.kivali@ubs.com
bound, which creates opportunities for non-HBM architectures to address memory +1-212-713 3945
constraints. Gianmarco Vella
Associate Analyst
SRAM-based Architectures: Optimizing for Low-Latency Inference gianmarco.vella@ubs.com
SRAM-based architectures tightly couple compute and offer a structural advantage for +1-415-352 4555
inference decode workloads, where performance is dominated by memory bandwidth Aaryan Wadhwa
and latency rather than compute/FLOPs. By colocating high-speed SRAM directly with Associate Analyst
compute, these systems eliminate reliance on external HBM for critical paths, materially aaryan.wadhwa@ubs.com
reducing data movement and control latency. The result is longer uninterrupted +1-212-821 6481
execution sequences and meaningfully improved per-user interactivity. However, SRAM
本摘录由系统从所标注的 PDF 证据页直接提取并保留英文原文,不做批量翻译;登录后在阅读器切换中文时才按需翻译。
打开研报阅读器