实时全球研报
US Semiconductors and Hardware: Weekly Investor Pulse: RAISE Summit Takeaways
研报英文原文证据摘录
US Semiconductors and Hardware: Weekly Investor Pulse: RAISE Summit Takeaways
US Semiconductors and Hardware
12 July 2026 Citi Research
broader electronics ecosystem. Major hyperscalers are already pursuing multi-year supply agreements to secure production
capacity. Industry supply-demand dynamics suggest that meaningful relief may not arrive until at least 2028, when new
manufacturing capacity is expected to come online.
Hyperscaler Custom Silicon: Traditional CPU and GPU architectures are increasingly constrained by networking, power, and
efficiency bottlenecks. In response, hyperscalers are accelerating investment in custom silicon purpose-built for AI
workloads. These chips are optimized around specific models, inference tasks, and agentic workflows running on internal
infrastructure. Google’s latest roadmap illustrates this specialization, with separate silicon platforms dedicated to low-
latency inference (8i) and large-scale model training (8t), highlighting the industry’s broader move toward workload-
specific compute architectures.
The KV Cache Bottleneck: As the cost of multi-turn agentic workflows continues to rise, hyperscalers are increasingly
focused on optimizing context memory management. A key strategy involves offloading KV cache data from expensive GPUs
to lower-cost storage layers, reducing redundant computations and improving system utilization. These techniques can
deliver cache-hit rates exceeding 95%, reduce time-to-first-token latency by 20x, and materially lower inference costs.
Moreover, because flash is in shortage, hybrid storage architectures including hard disk drives are seeing a resurgence for
massive AI datasets.
Latency as a Competitive Advantage: Within enterprise environments, inference speed is becoming a critical determinant of
本摘录由系统从所标注的 PDF 证据页直接提取并保留英文原文,不做批量翻译;登录后在阅读器切换中文时才按需翻译。
打开研报阅读器