REAL-TIME GLOBAL RESEARCH
US Semiconductors and Hardware: Weekly Investor Pulse: RAISE Summit Takeaways
Research evidence excerpt
US Semiconductors and Hardware: Weekly Investor Pulse: RAISE Summit Takeaways
US Semiconductors and Hardware
12 July 2026 Citi Research
broader electronics ecosystem. Major hyperscalers are already pursuing multi-year supply agreements to secure production
capacity. Industry supply-demand dynamics suggest that meaningful relief may not arrive until at least 2028, when new
manufacturing capacity is expected to come online.
Hyperscaler Custom Silicon: Traditional CPU and GPU architectures are increasingly constrained by networking, power, and
efficiency bottlenecks. In response, hyperscalers are accelerating investment in custom silicon purpose-built for AI
workloads. These chips are optimized around specific models, inference tasks, and agentic workflows running on internal
infrastructure. Google’s latest roadmap illustrates this specialization, with separate silicon platforms dedicated to low-
latency inference (8i) and large-scale model training (8t), highlighting the industry’s broader move toward workload-
specific compute architectures.
The KV Cache Bottleneck: As the cost of multi-turn agentic workflows continues to rise, hyperscalers are increasingly
focused on optimizing context memory management. A key strategy involves offloading KV cache data from expensive GPUs
to lower-cost storage layers, reducing redundant computations and improving system utilization. These techniques can
deliver cache-hit rates exceeding 95%, reduce time-to-first-token latency by 20x, and materially lower inference costs.
Moreover, because flash is in shortage, hybrid storage architectures including hard disk drives are seeing a resurgence for
massive AI datasets.
Latency as a Competitive Advantage: Within enterprise environments, inference speed is becoming a critical determinant of
The English excerpt is extracted automatically from the cited source page and may contain layout or recognition errors. It is never batch translated.
Open report viewer