REAL-TIME GLOBAL RESEARCH
Must C
Research evidence excerpt
Must C
However, constraints cover the full stack of energization —
transformers, switchgear, turbines, fuel cells, substation work, and
skilled labor. That is why the market is increasingly bifurcating between
those who can secure physical capacity and those who must buy access
from them. This strengthens the hand of hyperscalers and infrastructure
owners relative to model labs that do not control cloud, chips, sites, or
power procurement.
Improving energy efficiency is a key focus, through both better
performance per watt and power usage effectiveness (PUE). Better
performance per watt takes place at the chip and software level through
improvements in liquid cooling designs, re-architecting software and
newer silicon generation to increase the output per watt. Put differently,
PUE improvements occur at the facility level before electricity even gets to
the servers.
As AI costs rise, we note a shift under way in select use cases toward
open-weights (partially open) models, reflecting their lower inference
costs relative to frontier closed-source models.3 (Toms Hardware, May
23) However, lower inference costs would drive higher overall compute
demand, increasing data center utilization — consistent with Jevons
Paradox, where falling costs lead to greater usage. (Bloomberg, June 10)
Model use in some areas would partly migrate from frontier models such as
GPT-5.5 and Opus 4.8 to those such as GLM 5.2 or V4 Pro or their
equivalents in the future. The scaling law from Hoffman (2022), which had
shown that optimal performance under a fixed compute budget is
achieved by scaling parameters and training tokens jointly and
approximately linearly4 (and which has helped justify vast investments in
the past), has since been superseded by over-training smaller architecture
The English excerpt is extracted automatically from the cited source page and may contain layout or recognition errors. It is never batch translated.
Open report viewer