GLOBAL RESEARCH ARCHIVE
China AI: DS Brings Higher Inference Efficiency; More Price Competition
Research evidence excerpt
China AI: DS Brings Higher Inference Efficiency; More Price Competition
China (PRC) | Technology EquityJuneResearch28, 2026
China AI: DS Brings Higher Inference
Efficiency; More Price Competition
DeepSeek’s (DS) new DSpark speculative decoding framework boosts
inference efficiency by raising V4's output speed and token generation per
GPU, improving user experience and lowering cost per token. Its release
of open-source toolkit DeepSpec for inference optimization would help
other LLM peers adopt similar techniques. This could intensify API price
competition but drive AI adoption. Compute/memory constraint is still the
key growth headwind for China AI.
DSpark increases output speed, GPU efficiency and lowers inference cost per token. On
June 27, DS released DSpark, a speculative decoding framework for inference optimization. It
uses a lightweight draft model to guess several next tokens first, then lets the DS-V4 target
model verify those guesses in batches. Thus the system can generate multiple correct tokens
from its target model run instead of decoding strictly one token at a time. GOOG released
something similar in 2023, but DS now achieved further improvement. DS also released
DeepSpec, an open-source training and evaluation toolkit that helps LLM developers train
and test draft models, making DSpark-like inference optimization easier to develop for other
AI players. DSpark could: 1) Increase output speed: compared with previous system that
guesses one token at a time, DSpark increases per-user generation speed by 60-85% for DS-V4-
Flash and 57-78% for V4-Pro, while the system was handling a similar amount of user traffic. 2)
Increase GPU efficiency/ lower cost per token: under normal production conditions, DSpark
The English excerpt is extracted automatically from the cited source page and may contain layout or recognition errors. It is never batch translated.
Open report viewer