普通外文研报
China AI: DS Brings Higher Inference Efficiency; More Price Competition
研报英文原文证据摘录
China AI: DS Brings Higher Inference Efficiency; More Price Competition
China (PRC) | Technology EquityJuneResearch28, 2026
China AI: DS Brings Higher Inference
Efficiency; More Price Competition
DeepSeek’s (DS) new DSpark speculative decoding framework boosts
inference efficiency by raising V4's output speed and token generation per
GPU, improving user experience and lowering cost per token. Its release
of open-source toolkit DeepSpec for inference optimization would help
other LLM peers adopt similar techniques. This could intensify API price
competition but drive AI adoption. Compute/memory constraint is still the
key growth headwind for China AI.
DSpark increases output speed, GPU efficiency and lowers inference cost per token. On
June 27, DS released DSpark, a speculative decoding framework for inference optimization. It
uses a lightweight draft model to guess several next tokens first, then lets the DS-V4 target
model verify those guesses in batches. Thus the system can generate multiple correct tokens
from its target model run instead of decoding strictly one token at a time. GOOG released
something similar in 2023, but DS now achieved further improvement. DS also released
DeepSpec, an open-source training and evaluation toolkit that helps LLM developers train
and test draft models, making DSpark-like inference optimization easier to develop for other
AI players. DSpark could: 1) Increase output speed: compared with previous system that
guesses one token at a time, DSpark increases per-user generation speed by 60-85% for DS-V4-
Flash and 57-78% for V4-Pro, while the system was handling a similar amount of user traffic. 2)
Increase GPU efficiency/ lower cost per token: under normal production conditions, DSpark
本摘录由系统从所标注的 PDF 证据页直接提取并保留英文原文,不做批量翻译;登录后在阅读器切换中文时才按需翻译。
打开研报阅读器