ReportGem ReportGem EN

普通外文研报

China AI: DS Brings Higher Inference Efficiency; More Price Competition

发布日期: 2026-06-28研究机构: Jefferies报告页数: 7原文语言: 英语证据页码: 1

研报英文原文证据摘录

China AI: DS Brings Higher Inference Efficiency; More Price Competition

China (PRC) | Technology EquityJuneResearch28, 2026

China AI: DS Brings Higher Inference

Efficiency; More Price Competition

DeepSeek’s (DS) new DSpark speculative decoding framework boosts

inference efficiency by raising V4's output speed and token generation per

GPU, improving user experience and lowering cost per token. Its release

of open-source toolkit DeepSpec for inference optimization would help

other LLM peers adopt similar techniques. This could intensify API price

competition but drive AI adoption. Compute/memory constraint is still the

key growth headwind for China AI.

DSpark increases output speed, GPU efficiency and lowers inference cost per token. On

June 27, DS released DSpark, a speculative decoding framework for inference optimization. It

uses a lightweight draft model to guess several next tokens first, then lets the DS-V4 target

model verify those guesses in batches. Thus the system can generate multiple correct tokens

from its target model run instead of decoding strictly one token at a time. GOOG released

something similar in 2023, but DS now achieved further improvement. DS also released

DeepSpec, an open-source training and evaluation toolkit that helps LLM developers train

and test draft models, making DSpark-like inference optimization easier to develop for other

AI players. DSpark could: 1) Increase output speed: compared with previous system that

guesses one token at a time, DSpark increases per-user generation speed by 60-85% for DS-V4-

Flash and 57-78% for V4-Pro, while the system was handling a similar amount of user traffic. 2)

Increase GPU efficiency/ lower cost per token: under normal production conditions, DSpark

本摘录由系统从所标注的 PDF 证据页直接提取并保留英文原文,不做批量翻译;登录后在阅读器切换中文时才按需翻译。

打开研报阅读器