Skip to main content
Models & Research

OpenBMB's MiniCPM5-2B Gets a 324M Draft Model That Accepts 6 Tokens at Once

Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.

By Precis Daily Newsroom2 min read333 words
Illustration for: OpenBMB's MiniCPM5-2B Gets a 324M Draft Model That Acce
Illustration
Key points
  • Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B.
  • Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.
  • Math and code domains hit 6.05 and 6.11 accepted tokens per verification step respectively.

Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0. Math and code domains hit 6.05 and 6.11 accepted tokens per verification step respectively. Uses DSpark semi-autoregressive drafting: parallel backbone plus lightweight sequential head to prevent suffix decay. Served via SGLang with --speculative-algorithm DSPARK and block size 7, Apache 2.0 licensed. Trained on 7.05B tokens, 6 epochs, with cross-entropy, L1, and confidence loss objectives. OpenBMB adds a 324M DSpark draft to MiniCPM5-2B OpenBMB has released the DSpark checkpoint , a 323.8M-parameter draft model trained specifically to accelerate MiniCPM5-2B through speculative decoding. SGLang uses the smaller model to propose seven-token blocks, then asks the 2.52B-parameter target to verify those proposals in one pass. MiniCPM5-2B is a dense next-token model with 42 layers, grouped-query attention, and a native context window of 131,072 tokens. Its LlamaForCausalLM architecture allows established inference engines to load the model without custom attention kernels. A published benchmark overview reports a 53.9 average across 34 evaluations, compared with 51.1 for Qwen3.5-4B. Reported results include 69.1 on LiveCodeBench v6 and 97.1 on τ2-Bench Telecom. Aggregate scores combine distinct tasks, so deployment tests remain necessary for any specific workload. Speculative decoding lets a smaller model propose upcoming tokens at lower computational cost. The target evaluates the proposed block in parallel, accepts tokens that satisfy its verification rule, and regenerates from the first rejected position. The target remains responsible for the final token sequence, while draft accuracy determines how much work each verification pass completes. DSpark addresses the acceptance decay common to parallel drafters, whose later proposals often lack enough information about earlier tokens in the same block. Its parallel backbone generates the initial proposals, a lightweight sequential module adds dependencies within the block, and a confidence scheduler chooses how many proposals to verify based on estimated prefix survival and the serving engine’s throughput profile.

Sources

Summarized from the linked originals.

Related stories

Illustration for: Predictive Analytics Meets AI
Models & Research

The predictive analytics landscape has evolved significantly in recent years, with AI-powered analytics driving growth and innovation across various industries.

MIT Technology Review: AI1 min