OpenBMB's MiniCPM5-2B Gets a 324M Draft Model That Accepts 6 Tokens at Once
Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.

- Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B.
- Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.
- Math and code domains hit 6.05 and 6.11 accepted tokens per verification step respectively.
Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0. Math and code domains hit 6.05 and 6.11 accepted tokens per verification step respectively. Uses DSpark semi-autoregressive drafting: parallel backbone plus lightweight sequential head to prevent suffix decay. Served via SGLang with --speculative-algorithm DSPARK and block size 7, Apache 2.0 licensed. Trained on 7.05B tokens, 6 epochs, with cross-entropy, L1, and confidence loss objectives. OpenBMB adds a 324M DSpark draft to MiniCPM5-2B OpenBMB has released the DSpark checkpoint , a 323.8M-parameter draft model trained specifically to accelerate MiniCPM5-2B through speculative decoding. SGLang uses the smaller model to propose seven-token blocks, then asks the 2.52B-parameter target to verify those proposals in one pass. MiniCPM5-2B is a dense next-token model with 42 layers, grouped-query attention, and a native context window of 131,072 tokens. Its LlamaForCausalLM architecture allows established inference engines to load the model without custom attention kernels. A published benchmark overview reports a 53.9 average across 34 evaluations, compared with 51.1 for Qwen3.5-4B. Reported results include 69.1 on LiveCodeBench v6 and 97.1 on τ2-Bench Telecom. Aggregate scores combine distinct tasks, so deployment tests remain necessary for any specific workload. Speculative decoding lets a smaller model propose upcoming tokens at lower computational cost. The target evaluates the proposed block in parallel, accepts tokens that satisfy its verification rule, and regenerates from the first rejected position. The target remains responsible for the final token sequence, while draft accuracy determines how much work each verification pass completes. DSpark addresses the acceptance decay common to parallel drafters, whose later proposals often lack enough information about earlier tokens in the same block. Its parallel backbone generates the initial proposals, a lightweight sequential module adds dependencies within the block, and a confidence scheduler chooses how many proposals to verify based on estimated prefix survival and the serving engine’s throughput profile.
Sources
Related stories

Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

Predictive Analytics Meets AI
The predictive analytics landscape has evolved significantly in recent years, with AI-powered analytics driving growth and innovation across various industries.

Mitsuba Squeezes a 27B Vision Model Into 7.3 GB on One GPU
Subtopic Small Models · Vision Language · Quantization Takeaways − Mitsuba is a ternary 1.58-bit quantization of Qwen3.8-27B, shrunk to 7.3 GB for a single 16 GB GPU. Purpose-built for ComfyUI: image to prompt generation for Stable Diffusion, Krea, and video pipelines.

China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB Locally
Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.