OpenBMB's MiniCPM5-2B Gets a 324M Draft Model That Accepts 6 Tokens at Once
Subtopic Small Models · Long Context · Inference Optimization OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.

- Subtopic Small Models · Long Context · Inference Optimization OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B.
- Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.
- Math and code domains hit 6.05 and 6.11 accepted tokens per verification step respectively.
Subtopic Small Models · Long Context · Inference Optimization
OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B.
Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.
Math and code domains hit 6.05 and 6.11 accepted tokens per verification step respectively.
Uses DSpark semi-autoregressive drafting: parallel backbone plus lightweight sequential head to prevent suffix decay.
Served via SGLang with --speculative-algorithm DSPARK and block size 7, Apache 2.0 licensed.
Trained on 7.05B tokens, 6 epochs, with cross-entropy, L1, and confidence loss objectives.
OpenBMB adds a 324M DSpark draft to MiniCPM5-2B
OpenBMB has released the DSpark checkpoint , a 323.8M-parameter draft model trained specifically to accelerate MiniCPM5-2B through speculative decoding. SGLang uses the smaller model to propose seven-token blocks, then asks the 2.52B-parameter target to verify those proposals in one pass.
MiniCPM5-2B is a dense next-token model with 42 layers, grouped-query attention, and a native context window of 131,072 tokens. Its LlamaForCausalLM architecture allows established inference engines to load the model without custom attention kernels.
A published benchmark overview reports a 53.9 average across 34 evaluations, compared with 51.1 for Qwen3.5-4B. Reported results include 69.1 on LiveCodeBench v6 and 97.1 on τ2-Bench Telecom. Aggregate scores combine distinct tasks, so deployment tests remain necessary for any specific workload.
Speculative decoding lets a smaller model propose upcoming tokens at lower computational cost. The target evaluates the proposed block in parallel, accepts tokens that satisfy its verification rule, and regenerates from the first rejected position. The target remains responsible for the final token sequence, while draft accuracy determines how much work each verification pass completes.
DSpark addresses the acceptance decay common to parallel drafters, whose later proposals often lack enough information about earlier tokens in the same block. Its parallel backbone generates the initial proposals, a lightweight sequential module adds dependencies within the block, and a confidence scheduler chooses how many proposals to verify based on estimated prefix survival and the serving engine’s throughput profile.
Sources
Related stories

Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

Mitsuba Squeezes a 27B Vision Model Into 7.3 GB on One GPU
Subtopic Small Models · Vision Language · Quantization Mitsuba is a ternary 1.58-bit quantization of Qwen3.8-27B, shrunk to 7.3 GB for a single 16 GB GPU. Purpose-built for ComfyUI: image to prompt generation for Stable Diffusion, Krea, and video pipelines.

Google's new Gemini tiers cut free users to its weakest model and lock $5/month subscribers out of Pro
Google's new Gemini tiers cut free users to its weakest model and lock $5/month subscribers out of Pro Starting in October 2026, Google will restrict access to its Gemini models for personal account users.

ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second
ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight. The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.

Aleph Alpha Releases Kolibri-1, an Open German Reasoning Model With 1M-Token Context
Subtopic Mixture Of Experts · Long Context Kolibri-1 is a 78B MoE with 3.46B active parameters, Apache 2.0, German and English focus. Context window validated up to 1,048,576 tokens, native 262,144, no position scaling tricks required.

Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

Google announces Gemini 4 Argon AI model, but you can't use it yet
Google promised Gemini 3.5 Pro in June, but it spent the summer trotting out smaller Flash models.

Kling 4.0 Debuts 30s AI Video Model - Briefs Finance
Warning : Undefined variable $stocks in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 448 Warning : foreach() argument must be of type arrayobject, null given in.