ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second
ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight. The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.

- ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight.
- The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.
- Per-tensor quantization is scored against BF16 on GSM8K, MMLU, LiveCodeBench, IFEval, BFCL and ACEBench.
ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight.
The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.
Per-tensor quantization is scored against BF16 on GSM8K, MMLU, LiveCodeBench, IFEval, BFCL and ACEBench.
Embedded MTP head adds less than 250 MB and delivers 1.28 to 1.66x decode speedup.
External DFlash 2 draft reaches 1.34 to 2.10x speedup, text-only, needs llama.cpp b10658+.
Apache 2.0 license, vision capable, runs via llama.cpp, Ollama, LM Studio, vLLM and SGLang.
ByteShape compresses Qwen3.8-27B below 14 GB
ByteShape has released five GGUF quantizations of the Qwen3.8-27B vision-language model. The 8.8 GB to 13.1 GB files use ShapeLearn, a per-tensor quantization system designed to preserve sensitive weights at low bit depths. In ByteShape’s RTX 5090 tests, speculative decoding reached 176 tokens per second with the smallest checkpoint.
A 27-billion-parameter model stored in BF16 requires roughly 54 GB for weights before runtime overhead. Quantization cuts that requirement by storing weights with fewer bits. ShapeLearn selects the datatype and quantization method for each tensor, assigning more precision to weights its optimizer identifies as sensitive. GGUF packages those quantized weights for llama.cpp-compatible local runtimes.
The five GPU-oriented builds span 2.56 to 3.84 average bits per weight, abbreviated as bpw. Lower values reduce storage and memory use, usually at the cost of some model quality.
A 16 GB card has enough capacity for each checkpoint file, although total VRAM use also includes the KV cache, runtime buffers, vision projector, and any speculative draft model. Context length, cache format, and GPU offloading determine whether a particular configuration fits entirely in memory.
Filename tags such as IQ4_XS and IQ2_XXS provide compatible indexing on Hugging Face. The underlying files contain ShapeLearn’s hybrid mix of quantization methods, so the tag alone does not describe every tensor’s format.
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
Sources
Related stories

Mitsuba Squeezes a 27B Vision Model Into 7.3 GB on One GPU
Subtopic Small Models · Vision Language · Quantization Mitsuba is a ternary 1.58-bit quantization of Qwen3.8-27B, shrunk to 7.3 GB for a single 16 GB GPU. Purpose-built for ComfyUI: image to prompt generation for Stable Diffusion, Krea, and video pipelines.

Vulnerabilities in AI Agents Expose Trust Gap Flaw
Google and other organizations have acknowledged vulnerabilities in their AI agents, which exploit trust gaps in the Model Context Protocol.

Chinese AI models parrot state doctrine or refuse to answer on sensitive topics
Chinese AI models frequently toe the party line when asked politically sensitive questions, according to a study by Aleph Alpha. The company markets itself alongside Cohere as a provider of "sovereign AI" for governments, giving it a commercial interest in distinguishing its.

OpenBMB's MiniCPM5-2B Gets a 324M Draft Model That Accepts 6 Tokens at Once
Subtopic Small Models · Long Context · Inference Optimization OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.

UMG, Sony Music File Second Lawsuit Against Suno Over AI Music Generator’s New Model - The Hollywood Reporter
Universal Music Group and Sony Music have filed a second copyright infringement lawsuit against AI music generation platform Suno , alleging that the company s recently released v6 model which Suno said is trained exclusively on licensed music uses their recordings without consent.

Teaching Everyone to Fish for Tokens
Nvidia wants you building your own model, not buying from Anthropic/OpenAI. Nathan Lambert Aug 17, 2026 84 8 Housekeeping: No voiceover for this post as I’m traveling.

Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

Predictive Analytics Meets AI
The predictive analytics landscape has evolved significantly in recent years, with AI-powered analytics driving growth and innovation across various industries.