ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second
Takeaways − ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight. The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.

- Takeaways − ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight.
- The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.
- Per-tensor quantization is scored against BF16 on GSM8K, MMLU, LiveCodeBench, IFEval, BFCL and ACEBench.
Takeaways − ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight. The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB. Per-tensor quantization is scored against BF16 on GSM8K, MMLU, LiveCodeBench, IFEval, BFCL and ACEBench. Embedded MTP head adds less than 250 MB and delivers 1.28 to 1.66x decode speedup. External DFlash 2 draft reaches 1.34 to 2.10x speedup, text-only, needs llama.cpp b10658+. Apache 2.0 license, vision capable, runs via llama.cpp, Ollama, LM Studio, vLLM and SGLang. ByteShape compresses Qwen3.8-27B below 14 GB ByteShape has released five GGUF quantizations of the Qwen3.8-27B vision-language model. The 8.8 GB to 13.1 GB files use ShapeLearn, a per-tensor quantization system designed to preserve sensitive weights at low bit depths. In ByteShape’s RTX 5090 tests, speculative decoding reached 176 tokens per second with the smallest checkpoint. A 27-billion-parameter model stored in BF16 requires roughly 54 GB for weights before runtime overhead. Quantization cuts that requirement by storing weights with fewer bits. ShapeLearn selects the datatype and quantization method for each tensor, assigning more precision to weights its optimizer identifies as sensitive. GGUF packages those quantized weights for llama.cpp-compatible local runtimes. The five GPU-oriented builds span 2.56 to 3.84 average bits per weight, abbreviated as bpw. Lower values reduce storage and memory use, usually at the cost of some model quality. A 16 GB card has enough capacity for each checkpoint file, although total VRAM use also includes the KV cache, runtime buffers, vision projector, and any speculative draft model. Context length, cache format, and GPU offloading determine whether a particular configuration fits entirely in memory. Filename tags such as IQ4_XS and IQ2_XXS provide compatible indexing on Hugging Face. The underlying files contain ShapeLearn’s hybrid mix of quantization methods, so the tag alone does not describe every tensor’s format. You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
Sources
Related stories

Vulnerabilities in AI Agents Expose Trust Gap Flaw
Google and other organizations have acknowledged vulnerabilities in their AI agents, which exploit trust gaps in the Model Context Protocol.

Chinese AI models parrot state doctrine or refuse to answer on sensitive topics
Chinese AI models parrot state doctrine or refuse to answer on sensitive topics Chinese AI models frequently toe the party line when asked politically sensitive questions, according to a study by Aleph Alpha. The company markets itself alongside Cohere as a provider of "sovereign AI" for governments, giving it a commercial interest in distinguishing its models from Chinese competitors.

Apparently, OpenAI isn't trying to build "magic intelligence in the sky" anymore
Apparently, OpenAI isn't trying to build "magic intelligence in the sky" anymore OpenAI CEO Sam Altman is pushing back against religious analogies tied to AI models.

Tavus' AI looks, listens, and talks back live
AI Tavus' AI looks, listens, and talks back live PLUS: Is ChatGPT “Dots” worth upgrading to Pro for? Good morning, AI enthusiasts, and welcome to our 6,908 new readers.