Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

- Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B.
- Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.
- Base model is a hybrid Mamba-2 + MoE + Attention architecture with 3B active parameters.
Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B.
Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.
Base model is a hybrid Mamba-2 + MoE + Attention architecture with 3B active parameters.
Supports up to 1M token context, switchable reasoning, tool calling, and six languages.
Scores 83.4 on PinchBench agent tasks but lags on coding and hard reasoning benchmarks.
Deployable via vLLM, SGLang, or Docker on a single H100 or DGX Spark.
Red Hat cuts Nemotron 3.5 Lightning’s weight footprint with FP8
Red Hat AI has released an FP8-quantized checkpoint of NVIDIA’s Nemotron model . Converting selected weights and activations from BF16 to FP8 roughly halves their storage, making the 30-billion-parameter agent model easier to deploy on a single accelerator.
Hugging Face displayed more than 780,000 downloads for the checkpoint at the time of publication. That counter measures file downloads rather than unique users or production deployments, but it indicates substantial early interest in a lower-memory version.
Red Hat produced the checkpoint with LLM Compressor , applying static per-tensor FP8 quantization to weights and activations in supported linear operators. FP8 stores each quantized value in eight bits, compared with 16 bits for BF16.
Several precision-sensitive components remain unquantized: conv1d layers, embeddings, latent projections, mixture-of-experts gates, multi-token prediction layers, the final normalization layer, and the language-model head. Calibration used 512 UltraChat samples at 2,048 tokens each.
The resulting savings apply primarily to quantized tensors. Total runtime memory also includes unquantized parameters, activations, Mamba state, attention key-value caches, and vLLM workspace allocations, so overall GPU memory use falls by less than 50% in many serving configurations.
Hopper and Blackwell GPUs can execute FP8 matrix multiplication through dedicated Tensor Cores. End-to-end throughput depends on batch size, sequence length, memory bandwidth, expert routing, and the share of execution handled by unquantized operators. FP8 therefore raises the performance ceiling without guaranteeing a twofold increase in request throughput.
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
Sources
Related stories

As AI Grows More Complex, Model Builders Rely on NVIDIA
Unveiling what it describes as the most capable model series yet for professional knowledge work, OpenAI launched GPT-5.2 in December.

What Is Jev? A Guide to TypeSafe AI’s System One Model
What Is Jev? A Guide to TypeSafe AI’s System One Model Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results, and then continues in that loop until the task is complete.
![Illustration for: Claude’s New addTools() Can Reuse 98.7% of Your Next Request. Editing tools[] Reuses None.](/images/articles/claude--s-new-addtools-can-reuse-98-7-of-your-next-request-editing.webp)
Claude’s New addTools() Can Reuse 98.7% of Your Next Request. Editing tools[] Reuses None.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI.

ToolGrad: Efficient tool-use dataset generation with textual "gradients"
Zhongyi Zhou, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR ToolGrad is a data generation framework that reverses the traditional paradigm by first generating tool-use answers before user queries. We show this design enables LLMs to achieve.

SambaRack SN50 Benchmarked on MiniMax M2.7 by SemiAnalysis
SemiAnalysis Benchmarks SambaRack SN50 with Fast Inference on MiniMax M2.7 MiniMax M2.7 is a model used by many of our customers around the world that helps augment their coding and agentic workflows using the fast inference speed of SambaNova’s SN40 to accelerate their tasks..

Introducing Muse Spark: Scaling Towards Personal Superintelligence - AI at Meta
Introducing Muse Spark: Scaling Towards Personal Superintelligence Introducing Muse Spark: Scaling Towards Personal Superintelligence Today, we’re excited to introduce Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a.

NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to fit on more devices, giving builders more to run locally.

How to scale agentic applications without creating AI sprawl
As agentic applications become more autonomous, enterprises need shared infrastructure for context, model, and tool access; governance; evaluation; and observability, rather than rebuilding these capabilities for every agent.