Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

- Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B.
- Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.
- Base model is a hybrid Mamba-2 + MoE + Attention architecture with 3B active parameters.
Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights. Base model is a hybrid Mamba-2 + MoE + Attention architecture with 3B active parameters. Supports up to 1M token context, switchable reasoning, tool calling, and six languages. Scores 83.4 on PinchBench agent tasks but lags on coding and hard reasoning benchmarks. Deployable via vLLM, SGLang, or Docker on a single H100 or DGX Spark. Red Hat cuts Nemotron 3.5 Lightning’s weight footprint with FP8 Red Hat AI has released an FP8-quantized checkpoint of NVIDIA’s Nemotron model . Converting selected weights and activations from BF16 to FP8 roughly halves their storage, making the 30-billion-parameter agent model easier to deploy on a single accelerator. Hugging Face displayed more than 780,000 downloads for the checkpoint at the time of publication. That counter measures file downloads rather than unique users or production deployments, but it indicates substantial early interest in a lower-memory version. Red Hat produced the checkpoint with LLM Compressor , applying static per-tensor FP8 quantization to weights and activations in supported linear operators. FP8 stores each quantized value in eight bits, compared with 16 bits for BF16. Several precision-sensitive components remain unquantized: conv1d layers, embeddings, latent projections, mixture-of-experts gates, multi-token prediction layers, the final normalization layer, and the language-model head. Calibration used 512 UltraChat samples at 2,048 tokens each. The resulting savings apply primarily to quantized tensors. Total runtime memory also includes unquantized parameters, activations, Mamba state, attention key-value caches, and vLLM workspace allocations, so overall GPU memory use falls by less than 50% in many serving configurations. Hopper and Blackwell GPUs can execute FP8 matrix multiplication through dedicated Tensor Cores. End-to-end throughput depends on batch size, sequence length, memory bandwidth, expert routing, and the share of execution handled by unquantized operators. FP8 therefore raises the performance ceiling without guaranteeing a twofold increase in request throughput. You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
Sources
Related stories

Claude’s New addTools() Can Reuse 98.7% of Your Next Request. Editing tools[] Reuses None.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI.

ToolGrad: Efficient tool-use dataset generation with textual "gradients"
ToolGrad: Efficient tool-use dataset generation with textual "gradients" ToolGrad: Efficient tool-use dataset generation with textual "gradients" Zhongyi Zhou, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR ToolGrad is a data generation framework that reverses the traditional paradigm by first generating tool-use answers before user queries. We show this design enables LLMs to achieve better tool-use performance.

SambaRack SN50 Benchmarked on MiniMax M2.7 by SemiAnalysis
SambaRack SN50 Benchmarked on MiniMax M2.7 by SemiAnalysis SemiAnalysis Benchmarks SambaRack SN50 with Fast Inference on MiniMax M2.7 MiniMax M2.7 is a model used by many of our customers around the world that helps augment their coding and agentic workflows using the fast inference speed of SambaNova’s SN40 to accelerate their tasks.

Introducing Muse Spark: Scaling Towards Personal Superintelligence - AI at Meta
Introducing Muse Spark: Scaling Towards Personal Superintelligence Introducing Muse Spark: Scaling Towards Personal Superintelligence Today, we’re excited to introduce Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.