Skip to main content
Models & Research

Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8

Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

By Precis Daily Newsroom2 min read349 words
Illustration for: Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Mode
Illustration
Key points
  • Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B.
  • Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.
  • Base model is a hybrid Mamba-2 + MoE + Attention architecture with 3B active parameters.

Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights. Base model is a hybrid Mamba-2 + MoE + Attention architecture with 3B active parameters. Supports up to 1M token context, switchable reasoning, tool calling, and six languages. Scores 83.4 on PinchBench agent tasks but lags on coding and hard reasoning benchmarks. Deployable via vLLM, SGLang, or Docker on a single H100 or DGX Spark. Red Hat cuts Nemotron 3.5 Lightning’s weight footprint with FP8 Red Hat AI has released an FP8-quantized checkpoint of NVIDIA’s Nemotron model . Converting selected weights and activations from BF16 to FP8 roughly halves their storage, making the 30-billion-parameter agent model easier to deploy on a single accelerator. Hugging Face displayed more than 780,000 downloads for the checkpoint at the time of publication. That counter measures file downloads rather than unique users or production deployments, but it indicates substantial early interest in a lower-memory version. Red Hat produced the checkpoint with LLM Compressor , applying static per-tensor FP8 quantization to weights and activations in supported linear operators. FP8 stores each quantized value in eight bits, compared with 16 bits for BF16. Several precision-sensitive components remain unquantized: conv1d layers, embeddings, latent projections, mixture-of-experts gates, multi-token prediction layers, the final normalization layer, and the language-model head. Calibration used 512 UltraChat samples at 2,048 tokens each. The resulting savings apply primarily to quantized tensors. Total runtime memory also includes unquantized parameters, activations, Mamba state, attention key-value caches, and vLLM workspace allocations, so overall GPU memory use falls by less than 50% in many serving configurations. Hopper and Blackwell GPUs can execute FP8 matrix multiplication through dedicated Tensor Cores. End-to-end throughput depends on batch size, sequence length, memory bandwidth, expert routing, and the share of execution handled by unquantized operators. FP8 therefore raises the performance ceiling without guaranteeing a twofold increase in request throughput. You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Sources

Summarized from the linked originals.

Related stories

Illustration for: ToolGrad: Efficient tool-use dataset generation with textual
Models & Research

ToolGrad: Efficient tool-use dataset generation with textual "gradients" ToolGrad: Efficient tool-use dataset generation with textual "gradients" Zhongyi Zhou, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR ToolGrad is a data generation framework that reverses the traditional paradigm by first generating tool-use answers before user queries. We show this design enables LLMs to achieve better tool-use performance.

Google Research6 min
Illustration for: SambaRack SN50 Benchmarked on MiniMax M2.7 by SemiAnalysis
Models & Research

SambaRack SN50 Benchmarked on MiniMax M2.7 by SemiAnalysis SemiAnalysis Benchmarks SambaRack SN50 with Fast Inference on MiniMax M2.7 MiniMax M2.7 is a model used by many of our customers around the world that helps augment their coding and agentic workflows using the fast inference speed of SambaNova’s SN40 to accelerate their tasks.

SambaNova Blog5 min
Illustration for: Introducing Muse Spark: Scaling Towards Personal Superintell
Models & Research

Introducing Muse Spark: Scaling Towards Personal Superintelligence Introducing Muse Spark: Scaling Towards Personal Superintelligence Today, we’re excited to introduce Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

Meta AI6 min