Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples
Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples | NVIDIA Technical Blog Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples Do Inference Now (DIN) Deploy provides open-source C++ samples that combine ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux.

- Each sample separates model export from deployment by using a Python exporter to convert Hugging Face checkpoints into ONNX artifacts, then runs inference through a native C++ CLI built on ONNX Runtime session and tensor APIs.
- The repository supports automatic speech recognition with OpenAI Whisper, NVIDIA Parakeet TDT, and NVIDIA Nemotron ASR Streaming, as well as interactive segmentation with Meta SAM 2.1 and prompt-driven image generation with FLUX.2-klein-4B.
- Performance measurements on DGX Spark show GPU acceleration ranging from 39x real-time for Nemotron ASR streaming to 206x for Parakeet TDT, while SAM 2.1 achieves 38.3 FPS on GPU versus 0.5 FPS on CPU.
| NVIDIA Technical Blog Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples Do Inference Now (DIN) Deploy provides open-source C++ samples that combine ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux. Each sample separates model export from deployment by using a Python exporter to convert Hugging Face checkpoints into ONNX artifacts, then runs inference through a native C++ CLI built on ONNX Runtime session and tensor APIs. The repository supports automatic speech recognition with OpenAI Whisper, NVIDIA Parakeet TDT, and NVIDIA Nemotron ASR Streaming, as well as interactive segmentation with Meta SAM 2.1 and prompt-driven image generation with FLUX.2-klein-4B. Performance measurements on DGX Spark show GPU acceleration ranging from 39x real-time for Nemotron ASR streaming to 206x for Parakeet TDT, while SAM 2.1 achieves 38.3 FPS on GPU versus 0.5 FPS on CPU. The FLUX.2 sample demonstrates graphics interop using Vulkan and DirectX with ONNX Runtime 1.25 capabilities, and shows how post-training quantization with NVIDIA Model Optimizer produces a drop-in ONNX replacement requiring no application-code changes. Explore the DIN Deploy repository to access CMake presets and sample implementations. Learn more about TensorRT for RTX for hardware-accelerated inference on RTX GPUs. Review the NVIDIA blog series on model quantization to understand optimization techniques used in the samples. Powered by NVIDIA Nemotron. AI-generated content may summarize information incompletely. Verify important information. Learn more Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems. Do Inference Now (DIN) Deploy is an open-source collection of practical C++ samples that bridges that gap. It combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to help developers move from a model checkpoint to a native, hardware-accelerated application on Windows and Linux. The same ONNX Runtime API can also be accessed through WinML 2.0 . Each DIN Deploy sample starts with a Python exporter that downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact. The application side is a native C++ CLI built on ONNX Runtime (ORT). That split keeps model conversion separate from deployment logic, and developers can take an exported model into a local application without requiring a model-specific runtime. Most sample code uses ONNX Runtime session and tensor APIs in C++. Vendor-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths. Execution providers that support the required ONNX Runtime tensor APIs can run the shared code. ORT’s copy tensor API keeps data locality manageable without dedicated vendor API usage in shared code. For preprocessing and postprocessing around exported model inference, the FLUX.2 sample uses ONNX Runtime’s graphics interop capability , introduced in version 1.25, with Vulkan and DirectX for sampling. The repository provides CMake presets for Windows and Linux, including Arm64 variants. DirectX is available only on Windows. For automatic speech recognition (ASR), DIN Deploy supports offline and streaming pipelines. OpenAI Whisper covers offline transcription across model sizes, while NVIDIA Parakeet TDT and NVIDIA Nemotron ASR Streaming provide streaming pipelines. The samples show how to move audio through a native application and return transcription results while using GPU acceleration where it is available. Meta SAM 2.1 samples support interactive masking for images and video. They turn model outputs into segmentation masks that native applications can use for selection, tracking, and other computer-vision workflows. Table 1 compares GPU and CPU performance for selected DIN Deploy workloads measured on DGX Spark. Model GPU (DGX Spark) CPU (DGX Spark) openai/whisper-large-v3-turbo 58.5x 3.8x nvidia/nemotron-3.5-asr-streaming-0.6b 39.01x 3.24x nvidia/parakeet-tdt-0.6b-v3 206.41x 14.44x facebook/sam2.1-hiera-base-plus 38.3 FPS 0.5 FPS Table 1. DGX Spark GPU and CPU performance for selected DIN Deploy workloads. Results for audio are reported as multiples of real time (higher is faster) The FLUX.2-klein-4B sample provides prompt-driven image generation. The sample includes graphics-API interops with Vulkan and DirectX, allowing applications to integrate GPU-resident resources with a cross-vendor shader interface. It also shows how post-training quantization (PTQ) with NVIDIA Model Optimizer produces a quantized ONNX model. Quantization is hardware-dependent, but due to ONNX interfaces remaining unchanged, the quantized model is a drop-in replacement requiring no application-code changes. Figure 1. Quantization can drastically accelerate models on recent tensor core hardware Start with the repository’s CMake presets for Windows, Linux, x86-64, and Arm64. After configuring and building the project, export a model to ONNX and run the CLI with TensorRT RTX, or copy the code into your own application and use the pipeline implementations. CMake downloads ONNX Runtime and TensorRT RTX by default. Learn more about TensorRT for RTX , NVIDIA Local AI , and the DIN Deploy repository , and the NVIDIA blog series on model quantization . Developer Tools & Techniques | General | TensorRT | Intermediate Technical | Benchmark | C++ | CUDA | DGX Spark | featured Luca is a developer technology engineer for professional visualization at NVIDIA. With a passion for deep learning and computer vision, he helps partners leverage NVIDIA technologies. Previously, Luca studied computational engineering at RWTH Aachen University in Germany. Maximilian Müller is a developer technology engineer for Professional Visualization at NVIDIA. His passion for computer vision and deep learning helps external partners and internal teams to best leverage the performance of NVIDIA accelerators using CUDA. Before joining NVIDIA, he acquired an MSc in electrical and computational engineering at RWTH Aachen.
Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.