SDSignal Desk

Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples

Oct 1, 2026, 10:59 AM · NVIDIA Developer

Image: NVIDIA Developer

DIN Deploy open-sources C++ ONNX Runtime + TensorRT RTX samples — Whisper to FLUX.2 — aiming at real local apps, not notebook demos.

Why it matters

NVIDIA’s Do Inference Now (DIN) Deploy repo ships C++ samples that export Hugging Face checkpoints to ONNX in Python, then run native CLIs on ONNX Runtime with the TensorRT RTX execution provider on Windows and Linux (including Arm64 presets). Workloads include Whisper, Parakeet TDT, Nemotron ASR Streaming, Meta SAM 2.1, and FLUX.2-klein-4B with Vulkan/DirectX interop via ORT 1.25.

DGX Spark measurements in the post show large GPU speedups — e.g. Parakeet at 206x real-time vs 14x on CPU, SAM 2.1 at 38.3 FPS vs 0.5 on CPU. Quantized FLUX.2 paths claim up to 1.63x vs BF16 with drop-in ONNX.

From the desk

Local AI only wins if developers can ship it inside real applications. DIN Deploy’s split — Python export, C++ runtime — is the right seam. We’re for TensorRT RTX samples that keep shared code on portable ORT APIs and quarantine CUDA to optional paths.

Benchmark tables from the vendor’s Spark box are marketing until others reproduce them, but the direction is correct: ASR, segmentation, and image gen as copy-pasteable native pipelines. Graphics interop for FLUX.2 is especially telling — generative features that meet game/engine memory where it already lives.

I’m watching WinML 2.0 access and whether these samples stay maintained when ORT and TensorRT RTX versions jump. Useful local AI is a packaging problem as much as a model problem.

Context

Fits NVIDIA’s local-AI narrative alongside DGX Spark and RTX Spark PCs — developer enablement for on-device inference.

Who feels it

C++ app developers
A concrete path from HF checkpoint to accelerated desktop/edge binary.
Creators & ISVs
SAM/FLUX samples hint at shipping selection and generation features without a cloud round-trip.
NVIDIA platform
Tightens ORT+TensorRT RTX as the default local stack on RTX/Spark hardware.

What to watch

  1. Community issues/PRs on DIN Deploy after launch
  2. Cross-vendor EP performance vs TensorRT RTX claims
  3. How painful export+quantize remains for non-NVIDIA models

Read the original

Continue at the source.

NVIDIA Developer

Companies: NVIDIA

Also covering this