Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples | NVIDIA Technical Blog

Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples | NVIDIA Technical Blog

By Luca Spindler
Publication Date: 2026-10-01 17:59:00

Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems.

Do Inference Now (DIN) Deploy is an open-source collection of practical C++ samples that bridges that gap. It combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to help developers move from a model checkpoint to a native, hardware-accelerated application on Windows and Linux. The same ONNX Runtime API can also be accessed through WinML 2.0.

From ONNX export to C++ implementation

Each DIN Deploy sample starts with a Python exporter that downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact. The application side is a native C++ CLI built on ONNX Runtime (ORT). That split keeps model conversion separate from deployment logic, and developers can take an exported model into a local application without requiring a model-specific runtime.

Most sample code uses ONNX Runtime session and tensor APIs in C++. Vendor-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths. Execution providers that support the required ONNX Runtime tensor APIs can run the shared code.

ORT’s copy tensor API keeps data locality manageable without dedicated vendor API usage in shared code.

For preprocessing and postprocessing around exported model inference, the FLUX.2 sample uses ONNX Runtime’s graphics interop capability, introduced in version 1.25, with…