---
format: "aidr-story-markdown/v1"
id: "7f6b9d99dce00408b1e0627d6b38df3b3423d4da25596f39ba812fac39fa8b4a"
canonical_url: "https://aidr.today/7f6b9d99?lang=en"
title: "Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-01T17:59:27.000Z"
category: "Infra"
topics: ["nvidia","inference","dgx-spark","64gb","local-ai","chips","infra"]
source_urls: ["https://developer.nvidia.com/blog/build-local-ai-apps-with-c-and-nvidia-tensorrt-rtx-samples/","https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/"]
summary: "Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems. Do Inference Now (DIN) Deploy is an open-source collection of practical C++ samples that bridges that gap. It combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to help developers move from a model checkpoint to a native, hardware-accelerated application on Windows and Linux. The same ONNX Runtime API can also be accessed through WinML 2.0 . Each DIN Deploy sample starts with a Python exporter that downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact. The application side is a native C++ CLI built on ONNX Runtime (ORT). That split keeps model conversion separate from deployment logic, and developers can take an exported model into a local application without requiring a model-specific runtime. Most sample code uses ONNX Runtime session and tensor APIs in C++. Vendor-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths. Execution providers that support the required ONNX Runtime tensor APIs can run the shared code. ORT’s copy tensor API keeps data locali…"
---

# Build Local AI Apps with C\+\+ and NVIDIA TensorRT RTX Samples

> [Open the canonical story](<https://aidr.today/7f6b9d99?lang=en>)

**Published:** 2026-10-01T17:59:27.000Z
**Category:** Infra
**Topics:** nvidia, inference, dgx\-spark, 64gb, local\-ai, chips, infra

## Summary

Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems\. Do Inference Now \(DIN\) Deploy is an open\-source collection of practical C\+\+ samples that bridges that gap\. It combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to help developers move from a model checkpoint to a native, hardware\-accelerated application on Windows and Linux\. The same ONNX Runtime API can also be accessed through WinML 2\.0 \. Each DIN Deploy sample starts with a Python exporter that downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact\. The application side is a native C\+\+ CLI built on ONNX Runtime \(ORT\)\. That split keeps model conversion separate from deployment logic, and developers can take an exported model into a local application without requiring a model\-specific runtime\. Most sample code uses ONNX Runtime session and tensor APIs in C\+\+\. Vendor\-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths\. Execution providers that support the required ONNX Runtime tensor APIs can run the shared code\. ORT’s copy tensor API keeps data locali…

## Sources

- [Story source](<https://developer.nvidia.com/blog/build-local-ai-apps-with-c-and-nvidia-tensorrt-rtx-samples/>)
- [Story source](<https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/>)

