Dynamic Control Flow in CUDA Graphs with Conditional Nodes | NVIDIA Technical Blog
Post updated on February 3, 2025 with details about CUDA 12.8. CUDA Graphs can provide a significant performance increase, as the driver is able to optimize…
Post updated on February 3, 2025 with details about CUDA 12.8. CUDA Graphs can provide a significant performance increase, as the driver is able to optimize…
NVIDIA NeMo has consistently developed automatic speech recognition (ASR) models that set the benchmark in the industry, particularly those topping the Hugging…
The exponential growth of generative AI, large language models (LLMs), and high-performance computing has created unprecedented demands on data center…
A simple rule to avoid multithreading bugs in applications that run in parallel on multiple GPUs.
GPUs are specially designed to crunch through massive amounts of data at high speed. They have a large amount of compute resources…
Mixed precision combines different numerical precisions in a computational method. The Volta and Turing generation of GPUs introduced Tensor Cores…
As accelerated computing continues to drive application performance in all areas of AI and scientific computing, there’s a renewed interest in GPU optimization…
NVIDIA has achieved a world-record large language model (LLM) inference speed. A single NVIDIA DGX B200 node with eight NVIDIA Blackwell GPUs can achieve over 1…
In the first part of the series, we presented an overview of the IVF-PQ algorithm and explained how it builds on top of the IVF-Flat algorithm…
The latest release adds the highly-requested “Application Replay” feature and additional collection knobs to give you more control over what data you collect…