Search results for “site:developer.nvidia.com”

Page 8 of about 215 results

developer.nvidia.com blog › dynamic-control-flow-in-cuda-graphs-with-conditional-nodes

Dynamic Control Flow in CUDA Graphs with Conditional Nodes | NVIDIA Technical Blog

Post updated on February 3, 2025 with details about CUDA 12.8. CUDA Graphs can provide a significant performance increase, as the driver is able to optimize…

developer.nvidia.com blog › accelerating-leaderboard-topping-asr-models-10x-with-nvidia-nemo

Accelerating Leaderboard-Topping ASR Models 10x with NVIDIA NeMo | NVIDIA Technical Blog

NVIDIA NeMo has consistently developed automatic speech recognition (ASR) models that set the benchmark in the industry, particularly those topping the Hugging…

developer.nvidia.com blog › building-the-modular-foundation-for-ai-factories-with-nvidia-mgx

Building the Modular Foundation for AI Factories with NVIDIA MGX | NVIDIA Technical Blog

The exponential growth of generative AI, large language models (LLMs), and high-performance computing has created unprecedented demands on data center…

developer.nvidia.com blog › cuda-pro-tip-always-set-current-device-avoid-multithreading-bugs

CUDA Pro Tip: Always Set the Current Device to Avoid Multithreading Bugs | NVIDIA Technical Blog

A simple rule to avoid multithreading bugs in applications that run in parallel on multiple GPUs.

developer.nvidia.com blog › improving-gpu-performance-by-reducing-instruction-cache-misses-2

Improving GPU Performance by Reducing Instruction Cache Misses | NVIDIA Technical Blog

GPUs are specially designed to crunch through massive amounts of data at high speed. They have a large amount of compute resources…

developer.nvidia.com blog › using-nsight-compute-nvprof-mixed-precision-deep-learning-models

Using Nsight Compute or Nvprof to Show Mixed Precision Use in Deep Learning Models | NVIDIA Technical Blog

Mixed precision combines different numerical precisions in a computational method. The Volta and Turing generation of GPUs introduced Tensor Cores…

developer.nvidia.com blog › advanced-nvidia-cuda-kernel-optimization-techniques-handwritten-ptx

Advanced NVIDIA CUDA Kernel Optimization Techniques: Handwritten PTX | NVIDIA Technical Blog

As accelerated computing continues to drive application performance in all areas of AI and scientific computing, there’s a renewed interest in GPU optimization…

developer.nvidia.com blog › blackwell-breaks-the-1000-tps-user-barrier-with-metas-llama-4-maverick

Blackwell Breaks the 1,000 TPS/User Barrier With Meta’s Llama 4 Maverick | NVIDIA Technical Blog

NVIDIA has achieved a world-record large language model (LLM) inference speed. A single NVIDIA DGX B200 node with eight NVIDIA Blackwell GPUs can achieve over 1…

developer.nvidia.com blog › accelerating-vector-search-nvidia-cuvs-ivf-pq-performance-tuning-part-2

Accelerating Vector Search: NVIDIA cuVS IVF-PQ Part 2, Performance Tuning | NVIDIA Technical Blog

In the first part of the series, we presented an overview of the IVF-PQ algorithm and explained how it builds on top of the IVF-Flat algorithm…

developer.nvidia.com blog › announcing-nsight-compute-2020-2-and-nsight-visual-studio-edition-2020-2

Announcing Nsight Compute 2020.2 and Nsight Visual Studio Edition 2020.2 | NVIDIA Technical Blog

The latest release adds the highly-requested “Application Replay” feature and additional collection knobs to give you more control over what data you collect…