Search results for “site:nvidia.com”

Page 16 of about 551 results

developer.nvidia.com blog › calculating-video-quality-using-nvidia-gpus-and-vmaf-cuda

Calculating Video Quality Using NVIDIA GPUs and VMAF-CUDA | NVIDIA Technical Blog

Video quality metrics are used to evaluate the fidelity of video content. They provide a consistent quantitative measurement to assess the performance of the…

developer.nvidia.com blog › how-access-global-memory-efficiently-cuda-fortran-kernels

How to Access Global Memory Efficiently in CUDA Fortran Kernels | NVIDIA Technical Blog

In the previous two posts we looked at how to move data efficiently between the host and device. In this sixth post of our CUDA Fortran series we discuss how to…

developer.nvidia.com blog › dynamic-control-flow-in-cuda-graphs-with-conditional-nodes

Dynamic Control Flow in CUDA Graphs with Conditional Nodes | NVIDIA Technical Blog

Post updated on February 3, 2025 with details about CUDA 12.8. CUDA Graphs can provide a significant performance increase, as the driver is able to optimize…

developer.nvidia.com blog › accelerating-leaderboard-topping-asr-models-10x-with-nvidia-nemo

Accelerating Leaderboard-Topping ASR Models 10x with NVIDIA NeMo | NVIDIA Technical Blog

NVIDIA NeMo has consistently developed automatic speech recognition (ASR) models that set the benchmark in the industry, particularly those topping the Hugging…

developer.nvidia.com blog › building-the-modular-foundation-for-ai-factories-with-nvidia-mgx

Building the Modular Foundation for AI Factories with NVIDIA MGX | NVIDIA Technical Blog

The exponential growth of generative AI, large language models (LLMs), and high-performance computing has created unprecedented demands on data center…

developer.nvidia.com blog › cuda-pro-tip-always-set-current-device-avoid-multithreading-bugs

CUDA Pro Tip: Always Set the Current Device to Avoid Multithreading Bugs | NVIDIA Technical Blog

A simple rule to avoid multithreading bugs in applications that run in parallel on multiple GPUs.

developer.nvidia.com blog › improving-gpu-performance-by-reducing-instruction-cache-misses-2

Improving GPU Performance by Reducing Instruction Cache Misses | NVIDIA Technical Blog

GPUs are specially designed to crunch through massive amounts of data at high speed. They have a large amount of compute resources…

developer.nvidia.com blog › using-nsight-compute-nvprof-mixed-precision-deep-learning-models

Using Nsight Compute or Nvprof to Show Mixed Precision Use in Deep Learning Models | NVIDIA Technical Blog

Mixed precision combines different numerical precisions in a computational method. The Volta and Turing generation of GPUs introduced Tensor Cores…

developer.nvidia.com blog › advanced-nvidia-cuda-kernel-optimization-techniques-handwritten-ptx

Advanced NVIDIA CUDA Kernel Optimization Techniques: Handwritten PTX | NVIDIA Technical Blog

As accelerated computing continues to drive application performance in all areas of AI and scientific computing, there’s a renewed interest in GPU optimization…

developer.nvidia.com blog › blackwell-breaks-the-1000-tps-user-barrier-with-metas-llama-4-maverick

Blackwell Breaks the 1,000 TPS/User Barrier With Meta’s Llama 4 Maverick | NVIDIA Technical Blog

NVIDIA has achieved a world-record large language model (LLM) inference speed. A single NVIDIA DGX B200 node with eight NVIDIA Blackwell GPUs can achieve over 1…

Try “site:nvidia.com” on: Marginalia · Mojeek · Wiby · DuckDuckGo · Bing · Google · Wikipedia · Internet Archive