Search results for “site:developer.nvidia.com”

Page 12 of about 219 results

developer.nvidia.com blog › how-overlap-data-transfers-cuda-cc

How to Overlap Data Transfers in CUDA C/C++ | NVIDIA Technical Blog

In our last CUDA C/C++ post we discussed how to transfer data efficiently between the host and device. In this post, we discuss how to overlap data transfers…

developer.nvidia.com blog › scaling-deep-learning-training-nccl

Scaling Deep Learning Training with NCCL | NVIDIA Technical Blog

NVIDIA Collective Communications Library (NCCL) provides optimized implementation of inter-GPU communication operations, such as allreduce and variants.

developer.nvidia.com blog › cuda-refresher-cuda-programming-model

CUDA Refresher: The CUDA Programming Model | NVIDIA Technical Blog

This is the fourth post in the CUDA Refresher series, which has the goal of refreshing key concepts in CUDA, tools, and optimization for beginning or…

developer.nvidia.com blog › cublas-strided-batched-matrix-multiply

Pro Tip: cuBLAS Strided Batched Matrix Multiply | NVIDIA Technical Blog

There’s a new computational workhorse in town. For decades, general matrix-matrix multiply—known as GEMM in Basic Linear Algebra Subroutines (BLAS) libraries…

developer.nvidia.com blog › cuda-pro-tip-optimize-pointer-aliasing

CUDA Pro Tip: Optimize for Pointer Aliasing | NVIDIA Technical Blog

A CUDA pro tip about pointer aliasing and how to use the restrict keyword to avoid performance problems due to aliasing in C and C++ code on CPUs and GPUs.