An Easy Introduction to CUDA C and C++ | NVIDIA Technical Blog
This first post in a series on CUDA C and C++ covers the basic concepts of parallel programming on the CUDA platform with C/C++.
This first post in a series on CUDA C and C++ covers the basic concepts of parallel programming on the CUDA platform with C/C++.
News and tutorials for developers, scientists, and IT admins
Nsight Systems helps you tune and scale software across CPUs and GPUs.
In efficient parallel algorithms, threads cooperate and share data to perform collective computations. To share data, the threads must synchronize.
In our last CUDA C/C++ post we discussed how to transfer data efficiently between the host and device. In this post, we discuss how to overlap data transfers…
NVIDIA Collective Communications Library (NCCL) provides optimized implementation of inter-GPU communication operations, such as allreduce and variants.
This walkthrough summarizes the TREx workflow and highlight API features for examining data and TensorRT engines.
This is the fourth post in the CUDA Refresher series, which has the goal of refreshing key concepts in CUDA, tools, and optimization for beginning or…
There’s a new computational workhorse in town. For decades, general matrix-matrix multiply—known as GEMM in Basic Linear Algebra Subroutines (BLAS) libraries…
A CUDA pro tip about pointer aliasing and how to use the restrict keyword to avoid performance problems due to aliasing in C and C++ code on CPUs and GPUs.