Search results for “site:developer.nvidia.com”

Page 9 of about 219 results

developer.nvidia.com blog › advanced-nvidia-cuda-kernel-optimization-techniques-handwritten-ptx

Advanced NVIDIA CUDA Kernel Optimization Techniques: Handwritten PTX | NVIDIA Technical Blog

As accelerated computing continues to drive application performance in all areas of AI and scientific computing, there’s a renewed interest in GPU optimization…

developer.nvidia.com blog › blackwell-breaks-the-1000-tps-user-barrier-with-metas-llama-4-maverick

Blackwell Breaks the 1,000 TPS/User Barrier With Meta’s Llama 4 Maverick | NVIDIA Technical Blog

NVIDIA has achieved a world-record large language model (LLM) inference speed. A single NVIDIA DGX B200 node with eight NVIDIA Blackwell GPUs can achieve over 1…

developer.nvidia.com blog › accelerating-vector-search-nvidia-cuvs-ivf-pq-performance-tuning-part-2

Accelerating Vector Search: NVIDIA cuVS IVF-PQ Part 2, Performance Tuning | NVIDIA Technical Blog

In the first part of the series, we presented an overview of the IVF-PQ algorithm and explained how it builds on top of the IVF-Flat algorithm…

developer.nvidia.com blog › announcing-nsight-compute-2020-2-and-nsight-visual-studio-edition-2020-2

Announcing Nsight Compute 2020.2 and Nsight Visual Studio Edition 2020.2 | NVIDIA Technical Blog

The latest release adds the highly-requested “Application Replay” feature and additional collection knobs to give you more control over what data you collect…

developer.nvidia.com blog › nvidia-cupynumeric-25-03-now-fully-open-source-with-pip-and-hdf5-support

NVIDIA cuPyNumeric 25.03 Now Fully Open Source with PIP and HDF5 Support | NVIDIA Technical Blog

NVIDIA cuPyNumeric is a library that aims to provide a distributed and accelerated drop-in replacement for NumPy built on top of the Legate framework.

developer.nvidia.com blog › develop-custom-physical-ai-foundation-models-with-nvidia-cosmos-predict-2

Develop Custom Physical AI Foundation Models with NVIDIA Cosmos Predict-2 | NVIDIA Technical Blog

Building smarter robots and autonomous vehicles (AVs) starts with physical AI models that understand real-world dynamics. These models serve two critical roles…

developer.nvidia.com blog › r2d2-building-ai-based-3d-robot-perception-and-mapping-with-nvidia-research

R²D²: Building AI-based 3D Robot Perception and Mapping with NVIDIA Research | NVIDIA Technical Blog

Robots must perceive and interpret their 3D environments to act safely and effectively. This is especially critical for tasks such as autonomous navigation…

developer.nvidia.com blog › building-an-ai-agent-for-supply-chain-optimization-with-nvidia-nim-and-cuopt

Building an AI Agent for Supply Chain Optimization with NVIDIA NIM and cuOpt | NVIDIA Technical Blog

Enterprises face significant challenges in making supply chain decisions that maximize profits while adapting quickly to dynamic changes.

developer.nvidia.com blog › simplifying-gpu-application-development-with-heterogeneous-memory-management

Simplifying GPU Application Development with Heterogeneous Memory Management | NVIDIA Technical Blog

Heterogeneous Memory Management (HMM) is a CUDA memory management feature that improves programmer productivity for all programming models built on top of CUDA.

developer.nvidia.com blog › nvidia-omniverse-what-developers-need-to-know-about-migration-away-from-launcher

NVIDIA Omniverse: What Developers Need to Know About Migration Away From Launcher | NVIDIA Technical Blog

As part of continued efforts to ensure NVIDIA Omniverse is a developer-first platform, NVIDIA will be deprecating the Omniverse Launcher on Oct. 1.