NVIDIA GPU Infrastructure Fundamentals: A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations

NVIDIA GPU Infrastructure Fundamentals: A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations book cover

NVIDIA GPU Infrastructure Fundamentals: A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations

Author(s): Vivian Aranha (Author)

  • Publisher: Packt Publishing
  • Publication Date: August 31, 2026
  • Language: English
  • Print length: 236 pages
  • ISBN-10: 1808087496
  • ISBN-13: 9781808087493

Book Description

Decode the NVIDIA GPU ecosystem in one structured guide. Compare technologies, understand how platform layers interact, and build the judgment to evaluate infrastructure choices and trade-offs.

Key Features

  • Understand how GPUs, CUDA, networking, storage, and DPUs support AI workloads
  • Learn the roles of MIG, vGPU, DCGM, Kubernetes, Slurm, NGC, and Triton
  • Connect infrastructure components across the AI development and deployment lifecycle

Book Description

NVIDIA GPU infrastructure spans hardware, system software, networking, storage, orchestration, MLOps, and inference. Understanding how these components fit together, where their responsibilities overlap, and which distinctions matter requires a clear, structured path.

This book provides that path through one coherent narrative of the NVIDIA GPU infrastructure stack. It covers accelerated computing, CUDA, and the NVIDIA software ecosystem before comparing data center GPUs against workload characteristics. You will examine MIG, vGPU, and DCGM for resource sharing and monitoring; Kubernetes and Slurm for GPU scheduling; and the networking and storage layer, including Ethernet, InfiniBand, RDMA, GPUDirect Storage, BlueField DPUs, and DOCA.

Later chapters connect infrastructure to the AI lifecycle through Airflow, MLflow, and Kubeflow for MLOps, NGC for software delivery, and ONNX, TensorRT, and Triton for inference. You will also explore the Kubernetes components, monitoring technologies, scaling considerations, and diagnostic concepts that support production GPU clusters.

By connecting these technologies instead of presenting them as isolated products, the book helps you compare platform choices, understand component boundaries, discuss trade-offs, and develop a durable mental model of NVIDIA GPU infrastructure.

What you will learn

  • Distinguish AI, machine learning, and deep learning
  • Explain why GPUs accelerate modern AI workloads
  • Match NVIDIA GPUs to training and inference requirements
  • Select MIG or vGPU for common resource-sharing scenarios
  • Compare Ethernet and InfiniBand for distributed AI workloads
  • Map MLOps tools to the right stage of the AI lifecycle
  • Differentiate ONNX, TensorRT, and Triton in inference workflows
  • Trace GPU cluster issues across platform layers

Who this book is for

This book is for system administrators, cloud and DevOps professionals, data center and networking teams, solution architects, technical managers, presales professionals, and beginners who need a clear understanding of NVIDIA GPU infrastructure. It is especially relevant to professionals moving into AI infrastructure roles, evaluating GPU platform technologies, or collaborating across compute, networking, MLOps, and operations teams. Basic familiarity with IT, cloud, or data center concepts is helpful; programming, data science, and previous GPU experience are not required.

Table of Contents

  1. Understanding AI Workloads and Accelerated Computing
  2. CUDA and the NVIDIA AI Software Stack
  3. NVIDIA GPU Architecture and Platform Selection
  4. Sharing, Monitoring, and Scheduling GPU Resources
  5. Building High-Throughput Storage and Network Paths
  6. Virtualized GPU Infrastructure and DPU Offload
  7. Operating the AI Lifecycle with MLOps and Inference
  8. Delivering and Operating NVIDIA AI Platforms
  9. Preparing for the NCA-AIIO Exam and the Next Step

Editorial Reviews

Editorial Reviews

About the Author

Vivian Aranha is an AI educator, technology leader, and founder of School of AI, with over 20 years of industry experience. He earned a Bachelor’s degree in Information Technology in 2004 and a Master’s degree in Computer Science in 2006. His career spans web technologies, mobile app development for iOS and Android, blockchain solutions, and AI systems and applications. Vivian has worked with Fortune 500 organizations, including The Washington Post, Delta Air Lines, and IBM. An instructor since 2009, he has trained professionals worldwide and now teaches AI globally. His courses have attracted over 2.5 million enrollments, with more than 500,000 students learning through School of AI, Udemy, Skool, and Maven.

View on Amazon

未经允许不得转载:Wow! eBook » NVIDIA GPU Infrastructure Fundamentals: A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations