NVIDIA Certified Associate AI Infrastructure and Operations (NCA-AIIO) Cert Prep
4h 26mBeginner2026-03-16
Authors

Packt Publishing
Course details
Get ready for the NVIDIA Certified Associate AI Infrastructure and Operations certification with this comprehensive certification prep course. Explore the fundamentals of AI and gain insights into the workings of an AI-centric data center. Learn about the NVIDIA technology stack, including GPU cores, CUDA programming, and advanced networking with InfiniBand and RDMA. Build your understanding of AI workflows, model training, and inference, and master operational and management strategies through tools like NVIDIA SMI and DCGM. Enhance your understanding of AI-centric solutions and gain practical skills applicable to global-scale AI deployments. Tailored for data center technicians, DevOps engineers, IT managers, and solution architects, this course creates a solid foundation of knowledge and skills for people preparing to take the NVIDIA certification exam.
Concepts
Introduction
- Introduction
Certification Details
- NVIDIA certifications
- Topics covered in certification
Module 1 Fundamentals
- Drivers of AI evolution
- AI use cases across industries
- AI, ML, DL, Gen AI
- Analogy for AI, ML, DL, Gen AI
- Transformer model
Module 2 Inside an AI-Centric Data Center
- Inside an AI-centric data center
- Power usage effectiveness (PUE)
- The compute power
- CPU and GPU
- CPU vs. GPU - Architectural difference
- Beyond Moore's Law
- Data processing unit (DPU)
- Network inside an AI-centric data center
- Network fabric
- Ethernet vs. InfiniBand
- Converged Ethernet (CE)
- Storage inside an AI-centric data center
- Cloud vs. on-prem
Module 3 NVIDIA Technology Stack
- NVIDIA - Powering AI GPU innovation
- NVIDIA technology stack
- Layer 1 - Physical layer
- GPU on a graphics card
- DGX platform
- DGX SuperPOD
- ConnectX
- BlueField DPUs
- NVIDIA reference architectures
- Understanding GPU cores
- Comparing GPU cores
- NVIDIA DGX platform - Timeline
- DGX platform - Deployment options
- DGX A100 vs. H100
- Layer 2 - Data movement and I O acceleration
- NVLink
- InfiniBand
- InfiniBand vs. Ethernet
- DMA and RDMA
- GPUDirect RDMA
- GPUDirect storage
- Layer 4 - Core libraries
- Compute unified device architecture (CUDA)
- Installing CUDA
- NVIDIA collective communications library (NCCL)
- NVLink, NVSwitch, PCIe, RDMA vs. NCCL
- Layer 5 - Monitoring and management
- NVIDIA-SMI
- Data Center GPU Manager (DCGM)
- Base Command Manager
- Which one to use
- Layer 6 - Applications and vertical solutions
- Summary
- NVIDIA AI Enterprise
- NVIDIA AI Factory
- Quick comparison
- Layer 3 - OS, driver, and virtualization
- GPU drivers
- GPU virtualization
- vGPU vs. MIG, part 1
- vGPU vs. MIG, part 2
Module 4 AI Workflows
- AI workflows
- ML frameworks
- The NVIDIA differentiator
- Model training vs. model inference
- Job scheduling vs. container orchestration
- Slurm vs. Kubernetes
- NVIDIA integration
- ML Ops - Analogy
- Why ML Ops
- NVIDIA tools supporting ML Ops