AI On-Prem Deployment and Fine-Tuning by Pearson

AI On-Prem Deployment and Fine-Tuning by Pearson

4h 51mIntermediate2026-09-16

Authors

Pearson

Pearson

Course details

Running AI on your own infrastructure gives you control over your data, your costs, and your privacy. In this course, instructor Sander van Vugt draws on decades of open-source experience and shows you how to deploy and manage generative AI models on-premises. Explore the essential components of generative AI and the use cases that open up in private AI workflows. Learn how to host AI systems on platforms like Linux and Kubernetes and choose the environment that best fits your organization. Along the way, examine the architecture of large language models (LLMs) and how to fine-tune them for your needs without expensive resources. By the end of this course, you'll be prepared to host and scale your own AI models in low-connectivity environments while keeping data processing secure.

Learning objectives
Deploy generative AI models on-premises for greater data control and privacy.
Host AI systems on Linux and Kubernetes based on organizational needs.
Fine-tune large language models (LLMs) for specific needs without expensive resources.
Add private data to AI workflows using retrieval augmented generation (RAG) and adapters.
Scale AI models in low-connectivity environments while keeping data processing secure.

Concepts

Introduction

  • Course introduction

Introduction to Generative AI (GenAI)

  • Module 1 - Generative AI (GenAI) fundamentals introduction
  • Learning objectives
  • AI solutions overview
  • Generative AI use cases and features
  • Core components of a GenAI solution
  • Common platforms and ecosystems
  • Machine learning vs. large language models (LLMs)
  • Responsible AI
  • Key skills for hosting GenAI solutions
  • Hardware requirements for this course

Key GenAI Components

  • Learning objectives
  • What is an LLM
  • The inference engine
  • Agents and orchestration
  • Application programming interface (API) exposure and client utilities
  • Lab - Exploring Hugging Face
  • Lab solution - Exploring Hugging Face

Large Language Models (LLMs)

  • Learning objectives
  • Understanding LLM architecture at a high level
  • Choosing the right model
  • LLM families
  • Using Hugging Face
  • Downloading models from Hugging Face
  • Using llama.cpp as a simple inference runtime
  • Using Ollama for inference
  • Lab - Test-drive an LLM with llama.cpp
  • Lab solution - Test-drive an LLM with llama.cpp

Reducing LLM Usage System Requirements

  • Learning objectives
  • GPU or CPU
  • Choosing the right LLM parameters
  • Picking the right inference engine
  • Cold versus warm start
  • Running LLMs in Podman AI Lab
  • Lab - Running LLMs in Podman AI Lab
  • Lab solution - Running LLMs in Podman AI Lab

Hosting Platforms Overview

  • Learning objectives
  • Why it makes sense to host your own AI platform
  • Hosting GenAI on Linux
  • Hosting GenAI on Kubernetes
  • Hosting GenAI on a public cloud
  • Lab - Hosting GenAI on a public cloud
  • Lab solution - Hosting GenAI on a public cloud

Tweaking Linux for AI

  • Module 2 - Hosting GenAI on Linux Introduction
  • Learning objectives
  • Managing GPU drivers on Linux
  • Installing GPU drivers on RHEL
  • Adding GPU support for containers
  • GPU optimization basics
  • Tuning Linux for CPU-only inference
  • Lab - Running llama.cpp on GPU
  • Lab solution - Running llama.cpp on GPU

Running an Inference Server

  • Learning objectives
  • Llama.cpp versus vLLM
  • Running inference servers as containers
  • Requirements for using vLLM
  • Running vLLM
  • Basic configuration and tuning for vLLM
  • Using Open WebUI
  • Lab - Running vLLM
  • Lab solution - Running vLLM

Options for Adding Data to an LLM

  • Module 3 - Adding Data to an LLM Introduction
  • Learning objectives
  • The knowledge problem
  • Prompt-based injection
  • Retrieval-augmented generation (RAG)
  • Adapters
  • Fine-tuning
  • Lab - Running Qwen with an adapter
  • Lab solution - Running Qwen with an adapter

Using Retrieval-Augmented Generation (RAG)

  • Learning objectives
  • Understanding the right RAG setup
  • Running backend services for RAG
  • Configuring Open WebUI for RAG
  • Testing RAG
  • Lab - Adding RAG to a private LLM
  • Lab solution - Adding RAG to a private LLM

Preparing the Kubernetes Platform

  • Module 4 - Hosting GenAI on Kubernetes and Red Hat introduction
  • Learning objectives
  • Kubernetes GenAI architecture overview
  • Kubernetes core resources overview
  • Installing a simple on-premises Kubernetes cluster
  • Deploying GPU workloads with the NVIDIA plugin
  • Lab - Installing Kubernetes with GPU support
  • Lab solution - Installing Kubernetes with GPU support

Running GenAI on Kubernetes

  • Learning objectives
  • Offering persistent storage for GenAI
  • Running inference on Kubernetes
  • Making the inference engine accessible with Gateway API
  • Scaling and monitoring
  • Lab - Offering Kubernetes-based LLM inference
  • Lab solution - Offering Kubernetes-based LLM inference

Red Hat AI Solutions

  • Learning objectives
  • Red Hat AI products overview
  • Trying Red Hat AI Inference Server

Putting It All Together

  • Module 5 - Putting it all together introduction
  • Learning objectives
  • Putting it all together - Building a practical offline-capable AI assistant for low-connectivity environments
  • Case study - Preparing the platform
  • Case study - Getting the right LLM
  • Case study - Running the inference engine
  • Case study - Connecting the Open WebUI client
  • Case study - Configuring Open WebUI for RAG
  • Wrap-up with best practices

Conclusion

  • Course summary
100,000 Toman