StudioTV course
AI/ML Architect Course
This course focuses on architecture, scaling, reliability, and platform choices, with an "architect" style that explains both why and how. AI (Artificial Intelligence) and ML (Machine Learning) are used here in a broad sense: designing, training, and serving models at scale.
Course objectives
Objectives:
- Understand the full ML value chain and its dependencies.
- Design large-scale training and inference architectures that remain operable.
- Reason about cost, performance, complexity, and risk with concrete examples.
- Be able to explain the "why", not only the "how", in clear language.
How to read this course
Guidance: start with Section 0 if you need a refresher, then move to Sections 2-4 for the core architecture decisions.
Read each section with the same lens: constraints, trade-offs, and operability. Every recommendation is meant to be defendable in production, not just correct on paper.
Note: a global glossary is available at the end of the document (Section 10) for acronyms and key concepts.
Table of contents
- 0. Essential ML refresher
- 1. Overview and principles
- 2. Distributed Training
- 3. Orchestration: Kubernetes vs Slurm (crucial)
- 4. MLOps: from POC to production
- 5. Hugging Face and the ML ecosystem
- 6. Infrastructure as Code & DevOps
- 7. Tools and ecosystem (integration)
- 8. Platform layer and go-to-prod
- 9. Future: HPC/AI trends
- 10. Global glossary
0. Essential ML refresher
This refresher sets the minimum foundation before diving into distributed training, inference, and industrialization. The goal is to clarify vocabulary, intuitions, and classic pitfalls so the rest of the course is smooth.
0.1 Dataset, labels, and features
A dataset is a set of examples (observations). Each example contains features (input variables) and, in supervised learning, a label (the ground truth to predict: class, score, value). Features can be numeric, categorical, text, or images; labels can be binary, multi-class, multi-label, or continuous.
Simple example: for fraud detection, features are amount, country, time, card type; the label is 0/1 (fraud or not). For an LLM, the main feature is the token sequence, and the label is often the next token.
Watch-outs:
- Quality: noise, missing values, inconsistent labels; a "dirty" dataset destroys performance.
- Bias: a non-representative dataset produces a model that is unusable in production.
- Imbalance: if the rare class is 1%, you need suitable metrics (AUC/precision-recall).
0.2 Loss function (e.g., cross-entropy)
The loss function is the error function you minimize during training. The goal is not to "classify well" directly, but to minimize a continuous measure that guides learning.
Common examples:
- Classification: cross-entropy (measures the gap between predicted and true distributions).
- Regression: MSE/MAE (difference between prediction and true value).
You often add regularization (L2, dropout) to limit overfitting. The idea is simple: a very low training loss does not guarantee a good model in production.
0.3 Gradient descent (intuition, no heavy math)
Gradient descent is an optimization mechanism in small steps. At each iteration, you compute how the loss changes with respect to parameters, then update weights in the direction that reduces the loss.
For a full walkthrough, we also have a dedicated page on the site:https://studiotvai.com/gradient-descent
Intuition:
- The learning rate controls step size: too large = unstable, too small = slow convergence.
- You often use mini-batches: a trade-off between gradient accuracy and speed.
- Variants (SGD, Adam, AdamW) improve stability and learning speed.
Repeated over thousands of iterations, this progressively learns a function that generalizes.
0.4 Overfitting / underfitting
Two classic failures:
- Overfitting: the model "memorizes" the training set but fails on validation/test.
- Underfitting: the model is too simple and fails everywhere.
Typical signals:
- Overfitting = training loss goes down, validation loss goes back up.
- Underfitting = both losses stay high.
Levers:
- Overfitting: more data, regularization, early stopping, data augmentation, smaller model.
- Underfitting: richer model, more informative features, more training.
0.5 Train/validation/test split
You split data to measure generalization correctly:
- Train: learn parameters.
- Validation: choose hyperparameters (learning rate, architecture).
- Test: final evaluation, used once.
Golden rule: never "peek" at test during optimization. For time series or production-like settings, you often do a time-based split (train on the past, test on the future) to avoid data leakage.
Offline metrics vs Online metrics
Offline: metrics on a fixed dataset (AUC, accuracy, BLEU, perplexity). Online: real impact in production (CTR, churn, revenue, cost/req, latency).
Concrete example: a model improves offline AUC but worsens production CTR because it favors content that "ranks well" but converts poorly. Conclusion: optimization must be aligned with business metrics.
Data leakage: examples + how to detect it
Examples: using a post-event timestamp, a score computed after a decision, or a feature that is derived from the label. Simple detection: strict time split, remove suspicious features, and a "time-travel" test (is the feature available at prediction time?).
0.6 Why we need GPUs
Modern models rely on massive matrix multiplications and convolutions. GPUs were designed for this massive parallelism and have very high memory bandwidth.
Concretely:
- A GPU runs thousands of operations in parallel.
- VRAM stores weights and activations to avoid slow transfers to CPU.
- Without GPUs, training large models can take weeks instead of hours/days.
Edge case: for small tabular models, CPU is enough. As soon as matrices get large or data volume explodes, GPU becomes essential.
0.7 Survival mini-glossary (10-15 terms)
Objectives:
- Provide the minimal vocabulary needed for the next sections.
- Clarify common student confusions.
Key points:
- A term = an operational idea, not abstract jargon.
- Production meaning matters more than an academic definition.
Concrete example: "Epoch" = one full pass over the dataset; if training runs 5 epochs, it has seen each example 5 times (on average).
Essential terms:
- Dataset: set of examples used to train/evaluate.
- Feature: input variable used by the model.
- Label: ground truth to predict.
- Batch: a set of examples processed together.
- Epoch: one full pass over the dataset.
- Learning rate: optimization step size.
- Loss: error function to minimize.
- Baseline: simple reference model.
- Overfitting: good on train, poor generalization.
- Drift: shift in data distributions or user behavior.
- A/B test: controlled comparison in production.
- Cold start: lack of history for a request/model.
- SLO: measurable Service Level Objective.
1. Overview and principles
An industrial ML system is never "just a model". It combines data, a training pipeline, a deployment system, and monitoring that feeds back into data. Without these building blocks, the model remains a proof of concept and does not become a reliable service.
Data -> Features -> Training -> Validation -> Registry -> Inference -> Monitoring -> Feedback loop
Guiding principles:
- Reproducibility and auditability are required to understand what runs in prod and retrain quickly.
- Metrics (latency, cost per request, error rate) must be defined before optimization.
- Inference is often the main cost center, so optimization frequently starts there.
- Risk management requires canary rollouts, shadow traffic, and rollback plans.
- Technical choices are always a trade-off between cost, performance, and complexity.
What we measure
| Category | Examples |
|---|---|
| Model metrics | accuracy, AUC, BLEU, perplexity |
| System metrics | throughput, latency P95, GPU util |
| Business metrics | CTR, churn, revenue, cost / req |
Warning: a "better" model can be worse in production.
In practice, you will see fine-tuning (GPU, Graphics Processing Unit), training LLMs (Large Language Models) with network/latency/fault-tolerance constraints, batch inference (throughput-oriented), and real-time inference (P95/P99-driven). Each workload comes with different priorities: training optimizes global throughput, real-time inference optimizes latency and stability, batch inference optimizes cost per prediction.
Key concept - Throughput vs Latency
Throughput: amount of work per unit of time (tokens/s, samples/s).
Latency: individual response time (P95/P99 for inference). -> Training almost always optimizes throughput; real-time inference optimizes latency.
Quick examples: a real-time recommender ingests a stream, computes online features, and serves low-latency predictions; a batch scoring job takes a nightly snapshot, runs batch inference, and exports scores to a CRM (Customer Relationship Management).
Key concept - One-off training vs continuous inference
One-off training: scheduled batch executions, driven by throughput, cost, and iteration speed.
Continuous inference: 24/7 service, driven by latency, availability, and SLOs.
2. Distributed Training
At scale, the main challenge is not raw compute, but communication, memory, and failure recovery. This drives the choice of parallelism, network topology, and stability trade-offs. These techniques (DDP, ZeRO, TP/PP) are primarily training techniques; for inference, you usually focus on serving engines, replication, sharding, and quantization.
Key sentence
"At scale, the challenge is less about training itself and more about communication patterns, memory pressure, and failure recovery."
2.1 Core concepts (must-know)
Sharding (simple definition)
Sharding means splitting a large object (parameters, gradients, activations, or data) into pieces, then distributing those pieces across multiple GPUs or nodes. The goal is to reduce memory per GPU and enable scaling. Concretely, a weight matrix can be partitioned by columns or rows; each GPU stores a slice and computes part of the result, then a collective communication (all-reduce/all-gather) reconstructs the full activation. Sharding shows up in ZeRO (sharding optimizer states/gradients/parameters), in tensor parallelism (sharding matrices), and in data parallelism (sharding the data batch).
Data Parallelism
In data parallelism, each GPU has a full copy of the model and processes a different micro-batch. After backward, gradients are aggregated with an all-reduce to compute the global average, then each GPU applies the same update. All-reduce is a collective operation that sums (or averages) gradients across all GPUs and returns the result to each GPU; it is often implemented as a ring to maximize bandwidth. NCCL (NVIDIA Collective Communications Library) implements these GPU collectives (all-reduce, all-gather, broadcast) and optimizes data movement over NVLink, PCIe (PCI Express), or InfiniBand.
Key points: the bottleneck is communication. As soon as all-reduce time exceeds compute time, scaling degrades. Global batch size is batch_local * num_gpus, so you often adjust the learning rate with warmup. Warmup means starting with a smaller learning rate and ramping it up over the first iterations; it stabilizes training with large global batches and prevents overly aggressive updates. If memory is the limit, gradient accumulation increases global batch size without exceeding VRAM.
Key concept - VRAM vs RAM vs NVMe
VRAM: GPU memory, very fast but limited, used directly for weights and activations.
RAM: CPU memory, larger but slower, useful for offload and buffering.
NVMe: persistent storage, very large but much slower, useful for offload and checkpoints.
Typical use cases: vision or NLP (Natural Language Processing) when the model fits in VRAM, fine-tuning on 8-64 GPUs, or when you want throughput without the complexity of sharding. Concrete example: ResNet50 on 8xA100, batch_local=256 -> batch_global=2048; add warmup and optionally gradient accumulation if memory is limited.
Common frameworks include PyTorch DDP, Horovod, and JAX pmap for data parallelism.
PyTorch DDP example (simplified):
import os
import torch
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP
def setup():
dist.init_process_group(backend="nccl")
local_rank = int(os.environ["LOCAL_RANK"])
torch.cuda.set_device(local_rank)
return local_rank
local_rank = setup()
model = MyModel().cuda()
model = DDP(model, device_ids=[local_rank])
for batch in dataloader:
loss = model(batch)
loss.backward()
optimizer.step()
optimizer.zero_grad()Network Topology Awareness
In distributed training, network topology matters as much as the number of GPUs. Topology-aware placement prefers intra-node communication (NVLink/PCIe) and reduces inter-node traffic; 8 GPUs in one node share a fast interconnect, whereas 8 GPUs across 8 nodes pay slower network all-reduces. Cross-rack communication adds even more latency and can become the real bottleneck. In practice, you want to place jobs on "close" nodes (same rack/leaf) to reduce latency variance and stabilize throughput.
Model Parallelism
Model parallelism splits the model across GPUs; each GPU stores only part of the parameters and computes part of the activations. This is useful when the model does not fit on a single GPU, but it is harder to orchestrate because you must synchronize inside layers.
Tensor parallelism is a form of model parallelism where the matrices of a layer are sharded, i.e., partitioned into chunks across several GPUs (by columns or rows). Each GPU computes a fraction of the matrix product, then an all-reduce or all-gather reconstructs the full activation. This is common for attention layers (QKV, Query/Key/Value) and MLPs (Multi-Layer Perceptron).
Expert parallelism (MoE, Mixture of Experts) distributes experts across GPUs; a router sends each token to a subset of experts, which requires good load balancing.
Typical use cases: a 70B LLM on 8 GPUs with NVLink, large embeddings, or MLP layers that exceed VRAM. When a single layer does not fit on one GPU, model parallelism becomes mandatory even with small batch sizes. Concrete example: TP=4 on an 8-GPU node means each layer is split into 4 parts and computed in parallel on 4 GPUs. For example, an MLP projection of 16384 x 65536 is split into 4 blocks; each GPU computes one quarter of the product, then an all-reduce/all-gather reconstructs the full activation. This makes the layer fit in memory, but adds synchronization at every layer, making latency sensitive to bandwidth.
Common frameworks include Megatron-LM (tensor parallel), DeepSpeed (tensor/expert parallel), and FairScale. On the inference side, vLLM uses tensor parallelism to serve an LLM across multiple GPUs.
Pipeline Parallelism
Pipeline parallelism splits the model into consecutive stages; each stage runs on a GPU or node, and a micro-batch moves from stage to stage like an assembly line. The classic problem is bubble inefficiency: at the beginning and the end, some stages sit idle, reducing efficiency. You mitigate it by increasing the number of micro-batches and interleaving multiple batches.
Used in Megatron-LM / DeepSpeed, especially for very deep models or when per-layer memory is limited. Concrete example: a CNN (Convolutional Neural Network) for detection can be split as: GPU0 = stem + backbone blocks 1-2, GPU1 = backbone blocks 3-4, GPU2 = head (FPN, Feature Pyramid Network + detection). A batch is split into 8 micro-batches; while GPU0 processes micro-batch 2, GPU1 processes micro-batch 1 and GPU2 processes micro-batch 0, gradually filling the pipeline. For a 48-layer transformer, you can define 4 stages of 12 layers, which helps fit memory while amortizing communication.
Common frameworks include DeepSpeed Pipeline Parallel and Megatron-LM, with similar variants in FairScale.
Key concept - Data vs Model vs Pipeline Parallelism
Data parallelism: replicate the model and shard the batch; simple to scale if the model fits in VRAM.
Model parallelism: split the model (tensor/expert), useful when a layer does not fit on one GPU.
Pipeline parallelism: split into stages; efficient with micro-batches but sensitive to bubbles.
Hybrid Parallelism
Hybrid parallelism combines data, tensor, and pipeline parallelism to scale LLMs beyond 10B parameters. Tensor parallelism shards matrices inside layers, pipeline parallelism splits layers into stages, and data parallelism replicates those groups to increase throughput. We use PP (pipeline parallelism), TP (tensor parallelism), and DP (data parallelism) as shorthand.
Concrete example: 64 GPUs = PP=2, TP=4, DP=8. Each DP group sees the same data batch, while the model is sharded across PP * TP. A simple rule is: total GPUs = PP * TP * DP, which helps plan topology.
Common frameworks include Megatron-LM to combine PP+TP+DP, often integrated with DeepSpeed (ZeRO) to reduce memory.
2.2 Distributed training frameworks
PyTorch Distributed
PyTorch Distributed relies on torch.distributed and uses NCCL as the GPU backend. DistributedDataParallel (DDP) launches one process per GPU and synchronizes gradients via all-reduce. DataParallel is the older "simple" mode that splits a batch across multiple GPUs inside a single process; it is convenient for a quick prototype, but it does not scale as well. DDP is preferred over DataParallel because it avoids the GIL, enables more efficient communication, and works correctly in multi-node setups. The GIL (Global Interpreter Lock) is a Python lock that prevents multiple Python threads from executing bytecode in parallel within a single process, which limits performance when you try to parallelize with threads.
All-reduce typically works by splitting gradients and circulating them in a ring; each GPU adds received chunks and ends up with the final average. Increasing global batch size improves throughput, but you must adapt the learning rate (LR scaling rule + warmup) and use gradient accumulation if memory is tight. Throughput is the number of examples or tokens processed per second; it is the core metric for measuring efficiency in training or in an inference server.
Use case: DDP for stable multi-node or multi-GPU training; DataParallel is only useful for a single-node prototype.
Typical launch:
torchrun --nproc_per_node=8 --nnodes=2 --node_rank=0 train.pyElastic training (node failures)
torch.distributed.elastic allows you to survive a node loss and resume the job without restarting everything. It is very useful on spot/preemptible clusters, or for long trainings that must tolerate interruptions.
DeepSpeed
DeepSpeed introduces ZeRO (Zero Redundancy Optimizer) to reduce memory: ZeRO-1 shards optimizer state, ZeRO-2 shards gradients, and ZeRO-3 shards parameters. CPU/NVMe offload allows you to go beyond VRAM, while activation checkpointing reduces memory at the cost of recomputation. ZeRO is a training mechanism, not an inference technique; for inference you more often use tensor/pipeline parallelism, sharding, or quantization.
Example config (simplified):
{
"train_batch_size": 1024,
"zero_optimization": {
"stage": 3,
"offload_param": { "device": "cpu" },
"offload_optimizer": { "device": "cpu" }
},
"activation_checkpointing": {
"partition_activations": true
},
"bf16": { "enabled": true }
}Use case: 13B or 70B LLMs on a VRAM-limited cluster, or when you must exceed GPU memory at the cost of bandwidth. Typical question: "When would you use ZeRO-3 vs model parallelism?" Expected answer: ZeRO-3 when the model fits in total GPU memory but not per GPU and bandwidth is sufficient; model parallelism when some layers are too large even after sharding, or when you need to split compute within a layer. In practice, very large LLMs combine ZeRO-3 + Tensor Parallelism + Pipeline.
JAX
JAX provides pmap for simple data parallelism and pjit for more flexible sharding (tensor/model). XLA (Accelerated Linear Algebra) compilation fuses operations and stabilizes performance, which explains its popularity in research and on TPU (Tensor Processing Unit). You do not need to be a JAX expert, but you should be able to explain why some customers choose it.
Minimal example:
import jax
import jax.numpy as jnp
@jax.pmap
def step(x):
return x * 2
y = step(jnp.ones((8, 4)))2.3 Use cases and parallelism examples
The table below gives a rough idea of practical choices depending on model sizes and memory constraints.
| Scenario | Parallelism | Why |
|---|---|---|
| Vision, 8 GPUs | DDP | Model fits in VRAM, scales fast |
| 13B LLM on 8 GPUs | DDP + ZeRO-2 | Tight memory, shard gradients |
| 70B LLM on 8 GPUs | TP + ZeRO-3 | Layers too large per GPU |
| 175B LLM on 64 GPUs | PP + TP + DP | Scalability and memory |
| MoE 64 experts | Expert Parallel + DP | Expert routing |
2.4 Failure modes & incidents
Common failure modes
- Communication: NCCL deadlock, timeouts, or mismatched tensor sizes in an all-reduce.
- Memory: OOM, VRAM fragmentation, leaks that degrade throughput.
- Storage: corrupted/partial checkpoint, unstable storage latency.
- Scheduling: straggler GPU, noisy neighbor, non topology-aware placement.
- Execution: job dies at 80% (preemption, outage, bug).
Business impact
- Days of GPU time wasted and cost burned with no model gain.
- SLOs not met (latency or throughput), direct user impact.
- Production incidents, loss of trust, delayed delivery.
Mitigations
- Frequent checkpoints, integrity validation, multiple versions.
- GPU-level monitoring and alerts on collectives (NCCL timeouts).
- Topology-aware scheduling and isolation of noisy nodes.
- Explicit timeouts and automatic restart (elastic/restart).
2.5 Practical rules (distributed training)
Golden rule - Distributed Training
As long as the model fits on a single GPU, start with DDP. Introduce ZeRO or Tensor Parallelism only when you have a real memory constraint. (ZeRO is a training technique; for inference, prefer sharding/quantization first.)
Useful reminders:
- Measure communication vs compute first; if communication dominates, reduce parallelism or tune micro-batches.
- If scaling stalls, inspect network topology and inter-node placement.
- Document reproducibility (versions, seeds, checkpoints) before "optimizing".
Milestone
Now that the parallelism building blocks are set, we can formalize a decision method to pick the right combination given constraints.
2.6 Mental models & decision method
Objectives:
- Build a fast framework for ML architecture decisions in production.
- Separate data, model, system, and business constraints clearly.
- Speed up diagnosis when a signal (cost, latency, quality) degrades.
Key points:
- Any architecture decision is guided by a trio: SLO (Service Level Objective, a measurable quality-of-service target) + cost + data.
- Bottlenecks change with scale: 1 GPU = compute, multi-GPU = communication, production = data + reliability.
- A "perfect" offline model is useless if pipeline and latency break the business promise.
ML system diagram (architect view):
User -> API -> Feature Store -> Model -> Business Logic -> Output
^ |
| v
Data Pipeline --------> Monitoring/Feedback -> TrainingConcrete example: A real-time anti-fraud engine serves 200 req/s. Real-time features come from a feature store, the model outputs a score, then a business rule decides (block / request verification). If P95 latency exceeds 150 ms, fraud slips through in production: the SLO constraint drives architecture (dedicated GPU, caching, compact model).
How to choose (decision trees + tables):
Decision tree - Training
Does the model fit on 1 GPU?
├─ Yes -> DDP (simple, stable, good scaling)
└─ No -> Real memory constraints?
├─ Yes -> ZeRO (1/2/3) + offload if needed
└─ No -> Model Parallelism (TP/PP) if a layer is too big
Then: do you need to scale > 32 GPUs?
├─ Yes -> Hybrid (DP + TP + PP)
└─ No -> DDP/ZeRO is often enoughQuick table (signals -> likely choice):
| Signal | Likely choice |
|---|---|
| Model fits in VRAM, simple training | DDP |
| OOM, per-GPU VRAM limit | ZeRO-2/3 |
| A layer does not fit | Tensor Parallelism |
| Very deep model | Pipeline Parallelism |
| Massive multi-node scale | Hybrid (DP+TP+PP) |
Decision tree - Inference
Strict latency SLO? ├─ Yes -> Real-time, dedicated GPU, cautious batching └─ No -> Batch, shared GPU (MIG/MPS) if needed High QPS? ├─ Yes -> Caching + batching + autoscaling └─ No -> Keep it simple, minimize cost Model too large for the latency target? ├─ Yes -> Quantization / distillation / model sharding └─ No -> Standard model
Decision checklist (fast):
- What is the target latency (P95/P99) and the cost budget per request?
- Are features available in prod with the same freshness?
- Does the model fit in VRAM without offload?
- Is the bottleneck compute, communication, or data?
3. Orchestration: Kubernetes vs Slurm (crucial)
Kubernetes and Slurm coexist in most AI/ML environments. The first is a service platform; the second is an HPC scheduler (High Performance Computing). Being able to argue the choice is often a decisive interview point.
| Dimension | Kubernetes | Slurm |
|---|---|---|
| Primary use | Services, inference, MLOps | Massive HPC training |
| Multi-tenant | Excellent | Limited |
| Autoscaling | Native (HPA/KEDA) | Limited |
| GPU topology control | Less fine-grained | Very controlled |
| Long batch jobs | Possible but more complex | Natural |
The core difference is philosophy: Kubernetes optimizes lifecycle management, rolling updates, and autoscaling; Slurm optimizes performance, exclusive allocation, and predictability. Kubernetes is ideal for variable traffic and multi-tenancy, while Slurm is ideal for stable, long-running jobs with strict GPU topology.
3.1 Kubernetes (ML-platform centric)
Kubernetes is well-suited for inference, MLOps pipelines, and multi-tenant environments. It naturally handles horizontal scaling, rolling updates, and isolation via namespaces and quotas.
Must-know: how to configure GPU scheduling (via the NVIDIA device plugin), GPU node pools, affinity/taints/tolerations, and the difference between Jobs and Deployments. Helm is the standard packaging layer to parameterize ML deployments.
Key sentence
"Kubernetes excels for lifecycle management and multi-tenant inference, but it's not always optimal for tightly coupled HPC training workloads."
GPU scheduling makes GPUs visible as schedulable resources (nvidia.com/gpu) through the device plugin, then lets the scheduler place pods based on capacity and constraints. You combine GPU limits with node selectors and taints/tolerations to ensure GPU workloads land on the right nodes.
GPU node pools are groups of dedicated nodes with a given GPU type and configuration (drivers, CUDA, disk size, network). They isolate GPU workloads, control costs, and let you scale per GPU type (e.g., an A100 pool for training and an L4 pool for inference).
For autoscaling, you commonly see HPA (Horizontal Pod Autoscaler) and KEDA (Kubernetes Event-driven Autoscaling) to adjust pod count based on load.
Recommended Kubernetes stack for LLMs (inference + ZeRO training)
Objectives:
- Clarify the required building blocks to serve LLMs on Kubernetes.
- Separate the inference stack (vLLM, TGI, Triton) from the training stack (DeepSpeed/ZeRO).
Key points:
- Inference ≠ training: vLLM/TGI/Triton serve the model, while ZeRO is a training technique (DeepSpeed) to reduce memory; you do not use it for inference.
- GPU foundation: NVIDIA GPU Operator (drivers + device plugin + DCGM), CUDA-compatible base images, and container runtime.
- Scheduling: Node Feature Discovery, Topology Manager, MIG/MPS for GPU sharing, and Kueue/Volcano for batch queueing.
- Distributed training: PyTorch + DeepSpeed (ZeRO) + NCCL + RDMA/InfiniBand network for multi-node.
- Serving: vLLM for LLMs (throughput + KV cache), TGI for Hugging Face, Triton for multi-model serving.
- MLOps: Kubeflow Training Operator (PyTorchJob) or MPI Operator to launch distributed jobs.
Concrete example (minimal stack):
- Install NVIDIA GPU Operator + Node Feature Discovery.
- Deploy a vLLM server for inference (Deployment) and a DeepSpeed PyTorchJob for ZeRO training.
- Add Kueue/Volcano for queueing and Prometheus/Grafana for GPU metrics.
When to pick it: real-time inference with traffic spikes and autoscaling, multi-team ML platform with isolation and quotas, or a full MLOps stack with pipelines, model registry, and monitoring.
Concrete examples: production recommender API with QPS-based autoscaling; LLM serving via vLLM with horizontal scaling; pipeline orchestration via Kubeflow Pipelines and experiment tracking.
GPU Deployment example:
apiVersion: apps/v1
kind: Deployment
metadata:
name: model-server
spec:
replicas: 2
selector:
matchLabels:
app: model-server
template:
metadata:
labels:
app: model-server
spec:
containers:
- name: app
image: my-ml-image:latest
resources:
limits:
nvidia.com/gpu: 1Job example (batch):
apiVersion: batch/v1
kind: Job
spec:
template:
spec:
containers:
- name: trainer
image: my-ml-image:latest
resources:
limits:
nvidia.com/gpu: 4
restartPolicy: Never3.2 Slurm (HPC centric)
Slurm is designed for massive multi-node training, long batch jobs, and controlled HPC environments. It offers fine-grained control of resources, which is critical for GPU topology and network performance.
You should be able to explain: sbatch submits a job, srun launches tasks, GPU partitions segment resources, and exclusive allocation guarantees performance. Fault tolerance is limited (restart), and scaling is less dynamic than in Kubernetes.
Key sentence
"Some customers still prefer Slurm for large-scale training due to predictable performance and simpler GPU topology control."
When to pick it: multi-node LLM training with long and stable jobs, strict GPU/network topology control, or dedicated HPC clusters with experienced users. Concrete examples: LLM pretraining on 64/128 GPUs with InfiniBand; scientific training (climate, bio) with heavy batch jobs; hyperparameter sweeps on dedicated partitions.
Slurm example:
#SBATCH --job-name=llm-train
#SBATCH --nodes=4
#SBATCH --gres=gpu:8
#SBATCH --ntasks-per-node=8
#SBATCH --exclusive
srun python train.py3.3 Expert point: hybrid architectures
A strong argument is to use Kubernetes for inference and the MLOps layer (Kubeflow Pipelines, Katib, monitoring), and Slurm for massive multi-node training. A modern alternative is Kubernetes + Kueue/Volcano for HPC batch jobs, with Kubeflow for pipelines. The right choice depends on SLOs, total cost, and available ops skills. An SLO (Service Level Objective) is a measurable quality-of-service target, e.g., a maximum P95 latency, a minimum throughput, or an acceptable error rate.
Hybrid architecture example:
Slurm -> Massive LLM training -> Model registry K8s -> Inference + Feature store + Monitoring
3.4 Concrete use cases
- Multi-tenant B2B SaaS: Kubernetes serves inference with QPS-based autoscaling, while Slurm runs monthly training on a large dataset.
- Research lab: Slurm handles multi-node training with strict GPU topology; Kubernetes hosts supporting services (MLflow, Jupyter, data APIs).
- Enterprise unification: Kubernetes + Kubeflow Pipelines + Volcano/Kueue for batch training, and Kubernetes for inference (vLLM), CI/CD, and monitoring.
3.5 Practical rules
Golden rule - Orchestration
If the workload is service-oriented (inference, MLOps, multi-tenant), Kubernetes is the default choice. If training is HPC, long-running, and tightly coupled, Slurm remains the most predictable.
Useful reminders:
- Kubernetes for inference + MLOps, Slurm for massive training; combine when needed.
- Check network topology before adding more nodes or going multi-rack.
- Inference SLOs (latency) outweigh cost optimization when you have end users.
4. MLOps: from POC to production
4.1 Realistic pipeline and industrialization
MLOps (Machine Learning Operations) is the set of practices that industrialize a model from a POC (Proof of Concept) to a reliable production system. A realistic enterprise pipeline generally follows this flow and must be versioned, observable, and secured:
Data ingestion -> Feature engineering -> Distributed training -> Validation -> Packaging -> Deployment -> Monitoring -> Feedback
In practice, data ingestion must handle schema, quality, data contracts, and lineage (traceability). Feature engineering must cover batch vs streaming, feature store, and versioning. Orchestration (Kubeflow Pipelines / Argo) structures steps, metadata, and tracking. Distributed training must produce logs, checkpoints, and restart mechanisms. Validation must combine offline metrics, fairness, and drift detection. Packaging goes through a model registry and versioned images. Deployment must support canary, shadow, and rollback. Finally, monitoring tracks latency, cost per request, and drift.
Simple run logging example (MLflow):
import mlflow
with mlflow.start_run():
mlflow.log_param("lr", 3e-4)
mlflow.log_metric("auc", 0.91)
mlflow.log_artifact("model.pt")Concrete use case (churn): ingest weekly from a data lake, run distributed training on 8 GPUs, validate AUC (Area Under the Curve) and bias, export to ONNX (Open Neural Network Exchange), push to the model registry, canary-deploy to 10% of traffic, and monitor drift over 30 days.
Pipeline YAML example (pseudo):
stages:
- ingest
- validate
- train_distributed
- evaluate
- package
- deploy
- monitor4.2 Large-scale inference
Large-scale inference starts with the distinction between batch and real-time. In batch, you optimize throughput and unit cost; in real-time, you optimize latency and reliability. Note: techniques like ZeRO are designed for training, not for inference; in production serving you talk more about quantization, sharding, and inference engines.
Key concept - P95 / P99
P95: 95% of requests are faster than this value.
P99: 99% of requests are faster than this value. -> They capture tail latency, which is critical for real-time inference.
You should also cover GPU sharing (MIG, Multi-Instance GPU; MPS, Multi-Process Service), model sharding, cold start (model loading), autoscaling based on QPS (queries per second) and latency, and caching/batching to amortize latency.
Frameworks to name: Triton Inference Server, TorchServe, vLLM, and Hugging Face TGI (Text Generation Inference). vLLM is especially valued because it provides PagedAttention for KV (key/value) cache management, continuous batching to maximize throughput, and simple deployment on Kubernetes via Deployment/Helm.
Minimal vLLM on Kubernetes:
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-server
spec:
replicas: 1
selector:
matchLabels:
app: vllm-server
template:
metadata:
labels:
app: vllm-server
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
args: ["--model", "meta-llama/Llama-2-7b-hf", "--tensor-parallel-size", "1"]
resources:
limits:
nvidia.com/gpu: 1Concrete case: multi-node vLLM (8 GPUs, 2 nodes)
Objectives:
- Explain why a single model over 8 GPUs needs a distributed orchestrator.
- Provide a simple recipe for "real" multi-node on Kubernetes.
Key points:
- Kubernetes cannot split one pod across multiple nodes: a pod is single-node.
- For a single model on 8 GPUs, you need a distributed backend (Ray is the simplest with vLLM).
- On Kubernetes, KubeRay (Ray Operator) simplifies deploying and managing the Ray head + Ray workers.
- Blueprint: 1 Ray head + 2 Ray workers (4 GPUs each), then vLLM starts in distributed mode.
Concrete recipe (principle):
- 1 Ray head (often CPU-only) + 2 Ray workers (4 GPUs each).
- Start vLLM with:
--distributed-executor-backend=rayand--tensor-parallel-size=8.
Mini-diagram:
Node A (4 GPU) -> Ray worker Node B (4 GPU) -> Ray worker Ray head -> orchestration vLLM -> TP=8 across 2 nodes
Key sentence
"Inference optimization is often more critical than training cost in production environments."
HPA example (Horizontal Pod Autoscaler, metric-based autoscaling):
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: model-server
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: model-server
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: qps
target:
type: AverageValue
averageValue: "100"Concrete example: inference sizing (12 blades, 4x A100 40GB)
Objectives:
- Connect LLM model size to required inference techniques (sharding, quantization, replication).
- Show how to leverage 12 GPU blades with 4x A100 40GB each.
Key points:
- 40GB of VRAM is never "100% usable": leave headroom for runtime, KV cache, and fragmentation. We also have a dedicated VRAM calculator on the site:https://studiotvai.com/llm-vram-calculator
- The 4 GPUs on the same blade communicate faster (NVLink/PCIe) than inter-node; prefer intra-node placement when possible.
- In inference, scaling is often done via replication (more copies) rather than intra-model parallelism, unless the model does not fit in VRAM.
Quick assumptions (order of magnitude):
- BF16/FP16 ≈ 2 bytes/parameter -> 30GB ≈ 15B, 80GB ≈ 40B, 350GB ≈ 175B parameters.
- VRAM "available" for weights: ~32-35GB per GPU (the rest is runtime + KV cache).
- KV cache can consume 6-12GB per GPU depending on context length and load.
Quick table (inference only, no fine-tuning):
| Model | Size | Main technique | GPUs / replica | Recommended placement |
|---|---|---|---|---|
| ~30GB | ~15B BF16 | Simple replication | 1 GPU | 1 GPU per replica |
| ~80GB | ~40B BF16 | TP=4 (or INT8) | 4 GPUs | 1 blade per replica |
| ~350GB | ~175B BF16 | TP=8 + PP=2 | 16 GPUs | 4 blades per replica |
Details per case:
- ~30GB model (e.g., Llama-13B BF16): 1 GPU is enough, but keep headroom for KV cache. On one blade (4 GPUs), you can serve 4 replicas. On 12 blades (48 GPUs), up to 48 replicas to maximize throughput. If you need long context (8k/16k), reduce batch size or switch to INT8 to free memory. vLLM example:
--tensor-parallel-size 1+--max-model-len 4096/8192. - ~80GB model (e.g., 40B BF16): in practice TP=4 on the same blade (4 GPUs) is stable because 80GB / 4 ≈ 20GB per GPU, leaving room for KV cache. This yields at most 12 replicas (1 per blade). Variant: INT8 (~40GB) to go back to 1 GPU, but with stronger constraints on context and batching; TP=2 is viable only if the model is quantized or if context/batch are very limited.
- ~350GB model (e.g., 175B BF16): multi-node sharding is mandatory. A realistic plan is TP=8 + PP=2 (16 GPUs, 4 blades): each pipeline stage has 8 GPUs (2 blades) and memory is distributed cleanly. With 48 GPUs, you get 3 replicas (3x16 GPUs) for throughput. For latency, keep micro-batches small and use topology-aware placement (avoid cross-rack).
vLLM configuration example:
vllm serve meta-llama/Llama-2-70b-hf --tensor-parallel-size 4 --pipeline-parallel-size 1 --max-model-len 4096
4.3 Inference use cases and patterns
- Marketing scoring: nightly batch run with Spark + batch inference. Prioritize throughput and cost per prediction; latency is not critical.
- Real-time fraud detection: strict P99 latency and reliability. Typical stack: Kubernetes + Triton + cache; focus on tail latency, load shedding, and fast fallbacks.
- Customer support LLM: token streaming with vLLM + QPS autoscaling. Prioritize steady throughput, fast cold starts, and predictable cost per request.
- Multi-model serving: multiple versions in production with a model registry and routing logic. Prioritize safe rollbacks, A/B testing, and quality gates.
4.4 Practical rules (large-scale inference)
Golden rule - Large-scale inference
In real time, P95/P99 latency matters more than cost; in batch, cost per prediction matters more than latency.
Useful reminders:
- Start simple (Triton/TorchServe), then optimize (batching, caching, quantization) based on metrics.
- Size autoscaling based on QPS and latency, not only CPU.
- Watch cold starts and multi-GPU models (sharding, KV cache).
4.5 Data pitfalls in production
Objectives:
- Understand why most ML incidents come from data.
- Learn to diagnose drift or pipeline breakages quickly.
Key points:
- A robust model does not compensate for fragile data.
- Schema and freshness failures are often silent.
Concrete example: A churn model trained on complete data is deployed in prod where 15% of features arrive late. Result: performance drops without any model change.
Most production ML incidents come from data, not from the model. You can have an excellent offline model and still hurt production if data is biased, incomplete, or inconsistent.
Data leakage / feature leakage The model learns an information signal that is not available in production (future leakage, indirect labels, post-event variables). Result: overly optimistic offline metrics and a model that fails in prod. Classic example: using a "closure date" field to predict cancellation. Mitigations: strict time split, feature review, leakage tests, and validation with a business owner.
Train/prod skew (distribution shift) Distributions change between training and production: new segments, seasonality, behavioral drift, different data sources. Performance drops with no obvious alert. Mitigations: distribution monitoring (stats, histograms), drift alerts, and retraining triggered by thresholds.
Late data Events arrive late (daily batch, collection latency) while the model is expected to predict in real-time. You train with "future" data compared to what is actually available in prod, creating an illusion of performance. Mitigations: watermarking, consistent time windows, and "time-travel safe" features.
Breaking schema change / schema evolution A field changes type or meaning, or disappears, and the model keeps scoring with null or incorrect values. Bugs are silent and costly. Mitigations: schema registry, versioning, data contracts, and backward-compatibility tests.
Data quality Missing values, duplicates, outliers, and noisy labels create systemic errors. Mitigations: validation rules, acceptance thresholds, and quality dashboards (null rates, ranges, ratios).
Data contracts A data contract defines schema, semantics, frequency, SLAs, and ownership. Without contracts, every upstream change breaks downstream. With contracts, you block or version changes and avoid regressions.
Backfill / replay Recomputing features or replaying pipelines is essential after a bug or change, but can introduce inconsistencies if you mix pipeline versions. Mitigations: version features, store training datasets, and isolate backfills from production runs.
4.6 Security & governance
Objectives:
- Protect data, models, and services in multi-tenant environments.
- Guarantee traceability (code + data + config) and compliance.
Key points:
- RBAC/IAM: fine-grained access control by role, service accounts, least privilege.
- Secrets management: never keep secrets in plaintext; use a vault/secret manager.
- Encryption: at-rest and in-transit encryption, with key rotation.
- Audit/logs: traceability of access, deployments, and served models.
- Model governance: model cards, lineage, versioning, dev->prod approvals.
Concrete example: A credit-scoring model must be auditable. Each deployment references a versioned artifact, the training dataset (hash), and a model card. An audit must be able to reconstruct exactly which version served a given customer decision.
Mini "go-live security" checklist (10 items):
- RBAC/IAM in place with minimal roles.
- Secrets managed by a vault/secret manager.
- Data encrypted at rest and in transit.
- Centralized access/admin logs.
- Documented data access policies (PII).
- Multi-tenant isolation (namespace, quotas, network).
- Model registry with versioning and approvals.
- Lineage: code + data + config tied to each version.
- Canary/shadow validated before global rollout.
- Incident runbook + rollback tested.
5. Hugging Face and the ML ecosystem
Hugging Face is a widely used open-source ecosystem for ML, best known for its model hub and libraries that make it easy to share, fine-tune, and deploy models. In practice, it standardizes how models, datasets, and tokenizers are packaged and consumed, which is why it has become the default entry point for many LLM workflows.
LoRA example:
LoRA (Low-Rank Adaptation) fine-tunes a model by adding small trainable matrices to selected layers. This keeps most weights frozen, reducing VRAM and training time while still adapting to a new domain.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig, get_peft_model
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-2-7b-hf")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-hf")
config = LoraConfig(
r=8,
lora_alpha=32,
target_modules=["q_proj", "v_proj"]
)
model = get_peft_model(model, config)4-bit quantization example (concept):
Quantization stores weights in fewer bits (e.g., 4-bit) to reduce memory and speed up inference. It can slightly reduce quality, so it is typically used when VRAM or latency is the primary constraint.
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-7b-hf",
load_in_4bit=True
)5.1 Hugging Face use cases and patterns
Typical use cases: domain adaptation (legal, finance, support) via LoRA, private model hub for governance and audit, massive datasets via streaming from S3 (Simple Storage Service).
Datasets example (streaming):
from datasets import load_dataset
ds = load_dataset("json", data_files="s3://bucket/train.jsonl", streaming=True)Tokenizer example (add tokens):
tokenizer.add_tokens(["<SKU>", "<PRODUCT>"])6. Infrastructure as Code & DevOps
Terraform enables reproducible provisioning of GPU clusters, networking, and storage, with a centralized state for audit and rollback; it is a cloud-focused IaC (Infrastructure as Code) tool. It is not the equivalent of BlueBanquise: Terraform primarily targets cloud IaC, while BlueBanquise is an HPC toolkit for bare-metal deployment (OS, Operating System, networking, provisioning, and Slurm integration). Docker provides reproducible ML environments, notably via versioned CUDA (Compute Unified Device Architecture) images and multi-stage builds. Git and CI/CD (Continuous Integration / Continuous Delivery) handle pipeline versioning and dev -> staging -> prod promotion, with data, model, and infrastructure tests.
Terraform example (simplified):
module "gpu_cluster" {
source = "./modules/gpu-cluster"
node_count = 8
gpu_type = "A100"
network_cidr = "10.0.0.0/16"
}Minimal Dockerfile example:
FROM nvidia/cuda:12.1.0-cudnn8-runtime-ubuntu22.04 as base
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
FROM base
COPY . .
CMD ["python", "serve.py"]Key point
"Reproducibility is a core concern for enterprise AI."
6.1 DevOps use cases and checklists
Use cases: a regulated environment must reproduce a training run 6 months later; a multi-team organization needs IaC to standardize GPU/network/storage; frequent deployments require reliable CI/CD for models and infrastructure.
Quick checklist: pinned versions (CUDA, drivers, libraries), Docker images tagged by commit, centralized Terraform state, and full observability (logs, metrics, traces).
CI pipeline example (simplified):
stages:
- lint
- unit_tests
- data_checks
- build_image
- deploy_staging
- canary_prod7. Tools and ecosystem (integration)
The goal is not to recite a tool list, but to understand where each tool fits in an architecture and which trade-offs it imposes.
7.1 Programming languages (Python, Go, Java, C++)
Python
Python is the default for data work, training, experimentation, and MLOps scripting. It is fast to iterate and has a massive ecosystem (numpy, pandas, torch, transformers), but is less suited to ultra-low-latency services. You see it everywhere: from notebooks to production services.
Minimal example (data validation):
def validate(rows):
return all("label" in r and r["label"] is not None for r in rows)Go
Go is often used for high-performance services (gateway, feature service, control plane) because concurrency is simple and binaries are static. It typically sits in front of inference servers to handle routing, batching, or rate limiting.
Minimal example (health endpoint):
package main
import "net/http"
func main() {
http.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
w.Write([]byte("ok"))
})
http.ListenAndServe(":8080", nil)
}Java
Java remains very common in enterprises for ingestion, ETL (Extract, Transform, Load), business services, and streaming, with a mature CI/CD toolchain. It is often the data/ops layer around the model, e.g., consuming Kafka or orchestrating business workflows.
Minimal example:
public class App {
public static void main(String[] args) {
System.out.println("ok");
}
}C++
C++ is chosen for maximum performance, CUDA kernels, and custom ops. It is used for fast preprocessing, ultra-low-latency inference engines, or native bindings.
Minimal example:
#include <vector>
float dot(const std::vector<float>& a, const std::vector<float>& b) {
float s = 0.0f;
for (size_t i = 0; i < a.size(); ++i) s += a[i] * b[i];
return s;
}7.2 Orchestration (Kubernetes, Slurm)
Kubernetes (K8s)
K8s is excellent for inference, multi-tenant services, autoscaling, and lifecycle management. It supports Jobs/Deployments, GPU scheduling via the NVIDIA device plugin, and integrates naturally with CI/CD and Helm. Kubeflow adds ML pipelines (KFP, Kubeflow Pipelines), training (Training Operator), and metadata (Katib).
Slurm
Slurm is ideal for massive multi-node HPC training, with fine-grained GPU topology control and exclusive allocations. It is less suited to real-time scaling and multi-tenancy.
Typical integration: K8s serves inference and services (vLLM for LLMs), Slurm handles massive training, and a modern variant is K8s + Volcano/Kubeflow if you want to unify.
7.3 DevOps tools (Git, Docker, Helm)
Git
Git versions code, pipelines, and configuration; the commit hash helps trace a production model, and tags represent model releases.
Example:
git tag model-v1.2.0Docker
Docker makes training/inference environments reproducible, notably through versioned CUDA images and multi-stage builds. It is the foundation for portability across dev, staging, and prod.
Helm
Helm packages and parameterizes Kubernetes deployments, enabling a new version by changing values (values.yaml).
Example values.yaml:
image:
repository: my-ml-image
tag: v1.2.0
resources:
limits:
nvidia.com/gpu: 17.4 Infrastructure as Code (Terraform)
Terraform reproducibly provisions GPU clusters, networking, and storage, with centralized state for audit and rollback. The standard integration is: create infrastructure via Terraform, then deploy services via CI/CD, with everything versioned in Git.
Minimal example (pseudo code):
resource "gpu_cluster" "train" {
node_count = 8
gpu_type = "A100"
}7.5 ML frameworks and libraries
PyTorch is the de facto standard for training and research, with DDP/torch.distributed and natural DeepSpeed integration. TensorFlow remains very mature for production (TensorFlow Serving, TFX - TensorFlow Extended) and still exists in enterprise stacks. JAX is valued in research for XLA and pmap/pjit sharding. Hugging Face standardizes LLM workflows with Transformers, Datasets, Tokenizers, and PEFT. Scikit-learn is still ideal for quick baselines and interpretable tabular models.
Quick mapping:
| Need | Tool |
|---|---|
| Tabular baseline | Scikit-learn |
| General deep learning | PyTorch |
| LLM fine-tuning | Hugging Face + PyTorch |
| High-performance research | JAX |
| Legacy enterprise pipelines | TensorFlow |
Integrated stack example: Python + PyTorch/DeepSpeed for training, Git + Docker for versioning and packaging, Terraform to provision the cluster, Slurm for massive training, Kubernetes + Helm for inference and services, and Hugging Face for LLM + PEFT.
8. Platform layer and go-to-prod
This section covers the platform layer and what it takes to move from a POC to a production system. The goal is to show you can think like an architect, with an end-to-end view and operable criteria.
8.1 Platform layer (cloud + data + compute)
The platform layer is the set of cloud primitives that enable a large-scale ML system: GPU compute, high-performance networking, storage, security, observability, and governance. It must provide a stable foundation for distributed training, multi-tenant inference, and MLOps workflows.
On compute: GPU pools (A100, H100, etc.), isolation (MIG/MPS), quotas, and scheduling. On networking: low latency and high bandwidth (RDMA, Remote Direct Memory Access, InfiniBand or RoCE, RDMA over Converged Ethernet) with topology-aware placement to limit inter-node communication. On storage: object storage for datasets, parallel file systems for massive training, and NVMe cache to smooth throughput for the data loader.
Dtypes and modern GPUs (A100 vs H100)
Objectives:
- Clarify what dtype means for an LLM and why it changes performance.
- Connect dtypes to modern GPU hardware capabilities.
Key points:
- A dtype (FP32, FP16, BF16, TF32, INT8, FP8) is the data type used to encode numbers.
Analogy: it is a "precision label", like a barcode that encodes how fine the stored values are.
- A100 GPUs are optimized for FP16/BF16/TF32 via Tensor Cores; they do not support FP8 natively.
- H100 GPUs add FP8, which can accelerate some LLM workloads while reducing memory, at the cost of stricter numerical stability management.
Concrete example: An FP8 LLaMA-2 model could reduce memory by ~50% compared to FP16. On A100, you must convert to BF16/FP16 to run it, which removes part of the gain. On H100, FP8 is native: you can actually benefit from throughput and memory reduction.
Mini comparison (order of magnitude):
| GPU | Memory | Native AI dtypes | Strengths | Practical limits |
|---|---|---|---|---|
| A100 | HBM2e 40/80GB | TF32, FP16, BF16 | Mature, large ecosystem, very stable | No native FP8 |
| H100 | HBM3 80GB | FP8, BF16, FP16 | Very strong for LLMs, native FP8 | High cost, needs optimization |
| MI300 | HBM3 128/192GB | BF16, FP16, INT8 | Very large memory, good perf | Less standard software stack |
From a security and operations standpoint, the platform must integrate IAM (Identity and Access Management) and RBAC (Role-Based Access Control), at-rest and in-transit encryption, secrets management, audit logs, and observability (metrics, logs, traces). Cost controls (tags, quotas, showback/chargeback) are required to avoid overruns on expensive GPU clusters.
Questions to ask on the platform side: what latency/throughput SLO and what budget per request or per token, what GPU/network topology is required, how to isolate teams and prioritize critical jobs, and what level of observability is needed to diagnose degradations.
8.2 Go-to-prod (POC -> production)
Going from POC to production means formalizing reproducibility, model quality, deployment robustness, and continuous monitoring. It is not a simple "push" of the model, but a set of quality and ops gates.
In practice, you first want strict reproducibility: versioned code, data, and configuration; pinned environment (Docker); seeds; deterministic pipelines. Then evaluation must cover offline metrics, non-regression tests, robustness tests (drift, outliers, prompt injection for LLMs), and A/B comparisons. Packaging goes through a model registry, signed artifacts, a stable input schema, and clear documentation.
Deployment must support canary, shadow, and rollback, with SLOs and error budgets defined. Monitoring must track latency, cost per request, business performance, data quality, and drift, with actionable alerts. Finally, runbooks and incident management make the system operable over time.
8.3 Practical rules (go-to-prod)
Golden rule - Go-to-prod
No production without proven reproducibility (data/code), observability, and rollback.
Useful reminders:
- Pin versions (code, data, deps) before production rollout.
- Use canary/shadow and explicit SLOs before global rollout.
- Prepare a rollback plan and an incident runbook.
9. Future: HPC/AI trends
This section lists emerging technologies that reshape HPC/AI platforms. The goal is to show you follow the trends and understand where they fit.
9.1 Hardware and accelerators
The market moves toward denser GPUs (NVIDIA Blackwell B200/GB200 with HBM3e, High Bandwidth Memory), competitive alternatives (AMD MI300, Intel Gaudi3), and lower precisions (FP8/FP4 floating-point) with sparsity. DPUs (Data Processing Unit) and SmartNICs (e.g., BlueField), as well as CXL (Compute Express Link), become important for offload and memory pooling.
9.2 Network and interconnects
Interconnects evolve to InfiniBand NDR/XDR or 400/800G Ethernet with RoCEv2 (RDMA over Converged Ethernet v2), while NVLink/NVSwitch increases generations for intra-node scaling. NCCL/SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) offloads reductions and improves latency, but requires topology-aware scheduling.
9.3 Storage and data pipelines
Massive training relies on parallel file systems (Lustre, GPFS - General Parallel File System) or NVMe-oF (NVMe over Fabrics), with local caching to stabilize throughput. You also see more data streaming and GPU-side preprocessing (DALI, Data Loading Library) to avoid CPU bottlenecks.
9.4 Compilers and runtimes
Compilers like torch.compile (Inductor), XLA, Triton, or TVM fuse operations and reduce overhead. CUDA Graphs reduces per-iteration launch overhead, and distributed/async checkpointing strategies (FSDP2, Fully Sharded Data Parallel v2, ZeRO++) speed up recovery.
9.5 Training and models
Large-scale MoE improves with more stable routing, while long context relies on FlashAttention v2/v3 or ring attention. Quantization-aware training and advanced mixed precision become standard.
9.6 Inference engines
Modern engines include vLLM, TensorRT-LLM, TGI, and SGLang for LLM throughput. Speculative decoding and continuous batching reduce latency, and multi-model serving becomes common with dynamic routing based on cost/latency.
9.7 Orchestration and platform
K8s + Kueue/Volcano gains traction for HPC batch jobs, with Kubeflow for ML pipelines. GPU sharing progresses (MIG, MPS, time slicing) and multi-tenant policies (quota, priority, fair sharing) become platform requirements.
10. Global glossary
Terms (A-Z):
- A/B test - controlled comparison of two production variants.
- Activation checkpointing - recompute activations to reduce memory.
- AUC - ranking metric (area under ROC curve).
- Batch - set of examples processed in parallel.
- BF16 - reduced-precision format, stable for training.
- Canary - progressive rollout to a small fraction of traffic.
- Cold start - initial latency due to model loading.
- Data contracts - schema/SLA/semantics contract for data.
- Data drift - drift in data distributions in production.
- DDP - DistributedDataParallel (PyTorch), gradient synchronization.
- DP/TP/PP - data/tensor/pipeline parallelism.
- Elastic training - ability to survive node failures.
- Feature store - centralized system to store/serve features.
- Fine-tuning - adapt a pre-trained model to a domain.
- FSDP - Fully Sharded Data Parallel (state sharding).
- GIL - Python Global Interpreter Lock.
- GPU - Graphics Processing Unit for parallel compute.
- Gradient accumulation - accumulate before update to emulate a large batch.
- Hybrid parallelism - combination of DP + TP + PP.
- IAM/RBAC - identity management and role-based permissions.
- Inference - prediction phase in production.
- KV cache - key/value cache to speed up LLM inference.
- Latency P95/P99 - latency percentiles (tail).
- LoRA - parameter-efficient fine-tuning via low-rank matrices.
- MIG/MPS - GPU sharing (Multi-Instance / Multi-Process).
- Model registry - versioned repository of deployable models.
- MoE - Mixture of Experts.
- NCCL - NVIDIA library for GPU collectives.
- NVLink - fast intra-node interconnect.
- NVMe - fast storage for offload/checkpoints.
- OOM - out of memory.
- PagedAttention - vLLM technique for KV cache management.
- Pipeline - chaining of data/train/deploy steps.
- QPS - queries per second.
- Quantization - precision reduction to speed up inference.
- RDMA - direct memory access for low-latency networking.
- SLA/SLO - service agreement/objective targets.
- Shadow - parallel deployment without affecting production.
- Sharding - partitioning of data or parameters.
- Tensor parallelism - intra-layer partitioning.
- Throughput - amount processed per unit of time.
- TPU - Google ML accelerator.
- Warmup - gradual learning rate ramp-up.