Skip to content

Infrastructure & operations

Understand · 15 min · no code

Before this: Fine-tuning and training · After this: The build path Hands-on version: 11 Production · In depth: Code quality pipeline

Start with the hands-on module

For the buildable version of this material, see Production. That module covers retries, idempotency and budgets with a runnable lab. This page covers the wider MLOps picture — drift, monitoring, deployment topology.

Building an AI model is only half the challenge. Running it reliably, efficiently, and cost-effectively in production is the other half. This page covers MLOps (the operational practices for machine learning), model optimization, and the infrastructure decisions that determine whether your AI system scales or stalls.


MLOps: DevOps for machine learning

MLOps (Machine Learning Operations) applies the principles of DevOps — automation, monitoring, version control, CI/CD — to the machine learning lifecycle. It bridges the gap between data science experiments and production systems.

MLOps vs DevOps

Aspect DevOps MLOps
What is versioned Code Code + data + models + experiments
What is tested Application behavior Model accuracy + application behavior
What is deployed Application artifacts Model artifacts + serving infrastructure
What is monitored Uptime, latency, errors Uptime, latency, errors + model accuracy + data drift
What triggers redeployment Code changes Code changes + data changes + model degradation
Pipeline Build, test, deploy Ingest, train, evaluate, deploy, monitor

MLOps maturity

Most teams start at "manual everything" (level 0) and gradually automate. You do not need a fully automated MLOps pipeline on day one. Start with version control for data and models, then add automation incrementally.

The MLOps lifecycle

graph LR A["Data<br/>Collection"] --> B["Data<br/>Preparation"] B --> C["Model<br/>Training"] C --> D["Model<br/>Evaluation"] D --> E["Model<br/>Registry"] E --> F["Deployment"] F --> G["Monitoring"] G -->|"Drift<br/>Detected"| A style A fill:#0284c7,stroke:#0284c7,color:#fff style B fill:#0284c7,stroke:#0284c7,color:#fff style C fill:#0d9488,stroke:#0d9488,color:#fff style D fill:#0f766e,stroke:#0d9488,color:#fff style E fill:#0f766e,stroke:#0d9488,color:#fff style F fill:#0284c7,stroke:#0284c7,color:#fff style G fill:#16a34a,stroke:#16a34a,color:#fff

Key stages:

  1. Data Collection: Gather raw data from sources — databases, APIs, logs, user interactions.
  2. Data Preparation: Clean, transform, and feature-engineer the data. Track lineage so you know where every data point came from.
  3. Model Training: Train (or fine-tune) the model. Log hyperparameters, metrics, and artifacts.
  4. Model Evaluation: Compare the new model against baselines using defined metrics. Automated evaluation gates prevent bad models from reaching production.
  5. Model Registry: Store versioned models with metadata (who trained it, on what data, with what performance).
  6. Deployment: Serve the model via an API endpoint, batch pipeline, or edge device.
  7. Monitoring: Track model performance, data quality, and operational health in production. When quality degrades, trigger the cycle again.

Model drift and monitoring

Model drift is the gradual degradation of model performance over time. A model that was accurate at launch may become unreliable as the real world changes around it.

Types of drift

Data drift
The distribution of input data changes. For example, a customer sentiment model trained on pre-pandemic reviews may perform poorly on post-pandemic data because the language and topics shifted.
Concept drift
The relationship between inputs and outputs changes. For example, a fraud detection model may degrade as fraudsters develop new techniques that look nothing like historical fraud patterns.
Feature drift
The data pipeline changes, causing features to be computed differently or become unavailable. For example, a feature that previously held "days since last purchase" is now always zero due to a data pipeline bug.

Monitoring strategy

What to Monitor How Alert Threshold
Prediction accuracy Compare predictions to ground truth (when available) Accuracy drops below defined baseline
Input data distribution Statistical tests (KS test, PSI) comparing current vs training data Distribution shift exceeds threshold
Output distribution Track the distribution of model predictions over time Sudden changes in prediction patterns
Latency Measure end-to-end response time P95 latency exceeds SLA
Error rates Track failed predictions, timeouts, and exceptions Error rate exceeds baseline
Token usage Monitor tokens consumed per request Unexpected spikes in consumption

Monitoring is not optional

In production AI systems, monitoring is as critical as the model itself. Without it, you will not know your model is degrading until users complain — or worse, until bad decisions are already made.


Quantization and model optimization

Quantization reduces the precision of a model's numerical weights — for example, from 32-bit floating point (FP32) to 8-bit integers (INT8) or even 4-bit. This dramatically reduces model size, memory usage, and inference latency, often with minimal impact on quality.

How quantization works

Precision Bits per Weight Relative Size Typical Quality Impact
FP32 32 1x (baseline) None (full precision)
FP16 / BF16 16 0.5x Negligible
INT8 8 0.25x Minimal for most tasks
INT4 4 0.125x Noticeable on complex reasoning

Quantization methods

Post-training quantization (PTQ)
Applied after training is complete. No additional training data is needed. Fast and easy but may lose more quality than training-aware methods.
Quantization-aware training (QAT)
Simulates quantization during training, allowing the model to adapt its weights to lower precision. Produces better quality but requires training infrastructure.
GPTQ / AWQ / GGUF
Specialized quantization formats for LLMs. GPTQ and AWQ are GPU-focused, while GGUF (used by llama.cpp) is optimized for CPU and edge inference.

Start with INT8

For most deployment scenarios, INT8 quantization provides an excellent balance of size reduction and quality preservation. Only go to INT4 if you have strict hardware constraints and can tolerate some quality loss.


Edge AI and on-device inference

Edge AI runs models directly on local devices — laptops, phones, IoT devices, on-premise servers — rather than sending data to the cloud. This is increasingly practical with small language models and quantization.

When to use edge AI

Scenario Why Edge Makes Sense
Data privacy Sensitive data never leaves the device or local network
Low latency No network round-trip to a cloud API
Offline operation Works without internet connectivity
Cost at scale No per-query API costs for high-volume use cases
Regulatory compliance Data residency requirements mandate local processing

Edge AI technologies

Technology Description
ONNX Runtime Cross-platform inference engine, supports quantized models
llama.cpp C++ inference for LLMs, runs on CPU, supports GGUF quantization
TensorFlow Lite Google's on-device ML framework
Apple Core ML On-device inference optimized for Apple hardware
Windows ML ML inference on Windows devices using DirectML

Trade-offs

Edge AI is not free. You gain privacy, latency, and cost benefits, but you trade model capability. A 3B-parameter quantized model running on a laptop will not match the quality of GPT-4o running in the cloud. Choose edge deployment when the trade-off makes sense for your use case.


Serving models, and where latency comes from

If you call a hosted API, the provider handles this and you should still understand it, because it explains what you can and cannot make faster.

Generation happens one token at a time, and each token depends on the one before, so a response cannot be produced in parallel. That single fact drives almost everything about latency.

Term What it means Why you care
Time to first token Delay before the first character appears What users actually perceive as speed
Tokens per second Generation rate after the first token How long a long answer takes
Prefill Processing your prompt before generating Grows with prompt length; parallelisable
Decode Generating the response Grows with answer length; not parallelisable
KV cache Retained attention state so earlier tokens are not recomputed Why long conversations use memory, not just tokens

Three practical consequences:

  • Stream the response. It does not reduce total time, but it cuts perceived latency enormously. A user watching text appear after 300ms is having a better experience than one staring at a spinner for six seconds, even when the second finishes sooner.
  • Shorter outputs are the real speed lever. Prompt length affects prefill, which is fast and parallel. Output length affects decode, which is neither. "Answer in one sentence" does more for latency than trimming the prompt.
  • Batching trades latency for throughput. Serving stacks group concurrent requests to use the GPU efficiently. Good for cost per request, slightly worse for any individual one.

If you self-host, this is what a serving engine such as vLLM or NVIDIA NIM does for you: continuous batching, paged KV cache, and an OpenAI-compatible endpoint so your application code does not change. Running a model with a naive loop gives a small fraction of the throughput of the same hardware served properly.


Rate limits, quotas and failover

The operational surprise that hits most teams first, and it is rarely in the design document.

Providers limit both requests and tokens per minute, and the token limit usually binds first. A feature that works in testing fails at launch not because the code is wrong but because ten concurrent users with long prompts exceed a quota nobody checked.

What this demands:

  • Retry with exponential backoff and jitter, honouring the Retry-After header. Retrying immediately makes an overload worse.
  • Retry only what is safe to retry. A timeout does not tell you whether the work happened. If the call had a side effect, retrying can duplicate it — which is why idempotency keys exist. See production.
  • A queue for anything that can be asynchronous. Batch endpoints cost substantially less for work that does not need an immediate answer.
  • Know your fallback before you need it. A second deployment in another region, or a smaller model that degrades quality instead of failing. Decide which, and test it; a fallback path that has never run is not a fallback.
  • Cap per user and per tenant. Without it, one runaway loop consumes the quota for everyone, and the outage looks like a provider problem.

Model versions move underneath you

A model name is not a version. Providers update what a name points to, and deprecate versions on their own schedule. Behaviour can change without any deployment on your side.

Pin an explicit version where the provider allows it, subscribe to deprecation notices, and keep an evaluation set you can rerun to answer "did it get worse?" with evidence rather than impressions.


Cost management in AI

AI infrastructure costs can scale quickly if not managed carefully. Here are the main cost drivers and how to control them:

Cost drivers

Cost Driver Description How to Optimize
API tokens Pay-per-token for hosted model APIs Optimize prompts, cache responses, use smaller models for simple tasks
GPU compute Training and inference on GPU instances Use spot instances, right-size GPU SKUs, quantize models
Storage Vector databases, model artifacts, training data Compress embeddings, archive old model versions, use tiered storage
Data processing ETL pipelines, embedding generation, indexing Batch operations, incremental updates instead of full re-indexing
Monitoring Logging, tracing, evaluation Sample traces rather than logging everything, set retention policies

Cost optimization strategies

  • Remove unnecessary context from prompts
  • Use shorter system prompts
  • Cache common responses
  • Batch similar requests
  • Use SLMs for simple tasks (classification, extraction)
  • Use LLMs only for complex reasoning
  • Route requests to the cheapest capable model
  • Consider open-source models for high-volume workloads
  • Use auto-scaling to match demand
  • Deploy in regions with lower compute costs
  • Use spot/preemptible instances for training
  • Quantize models to reduce serving costs

Track cost per request

Establish a metric for cost per request or cost per user interaction. This helps you make informed decisions about model selection, prompt design, and infrastructure choices. Without this metric, costs tend to grow unnoticed.


Putting it all together

A production AI system brings together all of these concerns:

Layer Concerns Key Decisions
Model Selection, fine-tuning, quantization Which model? Cloud or edge? What precision?
Data Ingestion, embedding, indexing, freshness How often to re-index? What chunking strategy?
Application Orchestration, guardrails, caching What framework? What safety checks?
Infrastructure Compute, storage, networking GPU SKUs, auto-scaling, regions
Operations Monitoring, alerting, incident response What to monitor? What are the SLAs?
Cost Budgeting, optimization, chargeback Cost per request? Budget alerts?

Each layer has its own best practices, but they are deeply interconnected. A change in model selection (e.g., switching from GPT-4o to Phi-4) ripples through infrastructure (less GPU needed), cost (lower per-request), and application (may need prompt adjustments).


Go deeper

  • Azure Machine Learning — the full MLOps toolchain, most of which you do not need for LLM applications.
  • Azure AI Foundry — the parts you probably do need: deployments, evaluation and content filtering.
  • MLflow — experiment tracking and a model registry that works outside any one cloud.
  • ONNX Runtime quantization — the mechanics behind "make it smaller and faster", with the accuracy trade-off stated.
  • NVIDIA NIM — the realistic route to self-hosting a model behind an OpenAI-compatible endpoint.