NVIDIA Multi-Instance GPU (MIG) partitions supported NVIDIA GPUs into isolated GPU Instances (GIs), which can be subdivided into Compute Instances (CIs). Each GPU Instance receives dedicated slices of memory-system resources and compute capacity, giving workloads predictable isolation compared with ordinary time-sharing on one full GPU. MIG is supported on selected Ampere and later data-center GPUs, with available profiles varying by GPU generation and memory configuration.
Within NVIDIA AI Infrastructure, MIG is a capacity-allocation tool for workloads that do not need a full accelerator. GPU Scheduling Basics provides the scheduler context; MIG changes the resource presented to that scheduler from one physical GPU into several smaller schedulable devices.
Start with the workload shape, not the number of slices
Profile model memory, SM utilization, tensor/core usage, encoder/decoder engines, latency, and throughput before selecting MIG profiles.
A workload that needs 35 GB of memory cannot use a small partition simply because its average compute utilization is low.
Choose the smallest profile that meets peak memory and performance SLO with margin.
GPU Instances isolate memory-system paths
MIG assigns dedicated memory slices, L2/cache and memory-controller resources to a GPU Instance according to the supported profile.
This gives stronger performance and fault isolation than sharing one full GPU through uncoordinated processes.
It is especially useful for multi-tenant inference, notebooks, CI, or smaller training jobs.
Compute Instances subdivide a GPU Instance
A GPU Instance can be subdivided into Compute Instances where the hardware/profile supports it.
CIs share the memory resources of their parent GI while partitioning compute resources further.
Use this only when the application/runtime/scheduler stack understands the resulting device model and the memory-sharing behavior meets isolation expectations.
Profiles are hardware-generation specific
MIG profile names encode resource shape but differ across A100, H100/H200, Blackwell-family, and other supported accelerators.
Do not copy an A100 profile table into a newer GPU design.
Use the current MIG User Guide for the exact profiles available on the installed GPU and driver.
Enabling MIG changes the scheduling inventory
When a GPU enters MIG mode and instances are created, schedulers see MIG resources instead of the same full-GPU inventory.
Plan maintenance windows because changing partition layouts can require draining workloads and reconfiguring the GPU.
A dynamic capacity policy should not destroy running instances merely to satisfy a new small job.
GPU Operator can reconcile MIG state
Current GPU Operator deploys MIG Manager and supports strategies that expose MIG resources to Kubernetes.
NVIDIA GPU Operator should own the desired state when Kubernetes manages the fleet.
Keep node labels and MIG configuration policy under version control rather than editing individual nodes.
Benchmark noisy-neighbor behavior explicitly
MIG provides hardware isolation, but workloads can still share host CPU, PCIe, NIC, storage, power/cooling, and scheduler resources.
Run simultaneous workloads in neighboring instances and measure latency, throughput, host CPU, I/O, and network effects.
Do not market “fully isolated” performance if the rest of the server remains a shared bottleneck.
Monitoring needs MIG-aware entities
DCGM and scheduler telemetry can report physical GPU plus GPU Instance/Compute Instance entities depending on metric support.
Attribute utilization, memory, errors, and jobs to the MIG slice actually assigned.
NVIDIA DCGM Monitoring explains why physical-GPU-only dashboards can hide which tenant or slice is affected.
Use full GPUs for workloads that need interconnect scale
Large training jobs that rely on full memory capacity, NVLink/NVSwitch, GPUDirect RDMA, or high collective bandwidth may be better served by full GPUs.
MIG can improve utilization for smaller workloads but should not become a blanket policy that fragments every node.
Maintain separate full-GPU and partitionable pools when the workload mix warrants it.
Reconfiguration should be capacity planned
Track demand by requested memory/compute profile over time.
If jobs frequently queue because the fleet has the wrong MIG shapes, the partition policy is the bottleneck.
Use scheduled rebalancing or dedicated node pools rather than continuously reconfiguring busy GPUs.
MIG succeeds when partitions match predictable workload envelopes
The mature design profiles real workloads, chooses supported hardware-specific profiles, drains before reconfiguration, reconciles state through GPU Operator, monitors slices independently, and preserves full-GPU pools for communication-heavy jobs.
MIG increases utilization by creating right-sized accelerators; it should not create a new fragmentation problem that wastes capacity in smaller pieces.
MIG capacity planning should use profile demand histograms rather than average GPU utilization. Track how many jobs ask for each memory/compute envelope, how long they run, and how often they queue because the requested shape is unavailable. A fleet can show 40% aggregate utilization while many jobs wait because free capacity is fragmented into the wrong MIG profiles.
Profile selection should consider inference batch size and KV-cache growth for language models. Memory requirements can rise with sequence length, concurrency, and caching even when model weights fit comfortably. Benchmark the production request envelope, not a one-request smoke test, before committing a service to a small MIG slice.
Training on MIG should be evaluated for communication limits. Smaller partitions are excellent for independent jobs, but large distributed training often depends on full GPU memory and peer/interconnect behavior. Keep job classes that need full-device NCCL/NVLink performance away from aggressively partitioned pools unless the supported GPU/MIG configuration is validated for that training pattern.
GPU reset/reconfiguration behavior varies by hardware generation and deployment. Some MIG transitions can be disruptive and require all workloads off the device. Automate drain checks and block reconfiguration when active processes remain. The partition controller should fail safely rather than force a new layout through a busy GPU.
MIG identifiers should be treated as ephemeral scheduling resources, not stable business identity. After reconfiguration, instance IDs and device enumeration can change. Applications should request resource classes through Kubernetes/scheduler abstractions instead of storing one MIG UUID in long-lived configuration unless the platform guarantees that binding.
Security isolation should include host and container boundaries around MIG. Hardware partitions isolate key GPU resources, but tenants can still share kernel, driver, host OS, PCIe path, network and storage. Use Kubernetes namespace/pod security, network policy and storage isolation alongside MIG when the requirement is multi-tenant security rather than only performance isolation.
Monitoring should distinguish allocated from actually used slices. One team can reserve several MIG instances and leave them idle, causing apparent hardware fragmentation. Track allocation duration, active compute/memory use, and queue pressure so quota or preemption policy can reclaim capacity where business policy permits.
Device plugin strategy affects resource names. GPU Operator supports MIG exposure strategies that can present homogeneous or mixed resource types. Choose the strategy that matches scheduler expectations and workload manifests, and keep the cluster consistent so one namespace does not see a different resource contract after node replacement.
Maintenance windows should consolidate fragmentation when demand changes. If daytime inference uses many small slices and overnight fine-tuning needs full GPUs, scheduled reconfiguration can create time-based capacity pools. This is more predictable than constantly changing MIG geometry in response to individual queue arrivals.
MIG success metrics should include utilization per slice, queue time by profile, reconfiguration frequency, failed scheduling, workload SLO, and stranded capacity. The objective is not the maximum number of instances per physical GPU; it is more completed useful work per accelerator without unacceptable performance variance.
Workload owners should understand that a MIG profile is a capacity class, not a percentage of one generic GPU. Different generations expose different memory, SM, engine and bandwidth characteristics even when profile names look similar. Benchmark per hardware family before offering one abstract ‘small GPU’ service tier.
MIG and power/cooling policy interact. Packing many active slices on every physical GPU can increase sustained utilization and rack power compared with fleets where full GPUs often sit partially idle. Capacity planning should consider thermal and electrical consequences of higher utilization, not only the scheduler benefit.
Incident response should identify whether the fault affects one compute instance, one GPU instance, the full physical GPU, PCIe path or host. A job failing on one MIG slice does not automatically mean every neighboring tenant should be evicted. Use DCGM entity-level evidence and driver/XID scope before draining the full node.
Quota models should charge for the resource actually reserved. A tenant using one large MIG profile and another using several small profiles should consume quota in a way that reflects scarce memory/compute rather than one resource object each. Otherwise scheduler accounting can encourage fragmentation and unfair allocation.
Profile changes should be documented in platform release notes because they alter the resource contract seen by users. If `nvidia.com/mig-*` classes change or disappear, existing manifests can become unschedulable. Give application teams migration time and provide replacement resource classes before removing old layouts.
Use a fallback pool of full GPUs for workloads whose demand is hard to predict. Some inference models grow memory with context/concurrency and may outgrow a slice during peak traffic. Full-GPU fallback can preserve service while the team decides whether to resize the profile permanently.
Document which workload classes are allowed to request each MIG profile and which full-GPU pools remain protected for distributed jobs. Clear service tiers reduce ad hoc scheduling exceptions and make capacity forecasting easier.
Partitioning policy should consider memory demand, compute profile, latency sensitivity, and scheduling frequency. Stable MIG shapes make placement predictable, but excessive fragmentation can leave capacity stranded even while no individual job can obtain the slice it needs.