Kubernetes Storage: Persistent Volumes and Storage Classes

Persistent storage in Kubernetes separates how an application asks for storage from how a cluster administrator or storage system provides it. The current CKA exam includes Storage as a core domain. The fundamental resources are PersistentVolume (PV), PersistentVolumeClaim (PVC), and StorageClass, with CSI drivers and storage backends implementing the actual volume lifecycle.

The mechanisms described in Kubernetes PersistentVolumeClaims become much easier to reason about when the lifecycle is explicit: a workload references a PVC; the PVC requests capacity, access mode, and optionally a storage class; Kubernetes binds it to a suitable PV or triggers dynamic provisioning; the node then mounts or attaches the storage through the relevant driver.

The confusing cases appear when scheduling, topology, reclaim policy, defaulting, access modes, expansion, and backend capabilities intersect. Storage is therefore both a control-plane binding problem and a real infrastructure dependency.

A PV is cluster storage with its own lifecycle

A PersistentVolume represents storage available to the cluster and exists independently of any one Pod.

It can be statically provisioned by an administrator or dynamically created by a provisioner in response to a claim.

Deleting or replacing a Pod does not imply deleting the PV. This independence is the core reason state can survive workload rescheduling.

PV lifecycle should also distinguish the Kubernetes object from the external storage asset. Deleting a PV object, changing reclaim policy, or restoring cluster state does not always imply the external disk or filesystem has the same lifecycle. CSI provisioners translate Kubernetes intent into provider operations. During disaster recovery, inventory which side is authoritative for existence and identity so the team does not accidentally create duplicate disks, orphan data, or import an external volume under the wrong PV metadata.

A PVC is the workload’s storage request

A PersistentVolumeClaim asks for storage characteristics rather than naming a particular physical disk in the normal abstraction.

Capacity, access mode, storageClassName, selector, and other attributes influence which PV can satisfy the claim.

Debug a Pending PVC before debugging a Pod mount. If the claim never bound, the workload cannot successfully consume the requested persistent storage.

PVC design should consider requested capacity as a minimum request and backend behavior. The bound volume may have provider-specific allocation units or expansion semantics. Applications should not assume usable filesystem capacity changes the instant the PVC specification changes; controller expansion and node/filesystem expansion can occur in stages. Monitor the PVC, PV, CSI events, and filesystem state. Treat expansion as a lifecycle operation with backup and rollback considerations, especially for stateful databases where storage manipulation can be difficult to reverse.

StorageClass expresses an administrator-provided storage profile

A StorageClass specifies a provisioner and parameters and can also define reclaim policy, volume binding mode, expansion, and topology-related behavior.

The class name is part of the application/cluster contract: a workload requesting fast-rwo depends on administrators maintaining what that class is meant to represent.

Keep the names semantic enough for workload teams to choose correctly without embedding provider-specific details into every application manifest.

StorageClass names become an API offered by the platform team. Changing the parameters behind a class without a new class name can make newly provisioned volumes behave differently from older volumes while manifests look identical. For significant changes in performance, topology, encryption, backup integration, or reclaim behavior, consider whether a new class communicates the new contract more clearly. Application teams need stable semantics; platform teams need freedom to evolve backends. Versioned or purpose-driven class names can balance those needs.

Dynamic provisioning shifts creation to the claim path

When a PVC requests a supported StorageClass, a provisioner can create a matching PV automatically.

This removes manual PV precreation and introduces dependencies on the CSI/provisioner, cloud or storage API, quota, topology, and credentials.

A failed provisioning event should be traced through PVC events and driver/controller evidence rather than treated as a generic scheduler problem.

Dynamic provisioning can also fail after the PV is created but before the Pod mounts it. Cloud API success may be followed by attachment quota, node permissions, CSI node-plugin problems, device discovery, filesystem formatting, or mount errors. Keep controller-side and node-side CSI logs conceptually separate. The provisioner/controller usually handles creation and attach orchestration; node plugins participate in staging and mounting. The Pod event stream often identifies which phase failed, allowing operators to avoid deleting a perfectly healthy disk.

Binding mode can coordinate storage with scheduling

Immediate binding can provision or bind storage before the Pod is scheduled.

WaitForFirstConsumer delays binding/provisioning until a consuming Pod exists, allowing topology and scheduling constraints to influence where the volume is created.

This matters for zonal storage. A volume created in the wrong zone can make an otherwise schedulable Pod impossible to place on nodes that can access the storage.

WaitForFirstConsumer is particularly useful when storage topology must align with Pod scheduling, but it also couples two decisions. The scheduler considers topology and workload constraints while selecting a node, and the storage system then provisions or binds accordingly. Complex affinity, taints, and limited zone capacity can make both appear stuck. Inspect scheduler and PVC events together. Solving only one side may reveal the next constraint, which is expected behavior when Kubernetes is trying to satisfy both compute and storage placement simultaneously.

Access modes describe allowed attachment semantics

ReadWriteOnce, ReadOnlyMany, ReadWriteMany, and related modes express how the volume may be mounted or used, subject to driver/backend support.

These are not performance guarantees and do not automatically mean every filesystem or storage backend behaves identically.

Match workload topology to the storage system. A replicated application expecting shared writable storage needs a backend that actually supports the intended access pattern.

Access modes are not enforcement by themselves in every backend. They describe intended ways a volume can be mounted and rely on driver/storage capabilities. Applications that require strict single-writer behavior may need StatefulSet identity, application-level locking, or storage semantics in addition to the Kubernetes claim. Understand the backend guarantee. A configuration labeled ReadWriteOnce should not be used as the only protection against concurrent writers if the failure model requires stronger fencing than the storage system provides.

Reclaim policy determines what happens after the claim

Retain and Delete express different post-claim behavior for dynamically or statically provisioned storage, depending on class/PV configuration.

Delete can automate cleanup and destroy data when the claim lifecycle is misunderstood; Retain preserves storage and creates an explicit cleanup/reuse responsibility.

Choose the policy from data value and recovery process, not from convenience. Test deletion in a nonproduction environment so application owners understand the consequence of deleting a PVC.

Reclaim-policy choice should include backup and legal retention. Delete can be correct for ephemeral state and dangerous for records that must be recoverable after accidental namespace cleanup. Retain can preserve valuable data and accumulate expensive orphaned volumes nobody owns. Track ownership, age, and cleanup status for retained assets. Platform automation can notify application owners or move retained storage through a reviewed disposal process. The best policy is the one whose downstream operational process actually exists and is tested.

Storage failures often cross scheduler and node boundaries

A Pod can be Pending because the PVC is unbound, because topology cannot match, or because attach/mount cannot succeed on the selected node.

Inspect PVC/PV binding, StorageClass, scheduler events, node topology, CSI controller/node components, attachment objects where applicable, and kubelet events in that order.

The cluster relationships in Kubernetes anatomy matter because storage control plane and node-level mount/attach paths fail at different stages.

Storage topology can interact with node maintenance. Draining a node that hosts a Pod with a zonal ReadWriteOnce volume may reschedule successfully only to nodes in the same zone. If that zone has insufficient capacity, the workload remains Pending even though other zones are healthy. This is not a generic scheduling failure; it is the consequence of data locality. Resilience planning should include storage placement and replicated data architecture rather than assuming the scheduler can move every stateful Pod anywhere in the cluster.

A practical storage review follows data through failure

Create a stateful workload, claim storage through a class, verify binding and node placement, restart the Pod, move it where supported, expand or replace it according to the scenario, then test deletion and recovery expectations.

Record which object owns each transition and which backend evidence proves the operation.

Persistent storage becomes predictable when teams can distinguish request, binding, provisioning, topology, attachment, mount, data lifecycle, and cleanup—and know which of those states survives when the workload itself is replaced.

A storage incident should close only after the data path is validated, not merely when the PVC becomes Bound. Confirm the workload can read/write expected data, filesystem health is normal, performance is within baseline, and the volume has the intended backup/reclaim/encryption posture. Then reconcile any manual PV edits or emergency StorageClass changes. Stateful recovery has a higher bar than a stateless Pod restart because a superficially healthy mount can still expose stale, partial, or wrongly attached data.

Storage architecture should also define backup responsibility independently of PV/PVC lifecycle. A dynamically provisioned volume can be durable across Pod restarts and still have no backup or point-in-time recovery. CSI snapshots, provider backups, application-consistent backup, and disaster recovery each add different guarantees. Platform teams should state what the StorageClass promises and what it does not. Application owners should not infer that ‘persistent’ means ‘protected from deletion, corruption, operator error, or regional loss.’

Persistent-storage reviews should keep a record of which StatefulSets or critical workloads depend on each class and backend. Changing a default StorageClass or retiring a CSI driver can affect future claims without touching existing PVs, creating a mixed estate with different behaviors. Dependency inventory helps platform teams migrate deliberately and prevents a class rename or backend retirement from becoming a surprise only when the next replica or environment needs new storage.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!