OfferTransform Your Career with Expert-Led IT Training. Flat discounts active!Explore Now
OnlineITGuru Logo
Cloud Computing & DevOps

Architectural Foundations of Persistence in Kubernetes and OpenShift

Last updated on Oct 5, 2026

Copy Link:
Architectural Foundations of Persistence in Kubernetes and OpenShift

At first, container technology was adopted in businesses because of its ability to work with stateless microservices. In these projects, the containers were seen purely as temporary tools, that is, when an instance fails or is moved to another node, its local data is lost and a new instance can be created without any whereabouts of the previous one. Over time enterprises started using containerization for core elements of their operations, such as their databases, message bus, cache, machine learning applications, etc. Even though the process seems to be quite handy, deploying stateful applications in the cloud puts stress on both the principles of the cloud infrastructure and durability of the storage used.

Red Hat OpenShift addresses this issue by adding the storage features that Kubernetes lacks. The feature is that OpenShift separates the computation orchestration from storage characteristics. In this context, engineers need to comprehend the basics of the implementation: Persistent Volume, Persistent Volume Claim, Storage Layer, Dynamic Provisioners, and Container Storage Interface. Being acquainted with such production workflows typically implies undergoing practical training through a GoLeads redhat openshift course.

Persistent Volumes and Claims

The distinction between capacity provisioning and application consumption is seen through the lens of Persistent Volumes and Persistent Volume Claims. A Persistent Volume is an OpenShift cluster-scoped resource that encapsulates a physical storage asset, such as a SAN LUN, an EBS volume, a Ceph block image, or an NFS volume. It includes the storage plug-in, its connecting metadata, access credentials, and metadata relating to its lifecycle. Being cluster-scoped, Persistent Volumes are not directly visible to regular developers who work in specific namespaces.

What does the Persistent Volume Claim mean? It is simply a namespace-scoped request for storage. A developer has to provide a set of characteristics for the storage (the size, access mode, volume mode, and class of storage device) in order to receive such storage. OpenShift's control plane responds to such requests and automatically binds Persistent Volume Clauses with the corresponding volume configuration. If there is no available volume match, and if the storage class allows for automatic provisioning, then OpenShift will utilize the appropriate driver to create a necessary storage volume.

By ensuring proper categorization, this separation implements both restrict boundaries of multi-tenancy and the principle of least-privilege. As a result, developers may just ask for certain abstract features of the required storage service (e.g. durability or performance). Thus, there is no need for them to obtain the details about the infrastructure, such as the topology of the storage network or SAN zoning.

Storage Classes and CSI Drivers

In large modern company environments where pod numbers exceed hundreds, manual storage provisioning can be meaningless. Storage Classes manage the process of converting requests from applications into storage resource allocations. Each Storage Class description includes the provisioner’s name as well as its own volume expansion rules and policies, reclamation rules and policies, and backend-specific parameters, including RAID level, rate of replication, file system format, and encryption settings.

The Container Storage Interface is the technology behind modern OpenShift storage integration. Due to its existence, Kubernetes engine and software made by vendors are not coupled. In previous versions of the system all the storage plugins were written in the compilers of the system together with Kubernetes application, and the issues this approach caused were quite serious.

CSI abstracts storage connectivity into standard RPC-based processes. In general operations of CSI, the CSI operates as an operator in OpenShift, working with two separate control planes.

  • The Node Plugin runs as DaemonSet on the nodes of the cluster and interacts with the local kernel, performs device binding, starts iSCSI/Fibre channel connections, creates loop devices, formats the block device with the default file systems such as ext4 or XFS, and mounts the filesystem into the specified container namespace.

  • The Controller Plugin works as a replicated application on the control plane or dedicated infrastructure nodes and translates the requests of OpentShift's volumes into API calls for the vendor, performing actions for creating, deleting, attaching and detaching the storage, as well as for resizing and snapshotting the storage directly on the external storage system or on cloud API.

In-depth knowledge of the interaction of the CSI controller and node plugins during pod migrations is important to prevent file corruption and deadlocks regarding the storage connection during the cluster failover.

Considerations in Storage Types and Workload Compatibility

The choice of the architecture helps determine whether the application will process transactions quickly or be troubled with bottlenecks due to I/O limitations. Every type of enterprise application has its peculiarities in terms of access patterns, write intensities, and network interactions. OpenShift actually differentiates three basic forms of volume delivery: Block storage, File storage, and Object storage.

Block Storage

Block storage means that raw sectors without any formatting are sent directly to the OS. After that, the node either formats the device with the filesystem locally or provides it to a container engine that is responsible for its own low-level block storage implementation.

  • Attributes: Outstandingly minimal latency, top-notch IOPS, no overhead for network file systems, with direct integration of kernel-level page caches.

  • Drawbacks: Block volumes are restricted to being accessed from only one node at a time and cannot be mounted simultaneously at different nodes using self-sufficient pods in generic configurations.

  • Optimized Usage: For effective work, block storage is the underlying technology of the most efficient databases, write-heavy queues, and transactional storage engines, such as PostgreSQL, Oracle Database, Microsoft SQL Server, MongoDB, Apache Kafka, and Elasticsearch, because these engines depend on the predictable operations of fsync, write-ahead logging, and precise block allocation techniques that distributed filesystems compromise.

File Storage

File storage provides a hierarchical, shared file system namespace over network protocols like NFS or distributed parallel file systems such as CephFS .

  • Characteristics: Built-in support for concurrent shared read and write access from dozens or hundreds of individual worker nodes.

  • Trade-offs: Protocol-level latency and metadata overhead added. Every directory lookup, every file-locking call, every permission check means network round-trips. Concurrent writes by multiple clients require distributed lock managers, which can be a bottleneck under high contention.

  • Workload Alignment: File storage is a good fit for shared application runtimes, central configuration repositories, legacy monoliths sharing a persistent media upload directory, WordPress or Drupal asset directories, static web content distribution, continuous integration build caches, and shared home directories for scientific computing clusters.

Object Storage

Object storage uses flat namespaces to store data, decoupling content from the filesystem hierarchy, and allowing access to objects through a unique key. Applications access object stores via RESTful HTTP APIs, usually using the Amazon S3 standard.

  • Features: Unlimited horizontal scaling, geo-replication, strong eventual consistency or read-after-write consistency, built-in metadata enrichment, and complete independence from the host kernel mount procedures.

  • trade-offs: High time-to-first-byte latency compared to raw block access. No atomic file updates, append operations or in-place byte editing. To rewrite a single byte, re-upload the entire object payload .

  • Workload Alignment: Object Storage is the New Standard for Cloud-Native State It is the first tier of machine learning training datasets, raw data lakes, application backup repositories, long-term archive retention, document management systems, audio and video streaming repositories, and event-driven architectures.

Today many modern stateful design patterns are moving away from mounting bloated shared filesystems into container pods and toward refactoring applications to push unstructured blobs directly to S3-compliant endpoints.

Access Modes and Life Cycle Management

OpenShift has very strict access mode declarations for volume binding . These declarations specify the connectivity topology between worker nodes and underlying storage volumes. Reclaim policies control the fate of physical assets at the end of application lifecycles.

Profiles Mode Access

When a Persistent Volume Claim is created, operators set one of four standard access modes:

  • ReadWriteOnce (RWO): The volume can be mounted as read-write by all pods scheduled on a single node. That same physical or virtual node can be running multiple pods which all read and write to the volume , but pods on other nodes are not allowed to access it . This is the default mode of operation for block devices.

  • ReadOnlyMany (ROX): Many pods on many nodes can mount the volume as read-only. This profile is commonly used for reference datasets, read-only AI inference models and shared static binaries.

  • ReadWriteMany (RWX): Many pods can mount the volume as read-write, across the cluster, on different nodes. This mode requires distributed file systems like CephFS or enterprise NFS, or cloud services like AWS EFS or Azure Files.

  • ReadWriteOncePod (RWOP): To avoid an edge-case data corruption scenario, this mode ensures that only one pod in the entire cluster can mount the volume in read-write mode. Standard RWO permits two pods to share the same volume on a single node, whereas RWOP will not accept a second pod even if it is on the same node. This provides the best protection of single-instance database engines during canary or rolling deployments.

Orphan Lifecycle and Management Policies

A storage architecture needs to be forward-looking to application retirement, accidental deletion and disaster recovery. The Persistent Volume Reclaim Policy describes what happens to the storage backend when its claim is deleted:

  • Delete: The associated physical or cloud volume is automatically released and deleted as soon as the claim is deleted. This is the default for dynamic cloud storage classes. It's convenient in CI/CD environments, but has severe operational risk for production databases if a namespace or claim is accidentally removed.

  • Retain: The Persistent Volume remains in a Released state, locking the underlying data asset, when the claim is deleted. No other claim can be bound to the volume until an administrator explicitly audits, scrubs, or reclaims it. This is the recommended baseline for Tier 1 production environments.

  • Recycle: Does a basic data scrub (rm -rf) on the volume and makes it available for binding again. This policy is mostly deprecated in favour of more modern dynamic provisioners.

Standardising automated cleanup routines is necessary to reach operational maturity. Clusters use dynamic provisioning (Retain policy) which often results in storage sprawl i.e. abandoned storage volumes across cloud accounts leading to huge infrastructure costs hidden from view. Enterprise operations teams should implement scheduled auditing tools to identify unattached Persistent Volumes and cross-reference them with active namespaces and application ownership directories.

Integration with OpenShift Data Foundation (ODF)

Red Hat OpenShift Data Foundation (formerly OpenShift Container Storage) is the flagship software-defined storage architecture for organisations looking for a common storage platform that works natively on OpenShift across bare-metal, VMware, AWS, Azure and Google Cloud.

Internal vs External Operating Topologies

ODF is built on top of proven open source engines, Ceph for block, file and object abstractions, NooBaa (Multicloud Object Gateway) for advanced object virtualisation and deduplication and Rook as a Kubernetes native Ceph orchestration operator. ODF has two main deployment topologies:

  • Converged (Internal) Mode: ODF runs directly in the OpenShift application cluster. Worker nodes have dedicated NVMe or SSD local drives. The Rook-Ceph operator runs storage daemons (OSDs, MONs, MGRs, etc.) as containerised workloads on dedicated infrastructure nodes or co-located with application workloads. This topology reduces the infrastructure footprint and is ideal for edge deployments, remote branch offices and medium sized clusters where hardware economy is key.

  • External Mode: OpenShift connects to a separate, standalone Ceph cluster that is managed outside of the OpenShift compute plane with CSI plugins. This topology is best for large scale enterprise deployments. Platform teams can scale storage capacity, replace discs, and perform Ceph cluster maintenance without impacting OpenShift control plane nodes or triggering pod evictions by decoupling storage lifecycle management from Kubernetes upgrade cycles.

Unified Capabilities (UC)

ODF satisfies the classic enterprise multi-protocol requirement by abstracting physical drives into three dynamic Storage Classes:

  • Ceph RBD (RADOS Block Device) – fast, dynamic RWO block storage with instant snapshotting, fast cloning and thin provisioning.

  • CephFS – Provides highly available, scalable RWX shared storage as files across multiple nodes. Supports concurrent access patterns natively without external NFS hardware.

  • Multicloud Object Gateway (NooBaa) - Provides S3-compatible endpoints with dynamic bucket claims. This lets platform architects declaratively specify intelligent data placement policies, replicate buckets across public clouds, use bucket-level encryption keys, and manage data lifecycles all through OpenShift resource manifests.

Performance Optimisation and Tuning

Achieving enterprise storage performance on OpenShift requires optimization at every point of the I/O path, from the container engine to the Linux kernel to the network transport layer to the physical hard drive controllers. As fine-tuning each of these layers requires an extensive knowledge of underlying platform internals, system administrators typically train themselves by taking structured red hat openshift course related to advanced node configuration and software defined networking.

Tuning the Node and the Linux Kernel

Storage performance starts at the node operating system layer, Red Hat Enterprise Linux CoreOS. Containers share the host kernel, so misconfigurations in kernel parameters affect every workload running on that node.

Platform engineers should automatically tune stateful worker nodes using the OpenShift Node Tuning Operator:

  • I/O Schedulers: The default Linux schedulers (e.g. MQ-Deadline or BFQ) are tuned for traditional rotating media or mixed generic computing. The high performance NVMe backends should be set to none to bypass kernel queuing mechanisms and pass I/O requests directly to the drive controller hardware queues.

  • Filesystem Selection and Tuning: ext4 vs xfs and database behaviour XFS is usually more preferable for large datasets and parallel write operations because of its efficient allocation groups and dynamic inode management. Mount options such as the disabling of access time updates (noatime) can avoid the constant disc write overhead for basic read operations.

  • Virtual Memory Paging: High-throughput transactional workloads can flood host page caches leading to unpredictable I/O pauses as the kernel flushes dirty pages to disc. Adjusting sysctl parameters such as vm.dirty_background_ratio and vm.dirty_ratio allows the kernel to write dirty data to storage continuously in small chunks, avoiding sudden latency spikes during burst write operations.

Fabric Architecture and Network Performance

If the storage is not local, its performance is limited by the network bandwidth. Storage traffic should never be competing with application traffic on the same physical interfaces.

  • Dedicated Storage Networks: Stateful clusters should isolate internal Ceph replication traffic or SAN traffic from frontend pod networking. OpenShift’s implementation of Multus CNI allows pods and storage daemons to attach secondary network interfaces directly to isolated high performance VLANs or dedicated physical networks.

  • Jumbo Frames: Set the MTU to 9000 on all storage switches, host NICs, and CNI configurations to reduce CPU packet-processing overhead and maximise throughput for large block and sequential I/O.

  • Single Root I/O Virtualisation (SR-IOV): For the most extreme latency-sensitive workloads, SR-IOV allows pods to bypass the OpenShift Open vSwitch (OVS) software network layers completely, communicating directly with physical network interface cards to achieve near-bare-metal latency.

Node Failure Scenarios, Resiliency, and High Availability

In a live cluster, nodes will crash, network switches will fail, and hypervisors will reboot at unexpected times. Storage resiliency means making sure data is intact and applications recover quickly in the event of a failure.

Pod Failover and Storage Attachment Mechanics

When a worker node with a running stateful pod unexpectedly fails , OpenShift detects the failure using kubelet heartbeats . The node is marked NotReady after a defined timeout, and the scheduler starts to evict pods. This causes a complex reconciliation process for a stateful application attached to a RWO volume:

The original volume is still “attached” to the dead host at the cloud or SAN level. The volume single-node attachment contract enforced by the CSI driver prevents the new pod scheduled on a healthy node from mounting the volume immediately. The controller must detach from the CSI interface of the failed node before attaching to the new node.

If the Operating System of the dead node has hung and not released its storage target, the cluster can become paralysed in a volume attachment deadlock. To resolve this, platform operators must implement automated node fencing solutions, such as the Node Health Check Operator with Machine Deletion Remediation or self-node remediation with STONITH (Shoot The Other Node In The Head).

Fencing ensures that a bad node’s storage attachments are only disconnected after the node is totally powered off or isolated from the network. This enables the CSI controller to safely detach the volume and attach it to the remaining node without any risk of split-brain write operations or data corruption.

StatefulSet Orchestration Strategies

Deploying stateful applications as generic Kubernetes Deployments is an anti-pattern. Pods are interchangeable instances in a deployment. It starts new ones before it kills old ones. This causes instant conflicts in RWO volume attachment.

StatefulSets give you the structural control you need for persistent apps:

  • Predictable Identity: Pods get predictable, zero-indexed network hostnames and persistent ordinal identities that survive restarts and rescheduling.

  • StatefulSets allow you to decouple storage claim authoring from static manifests. In the StatefulSet spec, you can define the dynamic template, and OpenShift will automatically create a unique Persistent Volume Claim for each replica it creates. Claim zero is for pod zero claim one is for pod one

  • Ordered Execution: StatefulSets follow strict rolling sequence rules. During scaling and update operations, the controller updates pods one at a time, checking health on one before going to the next. This ensures quorum is maintained in distributed clusters like ZooKeeper, Raft, or Cassandra.

Day-2 Operations: Backup, Recovery, Disaster Recovery

Storage provisioning is only the beginning of the operational lifecycle. You have to have solid automated processes for volume resizing, snapshot orchestration, application-consistent backups and disaster recovery across multiple datacenters to ensure data integrity.

Dynamic Volume Expansion and Snapshots

In today’s business, storage volumes cannot remain stagnant at the original provisioned size. They need to be up and running 24/7. OpenShift supports dynamic volume expansion with CSI primitives:

  • Online Volume Expansion: An operator can increase the storage capacity request on a live Persistent Volume Claim. The node plugin will dynamically expand the underlying XFS or ext4 filesystem. The CSI controller will tell the backend storage array to resize the physical LUN. All of this happens while the container remains online and processing I/O. There is no universal support for shrinking volumes because of filesystem structural limitations and sizing strategies must always scale upward.

  • CSI Volume Snapshots VolumeSnapshot CRD is an abstraction for point-in-time storage snapshots for various storage hardware. Declarative manifests can be used by platform teams to trigger snapshots. Volume snapshots offer point-in-time recovery, but they only provide storage at a crash-consistent level. If the database has active memory buffers then snapshots need to be combined with freeze hooks at the application level (pre- and post-exec scripts that flush buffers to disc) to guarantee full transaction consistency upon restoration.

Application backup with OADP

The storage layer snapshots do not capture the complete application context. The storage volume is useless if the OpenShift namespace is completely deleted and the deployment manifests, secrets, configmaps, service accounts, and route definitions for the namespace are not restored.

OpenShift API for Data Protection (OADP) provides full application-aware backup and recovery:

  • OADP is built on top of the open source Velero engine and orchestrates backup of all Kubernetes resource configurations and the underlying persistent data volumes.

  • OADP calls CSI snapshot plug-ins to capture persistent storage, or uses low-level data movers such as Kopia or Restic to push file system level increments directly into object storage repositories.

  • Schedule OADP backups namespace-level and have disaster recovery plans automatically push encrypted metadata and volume archives to remote, immutability-locked S3 buckets.

Multi data center disaster recovery architectures

For mission critical operations that require survival against entire regional datacenter outages, OpenShift leverages Red Hat Advanced Cluster Management (ACM) in combination with ODF Disaster Recovery:

  • Metro-DR (Synchronous Replication): Used between datacenters interconnected by ultra-low latency network infrastructure (typically less than 5ms round trip time). Each write is synchronously replicated by the storage backend to both sites before returning an acknowledgement to the application. If a datacenter goes down, workloads fail over to the secondary site with a Recovery Point Objective (RPO) of zero meaning no data loss.

  • Regional-DR (Asynchronous Replication): Used across geographically remote regions separated by hundreds or thousands of miles where the latency of the speed of light makes synchronous writes infeasible. The platform asynchronously replicates block and file deltas to the secondary site on a periodic basis. In a disastrous regional failure, applications are recovered in the surviving region with an RPO defined by the replication interval (usually minutes) and a low Recovery Time Objective (RTO).

Creating, operating, and restoring distributed storage systems in business settings requires ongoing systemic expertise. For engineers, architects of cloud-based systems, and DevOps experts who want to develop and run these necessary systems for their businesses, there are openshift redhat training technology, which provides direct hands-on experience needed to run cloud-based applications safely.

Why Choose Us

Master Your Future with OnlineITGuru

We don't just provide courses; we build careers. From expert-led live training to dedicated placement support, discover why thousands of professionals trust us for their digital transformation journey.

200+

Partner Companies

$120K

Highest Package

75%

Average Hike

98%

Placement Rate

Reliable Career Partners

Google
Microsoft
Amazon
Meta
Netflix
Apple