
Ceph ensures high availability by replicating data threefold across distinct nodes and failure domains, enabling self-healing, and maintaining cluster brain with an odd-numbered quorum of monitors.
Build a three-node CIF cluster with replication and quorum-enabled monitors. Prepare AlmaLinux 9, static IPs, NTP, and passwordless SSH, then deploy, provision osds, and monitor via the CIF dashboard.
Design a two-network Ceph storage cluster with a public network for client traffic and a cluster network for OSD-to-OSD replication, heartbeats, and recovery.
Discover how Ceph's core components (OSD, MON, MGR, MDS, and RGW) work together to store data, manage the cluster map and dashboards, and enable CephFS and S3-compatible object storage.
Trace the flow of data from a client to OSDs, through the cluster map and crash placement groups, with parallel replication and MTU 9000, no hardware RAID.
Update AlmaLinux 9.7 servers, install epel, verify selinux and firewalld, verify storage with lsblk, and configure dual networks for a three-node Ceph deployment.
Configure Ceph network now by setting static hostnames, building local DNS with /etc/hosts, and securing Ceph traffic with firewalld trusted zones across three nodes.
Discover how to set up Ceph storage by scanning hosts for blank disks, converting them into active OSDs, and delivering a healthy, highly available Ceph cluster with nine OSDs.
Map a Ceph RBD image to a Linux client and to Proxmox, demonstrating manual mapping on AlmaLinux and automated storage provisioning in a Proxmox environment.
Connect Proxmox to a Ceph RPD pool, initialize it, and configure a Proxmox client with authentication; add Ceph storage and run VMs in RPD with replicated data for high availability.
Deploy and manage a high-availability CephFS shared file system using MDS containers across three nodes, creating data and metadata pools for scalable, POSIX-compliant storage.
Deploy and configure the RADUS gateway to enable S3-compatible object storage on a Ceph cluster, with multi-node RGW deployment, backend pools, and access key and secret key credentials.
Install the AWS CLI on a separate AlmaLinux server, configure credentials, and use an S3 backend to upload and verify a file in Ceph object storage via the RADUS gateway.
Mount s3 buckets as folders with s3fs to let legacy apps save to a Linux directory, while files are backed by the CIF storage cluster through an HAProxy load balancer.
This lab demonstrates end-to-end Ceph storage workflows: create and mount an rbd block device, mount a shared network folder, and upload data to an s3-compatible object store.
Control Ceph rebalancing speed to protect client IO during peak usage by limiting recovery traffic with osdmax backfill 1 and osd recovery slip 0.1.
Learn how modern, containerized Ceph enables zero-downtime upgrades via a ruling upgrade that updates monitors, managers, MDS, and OSD one at a time using CephADM.
Learn capacity planning for Ceph storage by calculating raw versus usable capacity, applying replication and erasure coding overhead, and enforcing the safe 85% rule to prevent outages.
Discover how Ceph performance hinges on the interaction of network, CPU, RAM, and disks, and distinguish IOPS from throughput while examining tail latency.
Tune osd performance by using BlueStore ram cache to reduce latency, with default 4 gb per osd, adjustable to 8 gb, and auto-tuning options.
Keep swap enabled as an emergency backup for Ceph clusters, but set swappiness to 1 and reserve memory with mainfree-kb to prevent latency spikes from swap.
Enable data-at-rest encryption in Ceph by deploying disks with the dash dash encrypted option; Ceph uses LUX and dm-crypt, with the monitor storing the key to unlock OSD on boot.
Apply a security checklist for Ceph storage, enforcing defense in depth with a dual network, restricted firewalls, RADOS gateway behind a reverse proxy, quotas, bucket policies, and encryption at rest.
Learn to diagnose common Ceph issues: recent crashes, OST full, degraded or misplaced pgs, and slow requests. Use Ceph health details and targeted commands to recover and maintain cluster health.
Master snapshot-based recovery in Ceph with Proxmox by creating and restoring VM snapshots in Ceph RPD storage and recovering individual files from CephFS via the hidden .snap directory.
Are you ready to master the future of Enterprise Distributed Storage?
In today’s data-driven world, traditional SAN and NAS hardware arrays are too expensive, too difficult to scale, and represent massive single points of failure. The modern enterprise requires storage that is software-defined, infinitely scalable, and completely self-healing.
Welcome to Ceph Storage.
Used by the world's largest telecommunications companies, cloud providers, and research labs, Ceph is the undisputed gold standard for open-source distributed storage.
About This Course:
Taught by a Senior Data Center System Administrator, this course skips the useless theory and focuses entirely on real-world, production-ready architecture. You will not just learn what Ceph is; you will build a highly available, multi-node cluster from scratch using cephadm.
Through intensive, step-by-step hands-on labs, you will learn how to:
Deploy and Bootstrap a highly available Ceph Storage Cluster.
Provision Block Storage (RBD) for Virtual Machines and databases.
Create Shared File Systems (CephFS) for enterprise file sharing.
Configure Object Storage (RGW) to create your own S3-compatible cloud storage.
Deploy HAProxy Load Balancing: Eliminate single points of failure and ensure true high availability for your S3 (RGW) storage endpoints.
Scale Dynamically: Add new nodes and disks (OSDs) without a single second of downtime.
Implement Disaster Recovery: Design Multi-Site asynchronous mirroring to survive a total data center failure.
Defend Against Ransomware: Use S3 Lifecycle Policies and Object Lock (Immutability).
Troubleshoot Like a Pro: Analyze daemon logs, manage CRUSH maps, and seamlessly recover from catastrophic hard drive failures.
Why take this course?
As an IT Architect, your ability to guarantee data integrity and uptime is your most valuable skill. This course is designed to take you from a complete beginner to a confident Storage Administrator capable of designing, deploying, and rescuing Petabyte-scale environments.
Stop relying on expensive proprietary hardware. Enroll today, and start building the ultimate software-defined data center.