Recover CEPH mon cluster after host failure

I recently had a Proxmox node’s root disk fail. I was able to bring it back up fairly quickly with a new Proxmox install but struggled with ceph mon process for that node. I encountered an issue where the mon process ate up all the RAM on the host. I left the cluster in this state overnight and by the morning the mon cluster had lost its mind and read/write was broken for the entire CEPH cluster.

The fix was to stop mon processes on all nodes. Then pick a healthy node and manually bring it back up after manually removing all other nodes from its config. Then on every node, manually destroy and re-create the ceph mon configuration. Here are the links I used to help me figure it all out:

https://forum.proxmox.com/threads/ceph-cluster-can-not-start-monitors.34689

https://forum.proxmox.com/threads/ceph-cannot-create-monitor-monitor-address-xxx-already-in-use-500.130932

https://forum.proxmox.com/threads/ceph-how-to-delete-dead-monitor.61172

Rescue unhealthy mon cluster

Taken from https://docs.ceph.com/en/latest/rados/operations/add-or-rm-mons/#removing-monitors-from-an-unhealthy-cluster

Stop all ceph-mon daemons on all Monitor hosts:

ssh {mon-host}
systemctl stop ceph-mon.target

Repeat this step on every Monitor host.

Identify a surviving Monitor and log in to the Monitor’s host:

ssh {mon-host}

Extract a copy of the monmap file by running a command of the following form:

ceph-mon -i {mon-id} --extract-monmap {map-path}

Here is a more concrete example. In this example, hostname is the {mon-id} and /tmp/monmap is the {map-path}:

ceph-mon -i `hostname` --extract-monmap /tmp/monmap

Remove the non-surviving or otherwise problematic Monitors:

monmaptool {map-path} --rm {mon-id}

For example, suppose that there are three Monitors—mon.a, mon.b, and mon.c—and that only mon.a will survive:

monmaptool /tmp/monmap --rm b
monmaptool /tmp/monmap --rm c

Inject the surviving map that includes the removed Monitors into the monmap of the surviving Monitor(s):

ceph-mon -i {mon-id} --inject-monmap {map-path}

Continuing with the above example, inject a map into Monitor mon.a by running the following command:

ceph-mon -i a --inject-monmap /tmp/monmap

Start only the surviving Monitors.

Verify that the Monitors form a quorum by running the command ceph -s.

The data directory of the removed Monitors is in /var/lib/ceph/mon: either archive this data directory in a safe location or delete this data directory. However, do not delete it unless you are confident that the remaining Monitors are healthy and sufficiently redundant. Make sure that there is enough room for the live DB to expand and compact, and make sure that there is also room for an archived copy of the DB. The archived copy can be compressed.

Additional procedures I took to regain full health

Remove dead mons from ceph.conf

vi /etc/pve/ceph.conf

Remove them both from mon_host line and [mon.<hostname>] stanzas

Remove systemd unit

systemctl stop ceph-mon@<hostname>.service
rm /etc/systemd/system/ceph-mon.target.wants/ceph-mon@<hostname>.service
systemctl daemon-reload

Remove ceph mon files

rm -r /var/lib/ceph/mon/ceph-<hostname>/

Re-create mon process

pveceph mon create

Reboot host. I only had to do this on one stubborn node that still wouldn’t start the mon process. It came back up fine after reboot.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.