Monitoring the node status
As you can guess, perhaps the first thing that you always need to check is the status of each node—whether they are online or offline. Otherwise, there is little point in proceeding with further availability and performance analysis.
If you have a network management system (such as Zabbix or Nagios) server, you can easily monitor the status of your cluster members and receive alerts when they are unreachable. If not, you must come up with a supplementary solution of your own (which may not be as effective or errorproof) that you can use to detect when a node has gone offline.
One such solution is a simple bash script (we will name it pingreport.sh, save it inside /root/scripts, and make it executable with chmod +x /root/scripts/pingreport.sh ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access