Chapter 3. Storage virtualization 139
Additionally, each server has a dedicated hang-detection thread that causes it
to restart should it become unresponsive. Operating system crashes or hangs
are detected by the resident IBM Director agent on each server, which
reboots the local machine if needed. The watchdog process restarts
automatically as part of the boot process and in turn restarts the MDS.
Hard faults
Hard faults are server failures for which recovery requires administrative
intervention. Hard faults are typically associated with hardware failures. They
have a greater impact than soft faults and require at least a machine reboot and
possibly physical maintenance for recovery.
The SAN File System detects hard faults by way of a hear ...