April 2003
Intermediate to advanced
1632 pages
43h 48m
English
The two most common metrics used to measure fault tolerance and avoidance are the following:
Mean time to failure (MTTF) The mean time until the device will fail
Mean time to recover (MTTR) The mean time it takes to recover once a failure has occurred
Although a great deal of time and energy is often spent trying to lower the MTTF, it’s important to keep in mind that even if you have a finite failure rate, if your MTTR is zero or near zero, this may be indistinguishable from a system that hasn’t failed. Downtime is generally measured as MTTR/MTTF, but because it can be prohibitively expensive to increase MTTF beyond a certain point, you should spend both time and resources on managing and reducing ...
Read now
Unlock full access