Architecture Security & Reliability

mean time between failures (MTBF)

/ M-T-B-F /

Suppose you watched a fleet of machines for a long time and wrote down how long each one ran before it broke. Average those running stretches and you get the mean time between failures, or MTBF — the typical length of an uptime spell. It is the headline reliability number you see on disks, power supplies, and servers, usually quoted in hours: an MTBF of 1,000,000 hours means that, on average, a working stretch lasts a million hours before the next failure.

The single most important thing to understand is what MTBF is an average over. It is not a promise about one unit's lifetime. A disk with a 1.2-million-hour MTBF will not run for 137 years; that number is computed from a large population over a short test and really describes a failure rate. The relationship is simple: failure rate is roughly 1 / MTBF. So 1,000 drives each with a 1,000,000-hour MTBF together produce about 1,000 / 1,000,000 = one failure per 1,000 hours of fleet operation — roughly one every six weeks. MTBF tells you how often to expect trouble across many units, not how long any one will last.

MTBF connects directly to availability. With MTTR the mean time to repair, availability = MTBF / (MTBF + MTTR), so a big MTBF and a small MTTR both push availability up (see availability). A subtle honesty: real components do not fail at a constant rate over their whole life — there is a 'bathtub curve' with extra early failures from manufacturing defects and extra late failures from wear-out, with a flat low-rate middle. MTBF figures describe that flat middle and quietly assume you are operating there, which is usually but not always true.

A datacenter runs 10,000 disks, each rated MTBF = 1,000,000 hours. Expected failures per hour = 10,000 / 1,000,000 = 0.01, i.e. about one disk failure every 100 hours, or roughly 88 failures per year. This is why large fleets always plan for routine disk replacement — failure is not an 'if' but a steady 'when'.

Scale turns a huge MTBF into frequent failures: 10,000 units make a million-hour MTBF mean a failure every 100 hours.

Never read MTBF as a guaranteed lifetime of one unit. It is a fleet-level rate; one disk with a million-hour MTBF can still die tomorrow, and across thousands of disks some will.

Also called
MTBFmean time to failureMTTF平均無故障時間