MTBF is one of the most misused metrics in mechanical engineering — a machine with a 10,000-hour MTBF does not mean it will run for 10,000 hours before failing; it means there is a 63% probability it will fail within 10,000 hours of operation.
Reliability engineering provides quantitative tools for predicting, measuring, and improving the probability that a machine will perform its required function for a specified time under specified conditions. For mechanical designers, a working knowledge of the key concepts — bathtub curve, MTBF, exponential distribution, series and parallel reliability — enables meaningful dialogue with reliability engineers, better design decisions, and more realistic life predictions. This guide covers the fundamentals without requiring deep statistical background.
The Bathtub Curve
The bathtub curve is a plot of failure rate (hazard rate, λ) versus time for a population of components or systems. It has three distinct regions:
Infant Mortality (Burn-In) Region: High initial failure rate that decreases with time. Failures in this region result from manufacturing defects, improper assembly, material flaws, and installation errors that reveal themselves under early operating conditions. Design and process improvements reduce infant mortality: tighter incoming inspection, burn-in testing (running components under stress before delivery), and improved manufacturing process control. Many consumer electronics and high-reliability industrial components are factory burn-in tested to eliminate infant mortality units before delivery.
Useful Life (Random Failure) Region: Roughly constant low failure rate. Failures in this region are random — they cannot be predicted for specific units but occur at a statistically constant rate for the population. This is the operating regime that MTBF calculations describe. Design improvements that reduce the random failure rate include using higher-rated components (operating well below rated capacity), improving environmental protection (sealing, thermal management), and using redundant systems.
Wear-Out Region: Increasing failure rate as components approach the end of their mechanical life through fatigue, wear, corrosion, or material degradation. Design life is defined by the onset of this region. Preventive maintenance and replacement schedules are designed to remove components before they enter the wear-out region.
MTBF: Definition and Correct Interpretation
Mean Time Between Failures (MTBF) is the average time between successive failures of a repairable system. For a population of N systems operating over time T with f failures: MTBF = (N × T) / f. For a single component following the exponential failure distribution, MTBF = 1/λ, where λ is the constant failure rate.
The critical insight about MTBF: under the exponential distribution assumption (constant failure rate — applicable in the useful life region), the probability of survival to time t is: R(t) = e^(−t/MTBF) = e^(−λt). At t = MTBF: R(MTBF) = e^(−1) = 0.368. This means there is only a 36.8% probability of surviving to the MTBF — or equivalently, a 63.2% probability of failing before MTBF is reached.
To find the time at which reliability equals a specified target R: t = −MTBF × ln(R). For 90% reliability: t = −MTBF × ln(0.9) = 0.105 × MTBF. A system with 10,000-hour MTBF has 90% reliability for only 1,053 hours. This is why MTBF alone is an incomplete specification — always specify the reliability target (R) and the time period (mission time, t) together.
Failure Rate and Its Relationship to MTBF
Failure rate (λ) is expressed in failures per hour, or more conveniently in FIT (Failures In Time) units where 1 FIT = 1 failure per 10⁹ component-hours. Electronic components are often specified in FIT; mechanical components are more commonly specified in MTBF or B10 life.
MTBF data for mechanical components is available from sources including:
• MIL-HDBK-217 (Electronic Components, including some mechanical) — US military reliability prediction handbook
• NSWC-11 (Mechanical Components Reliability) — US Navy handbook for mechanical component reliability prediction, covering bearings, seals, springs, gears, and actuators
• IEC 62061 / ISO 13849 — functional safety reliability data for safety system components
• Bearing manufacturers (SKF, NSK, FAG) — B10 life data, which can be converted to reliability at specific operating times
Series vs Parallel Reliability
Real systems consist of multiple components whose failure modes interact. Two fundamental configurations:
Series System: All components must function for the system to function (most common in mechanical systems — failure of any component causes system failure). System reliability: R_system = R₁ × R₂ × R₃ × … × R_n. This means system reliability is always lower than the least reliable component. With ten components each at 99% reliability: R_system = 0.99¹⁰ = 0.904 (90.4%). With 100 components at 99%: R_system = 0.99¹⁰⁰ = 0.366 (36.6%). This shows how complex systems can have much lower reliability than their individual components.
Parallel (Redundant) System: The system functions if any one of the parallel components functions. System reliability: R_system = 1 − (1−R₁) × (1−R₂) × … × (1−R_n). Two components each at 90% reliability in parallel: R_system = 1 − (0.1)(0.1) = 0.99 (99%). Redundancy dramatically improves reliability but adds cost, weight, and complexity. Common in safety systems (dual braking circuits, redundant power supplies, emergency stops).
B10 Life vs MTBF: Two Approaches to Life Specification
The B10 life (also L10 life for bearings) is the time by which 10% of a population of components will have failed — equivalently, 90% survive to B10. This is a percentile specification rather than a mean specification, and it is more useful for design purposes because it directly answers “at what point should I schedule replacement?”
For bearings (ISO 281), the basic rating life L10 is calculated from the dynamic load rating C, equivalent dynamic load P, and exponent p: L10 = (C/P)^p million revolutions, or in hours: L10h = L10 × 10⁶ / (60 × n), where n is speed in rpm. This is the industry-standard design life calculation for rolling element bearings from SKF, NSK, FAG, NTN, and other major suppliers.
MTBF and B10 life can be related under the exponential distribution assumption: B10 = 0.105 × MTBF (since R(B10) = 0.9 = e^(-B10/MTBF)). However, mechanical wear-out typically follows a Weibull distribution rather than exponential — the Weibull shape parameter (β) characterizes the failure mode: β < 1 (infant mortality), β = 1 (exponential/random), β > 1 (wear-out). Bearing manufacturers use β ≈ 1.1–3 for rolling element bearing life calculations.
Reliability Block Diagram (RBD)
A Reliability Block Diagram models a system’s reliability by showing which components are in series and which are in parallel. Building an RBD requires:
1. Define system success (what must function for the system to perform its purpose)
2. Identify all failure modes that cause system failure
3. Determine which failures are independent and which are common-cause
4. Arrange blocks in series (all required) and parallel (redundant alternatives)
5. Assign reliability values to each block from component data or reliability predictions
Software tools for RBD analysis include Relex (Windchill Quality Solutions), Isograph Availability Workbench, and open-source options like OpenReliability. For simpler systems, the series-parallel formulas above suffice for hand calculation.
Design Actions to Improve Reliability
Derating: Operating components well below their rated capacity reduces failure rate. For electronic components, operating at 50% of rated power reduces failure rate by a factor of 2–4. For mechanical components, the analogous approach is selecting bearings for L10h life of 3–5× the target service life, or using safety factors of 3–4 on fatigue.
Simplification: Fewer components mean fewer failure opportunities. The series reliability equation shows that eliminating even one 99%-reliable component improves system reliability by approximately 1%. Eliminating a 95%-reliable component improves system reliability by 5%. Apply DFA and part count reduction systematically.
Environmental protection: Sealing (IP65 or higher for outdoor/washdown environments), thermal management (keeping temperatures within rated ranges), vibration isolation, and corrosion protection all reduce the environmental stress that accelerates failure in the random failure region.
Failure mode elimination through design: FMEA (Failure Mode and Effects Analysis) identifies failure modes and their effects, then drives design changes that eliminate high-criticality failure modes. Per IEC 60812, FMEA assigns severity (S), occurrence (O), and detection (D) ratings to each failure mode, producing a Risk Priority Number (RPN = S × O × D) that prioritizes improvement actions.
Predictive maintenance enablers: Design in condition monitoring points — accelerometer mounting pads on bearing housings, oil sampling ports, temperature monitoring provisions — that allow maintenance to track degradation and intervene before failure occurs. The goal is transforming random failures into managed, predictable maintenance events.
Conclusion
Reliability engineering provides a rigorous quantitative framework for what engineers often address qualitatively as “robustness” or “overdesign.” Understanding MTBF as a reciprocal of failure rate (not a guaranteed life), using B10 life for maintenance interval planning, and applying series reliability analysis to identify the weakest links in complex systems are practical skills that improve design decisions. The most effective reliability improvement actions — simplification, derating, environmental protection, and failure mode elimination through FMEA — are all things the mechanical designer controls directly.



コメント