FMEA (Failure Mode and Effects Analysis) is either the most valuable reliability tool in a mechanical engineer’s toolkit or a bureaucratic paperwork exercise — depending entirely on how it’s conducted. Done properly, it surfaces failure modes before they happen and drives design improvements that save warranty costs, customer downtime, and sometimes lives. Done as a box-checking activity, it produces documents no one reads and prevents nothing.
This guide provides a step-by-step process for conducting a meaningful Design FMEA (DFMEA) for mechanical systems, using a conveyor drive system as the worked example throughout. I’ll cover the failure mode identification methodology, severity/occurrence/detection rating scales and common rating mistakes, RPN calculation and its limitations, action priority determination, and the key differences between Design FMEA and Process FMEA. This is a practical guide for engineers who want to conduct FMEA that actually improves their designs, not just check a compliance box.
- FMEA Fundamentals: What It Is and What It Isn’t
- Step 1: Define the System and Create the Function-Requirement Matrix
- Step 2: Identify Failure Modes
- Step 3: Identify Effects and Causes
- Step 4: Assign Severity, Occurrence, and Detection Ratings
- Step 5: Calculate RPN and Determine Action Priority
- Step 6: Define Corrective Actions and Track Closure
- Design FMEA vs. Process FMEA
- Conclusion
FMEA Fundamentals: What It Is and What It Isn’t
A Design FMEA is a structured analysis of each element of a design to identify: (1) the ways it can fail (failure modes), (2) what happens to the system when it fails (effects), (3) why it fails (causes), and (4) what existing controls detect or prevent the failure. From this analysis, a Risk Priority Number (RPN) is calculated for each failure mode, and design improvements are prioritized based on this risk assessment.
FMEA is not a one-time analysis completed at a project’s end — it’s a living document that should be initiated during conceptual design and updated throughout the design development process. The earlier in the design process that failure modes are identified, the cheaper they are to address. A failure mode identified during concept selection costs far less to resolve than the same failure mode discovered during field operation.
FMEA is most effective as a team activity, not a solo exercise. The design engineer brings structural knowledge; the manufacturing engineer brings process capability context; the service engineer brings field failure data; the test engineer brings validation data. A FMEA conducted by one person in isolation misses the cross-functional knowledge that makes it valuable.
Step 1: Define the System and Create the Function-Requirement Matrix
Before identifying failure modes, define what the system must do. For our example — a conveyor drive system — the primary functions are: (1) transmit motor torque to the conveyor belt drive roller at the required torque and speed; (2) maintain constant belt speed under variable load; (3) withstand starting torque transients (typically 2–3× running torque); (4) operate reliably for a 10-year design life at specified duty cycle; (5) allow belt tension adjustment.
Define the system boundary: what is included in this FMEA (motor coupling, gearbox, drive shaft, bearings, drive sprocket, chain, driven sprocket) and what is excluded (the motor itself, which has its own FMEA; the belt conveyor structure; the control system). Clear system boundary definition prevents both double-counting (analyzing the same failure mode at two levels) and gaps (missing failure modes at system boundaries).
Step 2: Identify Failure Modes
For each component, identify all physically possible failure modes — not just the most likely ones. Use functional language: “fatigue fracture of the shaft,” “wear of gear teeth beyond allowable pitch circle modification,” “seizure of bearing inner race,” “corrosion of chain link pins,” “loosening of coupling fasteners.” Each failure mode should describe a specific, physical degradation or failure of the component.
Common sources for failure mode identification: historical field failure data for similar components (FMEA is most powerful when past failures are the seed data), physics of failure analysis (what are the dominant stress modes for this component?), previous FMEAs for similar systems, supplier reliability reports, and engineering standards (AGMA gear failure mode classification, for example).
For the conveyor drive shaft, identified failure modes might include: fatigue fracture at the keyway stress concentration; wear of the shaft surface at the oil seal interface; fretting corrosion under the coupling hub bore; excessive deflection under transient load. Each of these is a specific failure mode with a different cause, effect, and detection path — they must be analyzed separately.
Step 3: Identify Effects and Causes
For each failure mode, identify: the effect on the next higher assembly and on the end user. A drive shaft fatigue fracture (failure mode) → loss of torque transmission → conveyor stops (effect on system) → production downtime, potential product spillage if loaded (effect on end user). Effects should be described in terms the customer or end user would recognize, not just engineering language.
Causes are the specific physical root causes of each failure mode. For drive shaft fatigue fracture: cyclic bending stress exceeding material endurance limit due to shaft misalignment; stress concentration at keyway exceeding design intent due to burr left by machining process; inadequate shaft diameter for transient overload torque. Multiple causes should be listed for each failure mode and analyzed separately — different causes lead to different prevention and detection strategies.
Step 4: Assign Severity, Occurrence, and Detection Ratings
The RPN (Risk Priority Number) = Severity × Occurrence × Detection, where each factor is rated on a 1–10 scale. The rating scales matter enormously for FMEA credibility — inconsistent or poorly calibrated ratings produce an analysis that prioritizes the wrong risks.
Severity (S): Rates the consequence of the failure mode on the end user. Rating scales vary by standard (AIAG FMEA-4, VDA FMEA, AIAG-VDA FMEA 1st Edition) but typically: 1–2 = no noticeable effect; 3–4 = minor degradation, workaround available; 5–6 = moderate performance loss, customer dissatisfied; 7–8 = significant loss of function, potential machine damage; 9 = safety issue with warning; 10 = safety issue without warning or regulatory non-compliance. Severity is rated for the effect, not the failure mode itself — a shaft fracture affecting only conveyor speed (Severity 5) is rated differently from a shaft fracture causing mechanical hazard to personnel (Severity 9–10).
Occurrence (O): Rates the likelihood of the cause occurring within the design life. Not the probability after all prevention controls are in place — that’s captured separately as detection. Typical scale: 1 = almost impossible (<1 in 1,000,000); 3 = remote (1 in 100,000); 5 = moderate (1 in 2,000); 7 = high (1 in 100); 9 = very high (>1 in 10). Rating occurrence without supporting data (field failure rates, similar part experience, fatigue life calculations) is a common FMEA credibility problem. Where hard data isn’t available, document the basis for the estimate.
Detection (D): Rates the ability of current controls to detect the failure mode or its cause before it reaches the end user. Lower D = better detection. 1 = current controls will almost certainly detect it before shipment; 5 = moderate chance of detection; 10 = no current control can detect it. Common mistake: rating detection too optimistically. A visual inspection is a Detection 7–8, not a Detection 2–3 — visual inspection misses many failure modes and depends heavily on inspector attention. Automated dimensional inspection with 100% coverage is Detection 2–3.
Step 5: Calculate RPN and Determine Action Priority
RPN = S × O × D, ranging from 1 (minimum risk) to 1,000 (maximum risk). A common and problematic practice: using an RPN threshold (e.g., RPN > 100 requires action) as the primary prioritization tool. The limitation of RPN as a sole criterion: an S=10, O=1, D=1 failure mode has RPN=10 — the detection is perfect and occurrence is nearly impossible, but the severity is catastrophic. Should this truly receive lower priority than an S=3, O=6, D=7 item with RPN=126? The answer is clearly no for safety-critical failures.
Better practice: high-severity failure modes (S ≥ 9) require action regardless of RPN, with special priority given to reducing severity (redesign to change failure effect) or occurrence (add prevention controls). For lower-severity items, RPN is a reasonable triage tool for prioritizing the remaining action list. The AIAG-VDA FMEA (published 2019) replaced the single RPN with an Action Priority (AP) system that separately considers severity, occurrence, and detection thresholds rather than multiplying them — addressing some of the known limitations of RPN.
Step 6: Define Corrective Actions and Track Closure
For each high-priority item, define corrective actions that reduce risk by: (1) Reducing severity — redesign to change the failure effect (e.g., add a secondary restraint so that shaft fracture doesn’t cause conveyor discharge); (2) Reducing occurrence — add prevention controls or redesign to eliminate the cause (increase shaft diameter to provide higher fatigue safety factor, specify surface finish to reduce stress concentration at keyway); (3) Improving detection — add inspection or monitoring that catches the failure mode before it reaches the customer (add vibration monitoring that detects imbalance before bearing failure progression).
Assign each action an owner and a due date. Actions without ownership and deadlines don’t get completed. Re-evaluate the RPN (or Action Priority) after the corrective action is implemented to confirm that risk has been reduced to an acceptable level. This re-evaluation is the step most often skipped in practice — without it, the FMEA is a one-pass analysis that doesn’t confirm risk reduction.
Design FMEA vs. Process FMEA
Design FMEA (DFMEA) analyzes the design of a product — failure modes caused by design deficiencies, material selection, and specification errors. It’s owned by the design engineering team.
Process FMEA (PFMEA) analyzes the manufacturing process — failure modes caused by process variations, operator errors, and equipment capability limitations that can produce nonconforming parts. It’s owned by the manufacturing/process engineering team.
The two FMEAs are linked: a DFMEA that identifies a critical characteristic (dimension or property that significantly affects function) should trigger a PFMEA analysis of the manufacturing process that produces that characteristic. If the DFMEA identifies that shaft surface finish at the seal interface must be Ra ≤ 0.4 μm for reliable sealing life, the PFMEA should analyze the grinding process that produces this surface and ensure adequate detection controls are in place for surface finish variation.
Conclusion
A well-conducted FMEA is one of the most effective tools for improving design reliability before hardware is built. The process is only as good as the failure mode identification quality, the calibration of severity/occurrence/detection ratings against real data, the honesty of action priority assignment (including for high-severity, low-RPN items), and the discipline of tracking actions to closure and re-evaluating residual risk. Applied properly to the conveyor drive system or any other mechanical design, FMEA systematically surfaces the failure modes that would otherwise be discovered through field failures — a much more expensive way to learn the same information.



コメント