Failure Analysis¶
A retrospective engineering investigation that uses a failed physical item's evidence and service context to test explanations of how and why it failed.
Core Idea¶
Failure analysis, as used here, is the retrospective engineering investigation of an observed failure of a physical item or structure. The investigator begins with what function the item was meant to perform and what departure was observed, then brings together available physical traces and context—design, manufacture, loading, environment and service history—to compare possible mechanisms and contributing conditions. The output is a bounded technical explanation: what the evidence supports, what it rules out, and what remains uncertain. NASA JPL's component history and mechanism distinction and NIST's structural-evidence and uncertainty practices support this common pattern across unlike settings.[1][2]
The method travels beyond metallurgy without becoming a universal theory of every kind of failure. NASA's Jet Propulsion Laboratory describes failure analysis after an ASIC-related malfunction and cautions that the observed effect may arise in surrounding circuitry or operating conditions rather than the chip itself. NASA Kennedy's materials laboratory investigates both mechanical and electrical components. NTSB's bridge-collapse inquiry shows the same evidential logic at structural scale: a physical failure had to be related to the bridge's design and loading history, not simply labeled by a fracture appearance.[1][3][4]
This is not Failure Mode and Effects Analysis (FMEA). FMEA's live entry is a forward-looking design/risk method that enumerates possible failure modes and ranks their effects before they occur; the present method works backward from a particular observed failure. The live FMEA node nevertheless includes “failure analysis” as an alias. That is an exact lexical collision, so the two identities must be disentangled before this draft can be promoted; no canonical alias change is made here.
Structural Signature¶
Sig role-phrases: observed departure from expected function → physical item and design/service context → preserved traces and records → alternative failure-mode/mechanism/condition hypotheses → discriminating examinations → bounded technical finding; recommendations and attribution optional.
- Observed departure. A component, device or structure failed to meet an expected function or limit. Without an actual case, enumerating what could fail belongs to prospective risk analysis, not this retrospective method.[1][4]
- Typed physical carrier and history. The failed item is located in its design, construction or operating setting. The same visible damage can have different explanations under different loads or environments. NASA JPL accordingly looks beyond an ASIC to neighboring circuit components and operating conditions; NIST examines building design and maintenance records.[1][2]
- Physical and contextual evidence. Parts, surfaces, measurements, images, records and observations can each constrain an explanation. NIST's structural investigators use debris, documents, footage, accounts, testing and models where appropriate; no one material mark is a universal master key.[2]
- Competing hypotheses at distinct causal levels. A failure mode describes how function was lost; a mechanism proposes the physical process; a contributing design, process or service condition explains why the mechanism was possible then. Those levels can interact and should not be collapsed into a single label.[1][4]
- Discriminating examination. Tests and comparisons should bear on live alternatives. They may be nondestructive, partly destructive or destructive depending on the question, specimen and preservation needs. JPL provides a local ASIC workflow; NIST describes hypothesis-driven mixed testing and quantified uncertainty in structural investigations.[1][2]
- Bounded finding. The analyst states the best-supported technical cause or causes and the limits of the evidence. A recommendation may follow, as in JPL or NTSB practice, but a corrective action or legal assignment of responsibility is a separate downstream purpose, not a defining ingredient.[1][4]
What It Is Not¶
It is not FMEA under a shorter name. FMEA asks, before or alongside design and operation, what failure modes are possible and how to prioritize their risks. A particular cracked member, failed circuit or collapsed bridge supplies a different starting condition: observed damage and history that must be explained. The live FMEA alias collision is therefore a vocabulary repair issue, not evidence that these two methods are identical.
It is not merely naming the visible damage. A shorted circuit, ruptured plate or broken fastener can be the symptom or mode while the cause lies in a manufacturing defect, design margin, loading history, environment or interaction among these. JPL explicitly allows an apparent ASIC failure to originate outside the ASIC; NTSB separated a gusset-plate capacity failure from design and load contributors.[1][4]
It is not a universal laboratory sequence. Preserving evidence often matters, but “nondestructive before destructive” is not a constitutive law of the abstraction. NIST describes combining the two forms and managing uncertainty, while JPL's ordered categories are explicitly an ASIC example methodology. The relevant question is which test distinguishes hypotheses without needlessly losing other evidence.[2][1]
It is not a legal judgment or automatic corrective-action program. JPL and NTSB include recommendations in their own institutional workflows, but technical explanation has a different warrant from assigning blame. This entry is limited by editorial scope to the physical engineering settings evidenced here; it does not claim that software-only incident review cannot also be called failure analysis in another practice.[1][4]
Scope of Application¶
The direct scope is engineered physical artifacts and structures whose failure leaves inspectable traces or records: metal parts, electronic devices, mechanical assemblies and structural systems. NASA Kennedy describes one institutional laboratory spanning mechanical and electrical components and using varied characterization methods. JPL's ASIC guide illustrates an electronics-specific version; NTSB's I-35W bridge investigation illustrates structural analysis where the carrier and evidence are much larger and more heterogeneous.[3][1][4]
The narrow original Wikipedia candidate was Metallurgical failure analysis. This entry is broader because independent official sources support the shared retrospective method outside metals; metallurgy remains a specialist instantiation, not a deleted identity. The broader method here remains bounded by physical evidence and engineered function. A purely social, financial or software failure may invite analogous inquiry, but does not automatically instantiate this particular physical-engineering entry.[1][3][2]
Clarity¶
Begin by separating four questions: What was supposed to happen? What happened instead? By what physical mechanism? Under what enabling conditions? The first two define failure; the last two form causal hypotheses. NTSB's bridge finding is an instructive distinction: a structural member's inadequate capacity under actual loads was linked to design and review history. A report that said only “the gusset plate failed” would stop before the explanation.[4]
Also state the strength and limit of the finding. NIST notes that structural inquiries compare alternatives and quantify uncertainty at measurement and analysis steps. Some causes can be excluded; others can be identified only as likely given missing specimens or records. “Root cause” should not become a ritual demand for one final culprit when evidence supports several interacting conditions or a narrower finding.[2]
Manages Complexity¶
A failed system can present many traces and many plausible stories. The method organizes them by carrier, context and hypothesis: which observations support a proposed mechanism, which contradict it, and which further examination would genuinely discriminate? In JPL's ASIC frame, checking surrounding components and operating conditions prevents premature commitment to the chip as the cause. In a structural inquiry, design documents and load history can explain why a particular member failed at a particular time.[1][4]
That organization reduces premature closure but does not make every investigation exhaustive. NIST explains that large collapse investigations may consider many hypotheses and use physical tests plus models; JPL notes that cost limits which failed parts receive its more detailed analysis. A finite investigation should preserve an auditable trail of which alternatives were tested and why a conclusion was reached, rather than present one damage signature as self-interpreting.[2][1]
Abstract Reasoning¶
Use abduction constrained by tests: infer candidate processes that could have produced the observed effects, then ask what independent traces each would predict. Keep a mode (the failed behavior), a mechanism (the physical process), and an enabling condition (why that mechanism arose here) on distinct lines. A hypothesis that explains one photograph but conflicts with load records or neighboring component evidence should weaken; a coherent explanation should survive multiple evidence types.[1][4][2]
The method's scope is also a decision. An investigator may only need to confirm a component mechanism, while a public-safety board may ask for a wider causal chain and recommendations. Neither setting licenses technical evidence to settle every legal question. Test choice likewise depends on what uncertainty matters and what evidence an intervention could destroy. These decisions affect the strength of the conclusion but are not one immutable recipe for all components.[1][2][4]
Knowledge Transfer¶
The ASIC and bridge cases share a logic while not sharing a technique. In ASIC work, the visible electronic malfunction must be checked against chip, circuit and operating-condition alternatives; JPL distinguishes failure mode from mechanism. In the I-35W bridge case, NTSB linked gusset-plate load capacity to design error, long-term weight changes and day-of-collapse loads. Both move from an observed failure through contextual evidence toward a bounded explanation rather than from a list of generic future hazards.[1][4]
What does not transfer is a library of physical signatures, an exact test order or a required institutional output. Semiconductor electrical characterization cannot simply be transplanted into bridge engineering; NTSB's public-safety recommendations are not automatically the output of a laboratory component inquiry. The abstraction is a method of evidential organization across physical systems, not a claim that their physics or responsibilities are interchangeable.[1][4][2]
Examples¶
ASIC-related electronic failure¶
NASA JPL's official ASIC failure-analysis appendix describes an investigator beginning after a reported device or circuit failure. It instructs analysts to verify the failure, distinguish its mode from mechanism, and consider a cause in neighboring active/passive components, interconnect or out-of-spec operation. Its listed confirmation and laboratory examinations are an example institutional workflow, not a universal command for every failed part. The source describes a methodology rather than one named chip's final diagnosis.[1]
Mapped back: reported ASIC/circuit departure → ASIC and surrounding circuit context → failure history and relevant electrical/physical observations → chip-internal versus external mechanism hypotheses → case-selected examinations → supported cause or narrowed alternatives; corrective recommendation optional.
I-35W bridge collapse¶
The NTSB's completed investigation of the Minneapolis I-35W bridge found inadequate load capacity in U10 gusset plates due to design error, acting under increased bridge weight and traffic/construction loads. It also identified contributing design-review and inspection conditions. This public record shows why a physical failure event and an upstream causal explanation occupy different levels. It is a specific official probable-cause finding, not a universal model of bridge collapse.[4]
Mapped back: observed bridge collapse → U10 truss/gusset plates in design and load history → physical and documentary investigation → local plate failure, design and loading alternatives → evidence comparison → NTSB probable cause with contributors; later safety recommendations are a downstream institutional output.
Near miss: prospective FMEA¶
A design team lists ways an as-yet-unfailed assembly might malfunction and scores severity, occurrence and detection to prioritize design mitigations. That work may improve reliability, but it lacks an observed failed item and retrospective evidence-to-cause inference. It is FMEA, the distinct live catalog identity whose “failure analysis” alias is queued for lexical adjudication.
Structural Tensions¶
- Preserving evidence vs. testing it. A more invasive examination can expose internal conditions while altering a specimen that other tests might need. Conversely, refusing all intervention can leave decisive alternatives unresolved. Diagnostic: What evidence might this test consume, what hypothesis will it discriminate, and what documentation or separate sample preserves other lines of inquiry? There is no universal nondestructive-first rule.[1][2]
- Immediate mechanism vs. upstream conditions. The failed element is often visible, but its behavior may depend on circuit context, design, loading or quality systems; broadening the account can make it more useful yet also more speculative. Diagnostic: Which evidence supports the physical mechanism, which different evidence supports the wider condition, and where does technical explanation stop short of attribution?[1][4]
Structural–Framed Character¶
Carrier test: an engineered physical component, device or structure has actually departed from expected performance. Transformation test: physical and contextual traces are compared with alternative mechanisms and conditions. Invariant test: the finding must explain the observed departure at the warranted confidence while distinguishing mode from mechanism and wider causes. Failure test: a prospective risk worksheet or an unsupported single-signature story does not establish this retrospective identity. Transfer test: the same evidence-to-cause organization appears in electronic components and structural collapse despite different tests and scales.[1][4]
Its character: domain-framed despite a transferable diagnostic skeleton. On the five Structural–Framed criteria: vocabulary travels across physical engineering but needs function, mode, mechanism and specimen to retain precise meanings; evaluative weight enters through reliability, safety and loss but a technical finding itself is evidence-bound; institutional origin lies in engineering laboratories and accident-investigation bodies; human-practice bound is substantial because evidence preservation, test choices and reporting scope are investigator judgments; and import versus recognize favors literal recognition across physical artifacts but only analogy for software or organizational “failures.” The entry is domain-specific, not a Prime.
Structural Core vs. Domain Accent¶
The core is retrospective: an actual departure, a physical carrier in context, discriminating evidence and a bounded explanation. Metallurgical fractography, electronic fault localization and structural load modeling are accents that furnish particular kinds of traces. A universal “signature library” is not required; what matters is whether the chosen observations constrain the proposed mechanism in that case.[3][1][2]
Similarly, a laboratory report, safety recommendation and legal attribution are different possible downstream artifacts. JPL's local workflow includes an action recommendation, and NTSB issued safety recommendations after its bridge finding. Neither makes legal responsibility or a corrective-action plan part of every failure-analysis identity.[1][4]
Instantiates / Related Primes¶
This entry is a kind of Diagnostic Method.
The broader abstraction is the live Diagnostic Method: this is an evidence-acquisition and interpretation method for identifying or excluding a fault, specialized to an already failed physical engineered item and its history. FMEA is related but temporally and logically distinct—prospective enumeration rather than retrospective diagnosis. Fault Tree Analysis may organize combinations of causes for a chosen top event, but its Boolean tree is not obligatory to a component investigation. Reverse Engineering may reconstruct design intent; a failure analysis can explain why one item failed without recovering the whole design. This edge and the lexical FMEA collision both require independent review before promotion.
Relationships to Other Abstractions¶
Current abstraction Failure Analysis Domain-specific
Parents (1) — more general patterns this builds on
-
Failure Analysis is a kind of Diagnostic Method Domain-specific
Physical failure analysis specializes diagnostic evidence-to-fault inference to an already observed engineered-item failure.The method acquires physical and contextual evidence, distinguishes plausible fault mechanisms and conditions, and reports a supported technical diagnosis with uncertainty. It adds the retrospective failed-item, design/service-history and material-test constraints to the general Diagnostic Method identity; the parent can diagnose other conditions without a failed physical component.
Hierarchy path (1) — routes to 1 parentless root
- Failure Analysis → Diagnostic Method
Neighborhood in Abstraction Space¶
Failure Analysis sits in a sparse region of the domain-specific corpus (70th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Software & Systems Architecture (29 abstractions)
Nearest neighbors
- Reset (military) — 0.84
- Glitch Art — 0.84
- Engineering Critical Assessment — 0.84
- Function (engineering) — 0.84
- Aviation accident analysis — 0.83
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Metallurgical failure analysis: the narrower original candidate, focused on metal materials and their mechanisms.
- Failure Mode and Effects Analysis (FMEA): prospective failure-mode enumeration and risk prioritization; its current “failure analysis” alias is an unresolved collision.
- Fault Tree Analysis: decomposition of a defined undesirable event through Boolean causal combinations; one possible supporting tool, not this whole investigation.
- Root-cause or blame ritual: a demand for one culprit can exceed the evidence or the investigator's mandate.
- Software incident analysis: a neighboring retrospective practice outside this entry's deliberately physical-engineering scope, not a claim about the whole phrase's usage.
- A fixed nondestructive-to-destructive procedure: test choice and order depend on case, evidence and constraints.[1][2]
References¶
[1] NASA Jet Propulsion Laboratory, “Appendix Four: Failure Analysis”, original ASIC documentation, opening definition, Major Tasks and Laboratory Work Flow headings. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t ↩u ↩v ↩w ↩x ↩y
[2] Tanya Brown-Giammanco, NIST Director of Disaster and Failure Studies, “NIST’s Investigations of Structural Disasters: What We Do and Why They Can Take Years to Complete”, original NIST investigator account, evidence, alternative-hypothesis and testing paragraphs. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n
[3] NASA Kennedy Space Center, “Material Analysis Laboratory (MAL) Overview”, original facility description, Overview paragraphs on mechanical and electrical component failure analyses. registry ↩a ↩b ↩c ↩d
[4] National Transportation Safety Board, “Collapse of I-35W Highway Bridge,” investigation HWY07MH024, completed investigation page, “What We Found” and “What We Recommended.” registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q