A failure occurs when a system, process, person, component or institution does not deliver a required function or acceptable outcome.
A broken pump, missed deadline, failed examination strategy, software outage, collapsed handoff, unusable dataset and unsuccessful policy intervention can all be failures. The important question is not merely that something went wrong, but what function was lost, how much was lost, how far the failure travelled and whether the system recovered.
Quick answer: how should failures be categorised?
- Function: what required capability was lost?
- Severity: partial, major, critical or catastrophic?
- Detectability: obvious, latent, intermittent or hidden?
- Propagation: local, downstream or systemic?
- Duration: transient, persistent or permanent?
- Recovery: automatic, manual, degraded-mode, replacement or impossible?
- Recurrence: isolated, repeated or chronic?
- Cause: trigger, contributing factor, design weakness or external shock?
- Ownership: who restores service and who prevents recurrence?
- Evidence: what proves the failure and its extent?
This page complements How to Categorise Errors. An error is a deviation from a correct or expected state; a failure is the loss of required function or outcome that may result from one or many errors, conditions or shocks.
Failure is defined against function
A component can look damaged and still perform its required function. Another can appear normal while silently failing. Classification should therefore begin with the function that was expected, not appearance alone.
Partial failures preserve some capability
A system may become slower, less accurate or less available while continuing to operate. Degraded operation deserves its own category because response differs from total loss.
Total failures remove required capability
The relevant service, process or function becomes unavailable until restoration or replacement.
Intermittent failures are difficult to reproduce
They appear and disappear depending on timing, load, environment or interaction. Their evidence record should preserve conditions, not just the final symptom.
Latent failures remain hidden
A failed backup, expired credential or broken safety control may remain undiscovered until another event demands it.
Safe failures contain harm
Some systems are designed to stop, isolate or enter a safe mode when a component fails. Loss of normal service can therefore be an intentional protection mechanism rather than uncontrolled collapse.
Unsafe failures create unacceptable exposure
When failure removes protection, generates hazardous behaviour or hides its own condition, the classification should reflect the additional risk.
Local failures stay bounded
A failed component or task may have little effect beyond its immediate scope when redundancy and isolation work well.
Cascading failures travel through dependencies
One outage can propagate through shared infrastructure, resource dependencies or control relationships. Dependency mapping is therefore part of failure classification.
Systemic failures reveal structural weakness
When many independent-looking parts fail together because they share one hidden dependency, incentive or design assumption, the failure belongs at system level rather than only at component level.
Severity and duration are different
A severe one-minute outage and a mild six-month degradation require different responses. Store magnitude and duration separately.
Recoverable failures have a return path
Restart, rollback, repair, rework or replacement can restore acceptable function. Recovery time, cost and evidence of restoration should be recorded.
Irreversible failures change the future
Permanent data loss, destruction, severe harm or one-way institutional decisions may have no true restoration path. Such failures deserve stronger prevention and release controls.
Failure cause should be layered
Preserve triggering event, proximate mechanism, contributing factors, latent conditions and governance weaknesses. This avoids reducing complex failures to one convenient label.
External shocks can expose internal weakness
A storm, market shock or supplier outage may trigger failure while poor redundancy or capacity planning determines how severe the outcome becomes.
Single failures and repeated failures differ
An isolated defect may justify bounded repair. Recurrent failures suggest that the earlier correction did not remove the generating condition.
Near misses belong beside failures
A near miss reveals a pathway that could have produced failure but was contained in time. It can provide high-value learning without waiting for a damaging outcome.
Detection route matters
Self-test, monitoring, human review, customer report and downstream symptom reveal different weaknesses in the detection architecture.
Failure evidence should be independently reconstructable
Logs, measurements, screenshots, test results, witness records and version histories should show what was expected, what was observed and when the divergence occurred.
Recovery success is not proven by attempted repair
A successful restart command does not prove restored service. Verification must observe the required function after repair.
Failure ownership can be split
One role may own immediate recovery, another root-cause investigation, another prevention, and another communication to affected users.
High-severity failures require stronger review independence
Where safety, rights, money or critical service are affected, the team that performed the repair should not be the only source of verification.
Failure classes should improve prevention
If the category does not change monitoring, redundancy, testing, training, escalation or design, it may not be operationally useful.
AI can help cluster failure patterns
Models can group similar incidents and surface hidden correlations, but machine-generated root-cause claims should remain hypotheses until evidence supports them.
A practical failure record
- failure ID;
- affected function;
- expected state;
- observed state;
- severity;
- detectability;
- scope and propagation;
- start and recovery time;
- trigger and contributing causes;
- dependencies;
- recovery method;
- verification evidence;
- owners;
- recurrence class;
- prevention action and version.
The deeper idea
Failure is not simply the moment something breaks. It is the relationship between required function, lost capability, propagation and recovery.
To categorise a failure well is to know what function was lost, how severely, how far the loss travelled, how it was detected, how it recovered and what changed so the same pathway is less likely to return.
Final answer
Categorise failures by lost function, severity, detectability, propagation, duration, recoverability, recurrence, cause, ownership and evidence. Keep failure separate from error, distinguish degraded from total loss, and verify restored function rather than assuming that a repair attempt succeeded.
