How to Categorise Failures | Function, Severity, Detectability, Propagation, Recovery and Recurrence

A failure occurs when a system, process, person, component or institution does not deliver a required function or acceptable outcome.

A broken pump, missed deadline, failed examination strategy, software outage, collapsed handoff, unusable dataset and unsuccessful policy intervention can all be failures. The important question is not merely that something went wrong, but what function was lost, how much was lost, how far the failure travelled and whether the system recovered.

Quick answer: how should failures be categorised?

  • Function: what required capability was lost?
  • Severity: partial, major, critical or catastrophic?
  • Detectability: obvious, latent, intermittent or hidden?
  • Propagation: local, downstream or systemic?
  • Duration: transient, persistent or permanent?
  • Recovery: automatic, manual, degraded-mode, replacement or impossible?
  • Recurrence: isolated, repeated or chronic?
  • Cause: trigger, contributing factor, design weakness or external shock?
  • Ownership: who restores service and who prevents recurrence?
  • Evidence: what proves the failure and its extent?

This page complements How to Categorise Errors. An error is a deviation from a correct or expected state; a failure is the loss of required function or outcome that may result from one or many errors, conditions or shocks.

Failure is defined against function

A component can look damaged and still perform its required function. Another can appear normal while silently failing. Classification should therefore begin with the function that was expected, not appearance alone.

Partial failures preserve some capability

A system may become slower, less accurate or less available while continuing to operate. Degraded operation deserves its own category because response differs from total loss.

Total failures remove required capability

The relevant service, process or function becomes unavailable until restoration or replacement.

Intermittent failures are difficult to reproduce

They appear and disappear depending on timing, load, environment or interaction. Their evidence record should preserve conditions, not just the final symptom.

Latent failures remain hidden

A failed backup, expired credential or broken safety control may remain undiscovered until another event demands it.

Safe failures contain harm

Some systems are designed to stop, isolate or enter a safe mode when a component fails. Loss of normal service can therefore be an intentional protection mechanism rather than uncontrolled collapse.

Unsafe failures create unacceptable exposure

When failure removes protection, generates hazardous behaviour or hides its own condition, the classification should reflect the additional risk.

Local failures stay bounded

A failed component or task may have little effect beyond its immediate scope when redundancy and isolation work well.

Cascading failures travel through dependencies

One outage can propagate through shared infrastructure, resource dependencies or control relationships. Dependency mapping is therefore part of failure classification.

Systemic failures reveal structural weakness

When many independent-looking parts fail together because they share one hidden dependency, incentive or design assumption, the failure belongs at system level rather than only at component level.

Severity and duration are different

A severe one-minute outage and a mild six-month degradation require different responses. Store magnitude and duration separately.

Recoverable failures have a return path

Restart, rollback, repair, rework or replacement can restore acceptable function. Recovery time, cost and evidence of restoration should be recorded.

Irreversible failures change the future

Permanent data loss, destruction, severe harm or one-way institutional decisions may have no true restoration path. Such failures deserve stronger prevention and release controls.

Failure cause should be layered

Preserve triggering event, proximate mechanism, contributing factors, latent conditions and governance weaknesses. This avoids reducing complex failures to one convenient label.

External shocks can expose internal weakness

A storm, market shock or supplier outage may trigger failure while poor redundancy or capacity planning determines how severe the outcome becomes.

Single failures and repeated failures differ

An isolated defect may justify bounded repair. Recurrent failures suggest that the earlier correction did not remove the generating condition.

Near misses belong beside failures

A near miss reveals a pathway that could have produced failure but was contained in time. It can provide high-value learning without waiting for a damaging outcome.

Detection route matters

Self-test, monitoring, human review, customer report and downstream symptom reveal different weaknesses in the detection architecture.

Failure evidence should be independently reconstructable

Logs, measurements, screenshots, test results, witness records and version histories should show what was expected, what was observed and when the divergence occurred.

Recovery success is not proven by attempted repair

A successful restart command does not prove restored service. Verification must observe the required function after repair.

Failure ownership can be split

One role may own immediate recovery, another root-cause investigation, another prevention, and another communication to affected users.

High-severity failures require stronger review independence

Where safety, rights, money or critical service are affected, the team that performed the repair should not be the only source of verification.

Failure classes should improve prevention

If the category does not change monitoring, redundancy, testing, training, escalation or design, it may not be operationally useful.

AI can help cluster failure patterns

Models can group similar incidents and surface hidden correlations, but machine-generated root-cause claims should remain hypotheses until evidence supports them.

A practical failure record

  • failure ID;
  • affected function;
  • expected state;
  • observed state;
  • severity;
  • detectability;
  • scope and propagation;
  • start and recovery time;
  • trigger and contributing causes;
  • dependencies;
  • recovery method;
  • verification evidence;
  • owners;
  • recurrence class;
  • prevention action and version.

The deeper idea

Failure is not simply the moment something breaks. It is the relationship between required function, lost capability, propagation and recovery.

To categorise a failure well is to know what function was lost, how severely, how far the loss travelled, how it was detected, how it recovered and what changed so the same pathway is less likely to return.

Final answer

Categorise failures by lost function, severity, detectability, propagation, duration, recoverability, recurrence, cause, ownership and evidence. Keep failure separate from error, distinguish degraded from total loss, and verify restored function rather than assuming that a repair attempt succeeded.


Continue through the series

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.