A process can produce acceptable work today and still be drifting toward failure. It can also produce an occasional bad result even though its underlying system has not changed. The difficult question is not simply whether the latest measurement is high or low. It is whether the pattern of variation tells us that the process itself has changed.
Statistical process control (SPC) is a disciplined way to answer that question. It uses time-ordered data, control charts, knowledge of the process and explicit response rules to distinguish routine variation from evidence that something different may be happening. The central ideas include process stability, common-cause and special-cause variation, Shewhart control charts, process capability, Cp and Cpk, and more sensitive methods such as CUSUM and EWMA.
SPC matters because improvement becomes dangerous when every fluctuation is treated as a crisis, and complacency becomes dangerous when a real process change is dismissed as random noise. A control chart does not automate judgement, certify quality, or prove causation. It creates a structured conversation between data and process knowledge so that intervention happens for a reason.
This article is a practical owner-level guide to that conversation. It separates control limits from specification limits, stability from capability, monitoring from acceptance sampling, and statistical signals from root-cause conclusions. It also shows where measurement quality, subgroup design, autocorrelation, non-normal data and repeated signalling rules can break an apparently sophisticated chart.
For the wider foundations, use How Scientific Measurement Works, How Statistical Inference and Uncertainty Work, Data Quality and Project Quality Management. This article owns a different job: the time-ordered statistical monitoring and improvement of a process.
1. Why statistical process control exists
Imagine a process that fills containers to a nominal value of 50 units. The latest measurement is 50.4. Is that bad? Without context, the number is almost meaningless. If the process normally varies between 49.7 and 50.4, the result may be routine. If it normally varies between 49.98 and 50.02, the same 50.4 may indicate a large shift. If the measuring instrument itself has changed, the apparent process shift may not belong to the process at all.
SPC begins from a deceptively simple observation: variation is always present, but variation does not always have the same meaning. Some variation arises from the ordinary combined behaviour of many small influences in the present system. Other variation appears when something specific changes: a setting, batch, tool, environment, method, supplier, operator condition, data pipeline or measurement device.
The practical risk is symmetrical. If managers react to every routine fluctuation, they can make a stable process less stable through unnecessary adjustment. If they ignore genuine signals, they can allow a changed process to continue producing avoidable failure. SPC is therefore not merely a graphing technique. It is a discipline for deciding when not to tamper and when not to look away.
NIST’s Process or Product Monitoring and Control handbook presents SPC as part of a wider set of monitoring and control techniques. Its control-chart sections describe charts as tools for routinely monitoring process behaviour and signalling when investigation may be needed. ISO 7870-1:2019 likewise presents the general philosophy and families of control charts; ISO reports that this edition was reviewed and confirmed in 2025, while a new edition is under development as of 2026.
2. Variation is the object of study
Suppose a service team completes the same type of task many times. Completion time differs from case to case because people, inputs, queues, systems and circumstances differ. A production line behaves similarly: material properties, temperature, machine condition, calibration and timing vary. Even a highly automated process is not identical from one cycle to the next.
SPC does not ask whether variation exists. It asks whether the variation is consistent with the process that has been operating, or whether the evidence suggests that the process has entered a different state.
Common-cause variation
Common-cause variation is the variation produced by the ordinary system as it currently operates. It is not “good variation”. A stable process can be stably bad. It can consistently produce excessive delay, waste or defect. Stability means that the pattern is sufficiently consistent for the chosen model and monitoring method, not that customers should be satisfied.
If variation is common-cause, improvement usually requires changing the system: the method, design, equipment, training architecture, material flow, standard work, workload, environment or another persistent feature. Blaming the most recent individual operator may do nothing to the mechanism producing the variation.
Special-cause variation
Special-cause variation refers to evidence that the process may have been affected by a condition not represented by its established pattern. A sudden tool failure, wrong setup, new supplier batch, software release, sensor fault, unusual temperature event or transcription error can create such a signal.
A signal does not identify the cause. It says that the observed pattern is sufficiently unusual under the chart’s working model to deserve investigation. The cause might ultimately be found outside the physical process—for example, a changed measurement method or data extraction rule. The chart is a detector, not a detective.
The two wrong reactions
Overreaction: change the process after every high or low point. This adds intervention noise to ordinary variation. A process that would have returned naturally toward its usual level can be pushed away by repeated adjustment.
Underreaction: explain every unusual point as “just variation”. This protects a changed process from investigation and can normalize a genuine failure.
SPC exists between these errors. Its value comes from making the distinction explicit before pressure, hindsight and narrative bias enter the room.
3. What a control chart actually contains
A basic control chart places a process statistic in time order. The statistic might be one observation, an average, a range, a proportion, a count or another quantity chosen for the process. The chart then adds a centre line and statistically derived control limits.
NIST describes the typical chart as having a centre line representing the in-control process and upper and lower control limits chosen so that almost all points are expected to fall within them while the process remains in control. In the traditional Shewhart form, three standard-error units are commonly used for the control-limit distance.
- Centre line: an estimate of the process location or the expected value of the plotted statistic under the current baseline.
- Upper control limit: a boundary for unusually high values under the working model.
- Lower control limit: a boundary for unusually low values under the working model.
- Time order: the sequence that allows shifts, trends, cycles and other temporal structure to be visible.
- Response rule: the agreed action when a signal appears.
The last item is often neglected. A control chart without a response plan can become wall decoration. Before a signal occurs, decide who reviews it, what data are preserved, how the process is checked and who has authority to intervene. Otherwise the organisation may improvise a different meaning for the chart every time the result is inconvenient.
4. Control limits are not specification limits
This is one of the most important distinctions in quality engineering.
Control limits describe the process. They are estimated from process behaviour under a stated baseline and statistical method.
Specification limits describe a requirement or acceptable range. They come from design, customer, regulatory, engineering or other legitimate requirements. They should not be calculated from whatever the process happens to do.
A process can be statistically stable yet incapable of meeting specification. Imagine a stable distribution centred exactly on target but much wider than the permitted range. The control chart can show no special-cause signal while customers still receive unacceptable output. Conversely, a process can sit comfortably inside specifications while showing a statistically significant shift that deserves attention before it eventually reaches the specification boundary.
Control limits ask: “Has the process changed?” Specification limits ask: “Is the output acceptable for its intended requirement?”
Drawing specification lines and calling them control limits destroys this distinction. It converts a monitoring method into a pass/fail display and hides early evidence of change. Drawing control limits and calling everything inside them “good” creates the opposite error.
5. Stability comes before capability
NIST’s process-capability guidance explicitly compares the output of an in-control process with its specification limits. That ordering matters. A capability index attempts to summarise the relation between the stable process distribution and the requirements. If the process is moving between different states, one mean and one standard deviation may describe no real operating condition.
Consider a machine that runs near 10.00 in the morning, shifts to 10.12 after a setup change and returns to 9.96 after maintenance. Pooling the data could produce a standard deviation and a Cpk value. The arithmetic would be valid for the pooled numbers, but the interpretation “this is the capability of one stable process” would be doubtful. The first question should be why the system occupies different states.
Capability analysis is therefore not a substitute for time-order analysis. A histogram can hide a temporal shift. Two narrow states can combine into one wide distribution. A control chart can reveal the sequence that the histogram removes.
6. Rational subgrouping decides what the chart is sensitive to
When multiple observations are grouped into a subgroup, the grouping should have a process reason. Observations inside a subgroup are usually chosen so that they represent short-term variation under similar conditions, while differences between subgroup averages can reveal changes over time.
If a process produces five consecutive pieces every hour, those five may form a rational subgroup when the aim is to detect hour-to-hour changes while estimating short-term within-hour variation. If instead the five observations are deliberately drawn from five different machines, days and suppliers, the within-subgroup spread can absorb the very differences the chart was meant to detect.
There is no universal subgroup size detached from purpose. The right design depends on the sampling opportunity, speed of change, cost of measurement and failure mechanism. The chart is not chosen first and the data forced into it afterwards. The monitoring question comes first.
7. Choose the chart for the data-generating process
The familiar control-chart families are not interchangeable decorations. They make different assumptions about the statistic, subgrouping and data type.
| Situation | Common chart family | What is monitored |
|---|---|---|
| One continuous observation at a time | Individuals and Moving Range (I-MR) | Individual values and short-term successive variation |
| Small rational subgroups of continuous measurements | X-bar and R | Subgroup mean and subgroup range |
| Larger rational subgroups of continuous measurements | X-bar and S | Subgroup mean and subgroup standard deviation |
| Fraction nonconforming, varying sample size | p chart | Proportion nonconforming |
| Number nonconforming, constant sample size | np chart | Count of nonconforming units |
| Number of nonconformities per constant opportunity unit | c chart | Count of nonconformities |
| Nonconformities with changing opportunity/exposure | u chart | Rate of nonconformities per unit |
| Small sustained shift is important | CUSUM or EWMA | Accumulated or weighted departure from baseline |
| Several correlated characteristics jointly matter | Multivariate methods | A joint statistic rather than isolated univariate charts |
This table is a navigation aid, not a complete design standard. NIST’s control-chart material and the ISO 7870 series provide deeper method-specific treatment. The relevant ISO standard for Shewhart charts is ISO 7870-2:2023.
8. Worked example: an Individuals and Moving Range chart
Consider a fictional calibration-check process that produces one reference reading per run. The first nineteen readings are:
50.1, 49.8, 50.2, 50.0, 49.9, 50.3, 50.1, 49.7, 50.0, 50.2, 49.9, 50.1, 50.4, 50.0, 49.8, 50.2, 50.1, 49.9, 50.0.
A twentieth reading is 54.8. All values are invented for teaching. They do not describe a real instrument or acceptable calibration criterion.
For an Individuals chart, one common estimator of short-term variation uses the moving range between successive values. For the full twenty-point sequence, the average moving range is 0.500. Using the familiar Individuals-chart multiplier of approximately 2.66 gives a centre line of 50.275, a lower control limit of approximately 48.945 and an upper control limit of approximately 51.605.
The twentieth value, 54.8, lies far beyond the upper control limit. The chart therefore supplies a strong signal that the last observation is inconsistent with the preceding pattern under this working model. Notice what it does not tell us: it does not say whether the process itself changed, whether the instrument was misread, whether a transcription error occurred, or whether the reference condition was wrong.
The moving-range chart also matters. The final jump from 50.0 to 54.8 increases the moving range sharply. The average moving range for the full sequence is 0.500, giving an upper moving-range limit of approximately 1.634 under the standard two-point moving-range constant. The last moving range is 4.8, again an obvious signal.
Do not delete the point merely because it is inconvenient
The investigation might discover a documented transcription error: the original record says 50.8 and the database says 54.8. If correction is legitimate, preserve both the erroneous entry and its correction history as appropriate, then recalculate the baseline from the corrected data. If the reading was real and caused by a one-off identifiable setup error, the team may separately model the stable state after documenting why that episode should not define the ongoing baseline.
What the team should not do is remove any point outside the limits simply to make the chart look controlled. A signal is the reason to investigate. Automatic deletion turns the detector into a cosmetic filter.
What the earlier nineteen points imply
If the twentieth point is shown by evidence to belong to a distinct special cause and the earlier nineteen points form the appropriate baseline, their mean is approximately 50.037 and their average moving range approximately 0.261. The corresponding illustrative Individuals limits are approximately 49.342 and 50.731. These calculations show why a large special-cause point can widen limits when it is used to estimate the process that it may not belong to.
The purpose of a Phase I analysis is precisely to establish a credible baseline through process knowledge and data. It is not to assume that the first arbitrary batch of observations already represents a stable system.
9. Phase I and Phase II answer different questions
Control-chart work is often divided into two broad modes.
Phase I: use historical or initial process data to understand behaviour, identify obvious special causes, assess assumptions and establish a baseline suitable for monitoring. Limits may be revised as the team learns which data represent the intended process state.
Phase II: use the established baseline to monitor future behaviour. The limits now have an operational purpose: detect evidence that the process has departed from the state used to create them.
Confusing the phases causes trouble. If limits are continuously recalculated after every new point, a gradual deterioration can drag the limits with it and become normalised. If initial limits are frozen forever despite a verified system redesign, the chart can keep signalling an obsolete baseline. Governance of the baseline is part of SPC.
10. X-bar and R charts separate location from within-subgroup variation
Suppose five consecutive units are measured every hour. The team calculates the average of each group and its range. The X-bar chart monitors the subgroup averages; the R chart monitors within-subgroup spread.
This separation is useful because a process can change in different ways. The average may shift while within-subgroup variation remains similar. Variation may increase without a large mean shift. If only the mean chart is watched, an unstable dispersion can be missed. If only the range chart is watched, a sustained location shift can be missed.
Traditional formulas use constants based on subgroup size to convert the average range into control limits for both charts. Those constants are not magical properties of a machine; they follow from the statistical design. If subgroup size changes materially, the appropriate calculations and interpretation can also change.
Subgrouping can hide or reveal a shift
Imagine a process that alternates between two machines, one running high and one low. If each subgroup contains a deliberate mixture from both machines, the subgroup average can look stable while the range remains inflated. If each machine is monitored separately, the difference becomes obvious. Neither design is universally correct; the right design depends on whether the improvement question concerns the combined output or the individual machine mechanisms.
This is why the person building the chart needs to understand the production or service flow. Statistical software can compute a limit from almost any rectangular dataset. It cannot decide whether the rows represent a meaningful process state.
11. Attribute charts require a clear unit, opportunity and definition
Not every process measurement is continuous. A unit may be conforming or nonconforming. A document may contain zero, one or several defined errors. A transaction batch may have a varying number of opportunities for a particular failure.
The statistical chart should match that structure. A p chart monitors a proportion when sample size can vary. An np chart monitors a number of nonconforming units when the sample size is effectively constant. A c chart can monitor counts of nonconformities under a constant opportunity base. A u chart adjusts a count for changing opportunity or exposure.
The definitions are more important than the chart label. If one reviewer records a document with three mistakes as “one bad document” and another records “three defects”, the metric has changed. If transaction volume doubles but the organisation plots raw complaint counts as if exposure were constant, an apparent deterioration may partly reflect more opportunities for the event.
Zero is data, not proof of perfection
A run of zero recorded defects can mean the process improved. It can also mean the opportunity for detecting defects changed, reporting failed, or the definition became narrower. Attribute data often depend strongly on the detection system. The chart monitors the recorded process, including its measurement and classification apparatus.
This is the same principle developed in Data Quality: a clean number is only as trustworthy as the process that created and maintained it.
12. A point beyond the limits is not the only possible signal
Shewhart charts are often supplemented by pattern rules: long runs on one side of the centre line, trends, or unusual concentrations near or beyond warning zones. Such rules can increase sensitivity to smaller or sustained departures that may not push one point beyond a three-sigma limit.
But every additional rule creates additional opportunities to signal. If many rules are applied to every chart and every point, the overall false-alarm behaviour changes. A team should not treat a collection of pattern rules as if each retained the same standalone false-alarm interpretation.
This connects directly to How Multiple Testing and Sequential Analysis Work. Control-chart monitoring is inherently sequential: observations arrive repeatedly, and the decision system keeps looking. The operating characteristics should be understood at the level of the monitoring scheme, not one isolated test.
A signal rule needs an action rule
If the team uses “eight consecutive points above the centre line” as a shift rule, what happens on the eighth point? Who checks whether the data are correct? Which recent process changes are reviewed? Is output contained pending investigation? Does the decision depend on severity? Without an agreed response, adding signal rules can simply create more alarms without more learning.
13. Process capability asks a different question
Once a process is reasonably stable for the intended analysis, capability compares its natural variation with requirements. NIST describes capability as comparing the output of an in-control process with specification limits.
The familiar two-sided index is:
Cp = (USL − LSL) / (6σ)
where USL and LSL are the upper and lower specification limits and σ represents the relevant process standard deviation under the capability model.
Cp compares specification width with process spread. It does not penalise an off-centre process. Cpk adds the process mean and therefore responds to centring:
Cpk = min[(USL − μ)/(3σ), (μ − LSL)/(3σ)]
These expressions are useful only when their assumptions and estimators fit the process. They should not be interpreted as universal risk guarantees.
14. Worked capability example: the process can be wide enough and still badly centred
Consider a separate fictional stable process with a lower specification limit of 9.80, an upper specification limit of 10.20, a process mean of 10.05 and an assumed within-process standard deviation of 0.04. These values are constructed only to demonstrate the indices.
The specification width is 0.40. Six standard deviations equal 0.24. Therefore:
Cp = 0.40 / 0.24 ≈ 1.67.
The distance from the mean to the upper specification is 0.15, giving an upper-side capability of 0.15 / 0.12 = 1.25. The distance to the lower specification is 0.25, giving 0.25 / 0.12 ≈ 2.08. Therefore:
Cpk = min(1.25, 2.08) = 1.25.
The gap between Cp and Cpk tells us something useful: the potential spread relative to the specification width is better than the achieved capability on the limiting side because the process is shifted upward from the midpoint of 10.00. Recentering could improve Cpk without reducing variation, if the process mechanism permits it.
Do not turn 1.33, 1.67 or another capability threshold into a universal law. Organisations and domains use different requirements depending on consequence, measurement, customer expectations and contractual or regulatory context. The index is evidence for a decision, not the authority for the decision.
15. Five ways capability analysis becomes misleading
Trap 1: the process is not stable
A single standard deviation summarises a mixture of states. The index then describes the pooled dataset more than a reproducible operating condition.
Trap 2: specifications are mistaken for statistical limits
The process is declared “in control” because most points lie inside customer specifications. That statement does not follow. A process can drift materially while remaining within a wide tolerance.
Trap 3: non-normal behaviour is ignored
The familiar normal-distribution interpretation of six sigma around a mean may not represent a skewed, bounded, multimodal or heavy-tailed distribution. Depending on the process, transformation, distribution-specific modelling or a nonparametric approach may be more appropriate.
Trap 4: measurement error is treated as process variation
If the measurement system contributes a large fraction of the observed variation, capability can look worse than the underlying process—or changes in measurement can look like process shifts. Measurement adequacy should be investigated as part of process understanding.
Trap 5: a favourable index is treated as permission to stop monitoring
Capability is conditional on the process state. A previously capable process can shift. A high historical Cpk is not an immunity certificate for future production.
16. The measurement system is inside the control system
SPC often appears to begin after the measurements have been collected. In reality, the measurement process is one of the processes being relied upon. Resolution, repeatability, reproducibility, bias, calibration status, environmental sensitivity, rounding and data transfer can all affect the chart.
If a device reports only whole units while the process changes by tenths, the chart may show long flat runs and sudden jumps created partly by resolution. If a new instrument has a different bias, a step change may appear on the process chart even when the underlying product has not moved. If two inspectors classify the same defect differently, an attribute chart can signal a change in interpretation rather than production.
The response is not to distrust all data. It is to treat measurement as a causal component. For the broader conceptual foundation, see How Scientific Measurement Works. SPC becomes stronger when the measurement chain and the process chain are investigated together.
17. Autocorrelation can make ordinary limits behave unexpectedly
Many classical control-chart calculations work most simply when successive observations or subgroup statistics behave approximately independently under the baseline. Real processes can contain serial dependence: temperature carries over, queues accumulate, chemical concentrations evolve gradually, software load persists, or maintenance affects several subsequent cycles.
When observations are strongly autocorrelated, a standard chart can signal too often or too slowly depending on the structure. The apparent run pattern may reflect predictable time-series behaviour rather than a newly introduced special cause.
The repair can include changing the sampling interval, modelling the time-series structure and monitoring residuals, or using a method designed for the process dynamics. The important principle is not “always remove autocorrelation”. It is “model the mechanism that makes observations dependent instead of pretending the sequence is exchangeable when it is not.”
18. Non-normal data require thought, not ritual transformation
Cycle times, concentrations, waiting times and defect counts often have distributions that are skewed or bounded. Some charting methods are reasonably robust to moderate departures under suitable conditions; others rely more strongly on specific distributional assumptions. Capability indices can be especially sensitive when tail probabilities are interpreted through a normal model that does not fit.
A transformation can sometimes create a more useful model, but it changes the scale and can complicate interpretation. A distribution-specific model may be preferable when the mechanism suggests one. Nonparametric or empirical approaches may be appropriate in other settings. The decision should serve the monitoring job, not the desire to make a diagnostic plot look conventional.
The analyst should also distinguish a non-normal stable process from a mixture created by hidden states. A bimodal histogram may not call for a clever distribution fit at all; it may reveal two machines, two recipes, two shifts or two measurement systems that should be analysed separately.
19. CUSUM and EWMA trade simple visibility for sensitivity to smaller shifts
A Shewhart chart reacts strongly to a large isolated departure because each point receives most of the attention individually. Smaller persistent shifts can take longer to cross a conventional Shewhart limit. Two important alternatives accumulate evidence over time.
CUSUM
A cumulative-sum chart accumulates departures from a target or reference value. Repeated small deviations in the same direction build rather than repeatedly resetting to zero. This makes CUSUM useful when a sustained small shift matters operationally.
EWMA
An exponentially weighted moving average gives the current observation the greatest weight while retaining decreasing influence from older observations. The smoothing parameter controls the memory of the statistic. ISO 7870-6:2024 specifically addresses EWMA control charts for the process mean and notes their value for detecting small to moderate shifts in settings where sampling can be slow, expensive or consequential.
These methods are not automatically “better” than Shewhart charts. They optimise different detection tasks. A monitoring design should ask what size and duration of change matters, what false-alarm rate is tolerable and how quickly the process can be investigated. A detector that signals sooner but creates operationally unmanageable false alarms may not improve the whole system.
20. Multivariate monitoring is needed when variables move together
Many processes have several correlated characteristics. Temperature and pressure may move together. Two dimensions may be linked by the same setup. A service process may trade waiting time against throughput. Monitoring each variable independently can miss a change in their joint relationship or create a flood of separate alarms.
Multivariate control methods combine information across variables into a joint statistic. Their usefulness depends on a credible covariance structure, adequate data and interpretable follow-up. A single multivariate signal can tell the team that the joint pattern is unusual without immediately revealing which variable or combination caused it.
The investigation plan therefore matters even more. A mathematically compact detector can create an operationally opaque alarm if nobody knows how to decompose the signal. Statistical compression should not erase the process knowledge needed for repair.
21. SPC is not acceptance sampling
Acceptance sampling asks whether a lot or batch should be accepted under a sampling plan. SPC asks how the process is behaving through time. The two can coexist, but they answer different questions.
A lot can pass acceptance criteria while the process that produced it is drifting. A stable process can also produce a lot that fails an acceptance rule due to ordinary random variation. Converting an acceptance sample into a control chart after the fact does not automatically create a valid monitoring design because the sampling purpose and structure may differ.
NIST separates process monitoring/control from lot-acceptance sampling within the same handbook chapter for exactly this reason. The shared use of sampling should not blur the different decisions.
22. A control-chart signal starts root-cause work; it does not finish it
Suppose a chart signals at 14:20. The team should preserve the surrounding process state: material lot, machine settings, environmental conditions, maintenance events, software version, operator handoffs, upstream changes and the measurement system. The exact list depends on the domain.
The investigation should compare candidate explanations against evidence. “Operator error” is not a root cause merely because a person touched the process. “Machine problem” is not a root cause merely because the signal occurred on a machine. The team needs a mechanism capable of producing the observed pattern and evidence that the mechanism was actually present.
When a cause is found, distinguish containment from correction. Containment protects current output or users. Correction removes the immediate issue. Corrective action changes the system to reduce recurrence. These can occur on different timelines.
23. Tampering: how management can manufacture variation
Imagine a process naturally fluctuating around target. After a slightly low result, an operator increases the setting. The next result is slightly high, so the setting is reduced. The process is now being adjusted partly in response to random noise. Each adjustment can add variation instead of reducing it.
This is one of SPC’s deepest management lessons. A dashboard that shows every fluctuation in red or green can pressure teams to explain ordinary variation as if every movement had a special story. The result is an organisation that continuously acts without necessarily learning.
Not all feedback adjustment is tampering. Automated control systems legitimately respond to measured deviations when the process dynamics and control law justify the action. The distinction is whether the adjustment is part of a designed control mechanism or an improvised reaction to noise.
24. Stable does not mean finished: how SPC supports improvement
A stable process gives the team something valuable: a reproducible baseline. If the baseline is unacceptable, improvement can be designed and evaluated against it. Instead of chasing individual bad days, the team can change the system and ask whether the new system has a different centre, spread or failure pattern.
The improvement cycle might look like this:
- Define the process and the outcome that matters.
- Verify the measurement system and operational definitions.
- Collect time-ordered baseline data.
- Assess stability and investigate special causes.
- Decide whether the stable process meets the requirement.
- If not, identify a plausible mechanism for improvement.
- Change the process deliberately.
- Monitor the transition without blending old and new baselines.
- Establish the new stable state if the evidence supports it.
- Reassess capability and sustain monitoring.
The cycle deliberately separates stabilising from improving. Removing a special cause can restore the previous process. Redesigning the common-cause system can create a better process. Both are valuable, but they are not the same intervention.
25. SPC can work outside manufacturing—but only if the process is defined honestly
Statistical process control is widely associated with manufacturing because repeated physical processes make the logic visible. The same ideas can sometimes be useful in services, laboratories, logistics, software operations, healthcare operations and administrative work.
The challenge is that service cases may differ substantially in complexity. A five-minute password reset and a three-week investigation are both “tickets” but may not belong to one homogeneous process. A school assignment, hospital episode or legal case can differ in ways that matter to the outcome. A chart does not make unlike cases comparable by placing them in one column.
Before charting a service metric, define the unit, entry and exit points, relevant stratification, exposure and operational meaning. If case mix changes, the chart may reflect a changing population rather than a changed process. Risk adjustment or stratified monitoring may be needed in high-variation domains.
The transferable principle is not “put everything on a control chart”. It is “respect time order, variation and the mechanism that creates the data”.
26. Seven dashboard habits that break SPC
1. Sorting the data instead of preserving time order
A ranked table can show who is highest or lowest but removes temporal structure. The chart can no longer reveal whether the process shifted after a particular event.
2. Colouring every point against target
Red/green target colouring answers an acceptance question. It does not replace statistical limits that estimate process behaviour.
3. Recalculating limits every reporting period
If a deteriorating process repeatedly updates its own baseline, the system can redefine failure as normal.
4. Comparing unlike units without stratification
Different products, machines, case types or customer populations can create mixtures that one chart cannot interpret cleanly.
5. Treating every signal as proof of a cause
The signal says investigate. The investigation supplies the causal evidence.
6. Celebrating a stable but incapable process
Statistical control can coexist with poor customer outcomes. Stability is a prerequisite for some kinds of prediction, not the final quality objective.
7. Using a control chart when there is no stable repeatable process to monitor
One-off strategic decisions, rapidly changing experimental prototypes or incomparable cases may need other analytical tools. SPC is powerful when the concept of a continuing process is real.
27. A control chart needs governance
A mature SPC system should be able to answer operational questions that software alone cannot answer:
- Who owns the process?
- Who owns the measurement definition?
- Which data create the baseline?
- Who can approve a new baseline?
- Which chart and signal rules are authorised?
- What happens when a signal appears?
- How are suspected data errors handled?
- How are process changes recorded?
- When is capability recalculated?
- Which specification version applies?
- How are discontinued charts archived?
Without this governance, two teams can produce different “control charts” for the same process by choosing different date windows, exclusions and subgroup rules. Both graphs may be mathematically correct and operationally incompatible.
28. A complete worked reasoning chain
Return to the fictional 50-unit process. A strong investigation can be written as a chain rather than a collection of screenshots.
Step 1: define the quantity
The measurement is one reference reading per completed run under a specified check condition. The unit and measurement procedure are fixed for the baseline period.
Step 2: preserve sequence
The readings are ordered by actual run time, not by value or operator. The twentieth reading is 54.8 after nineteen values close to 50.
Step 3: detect the signal
The Individuals and Moving Range calculations show that the twentieth observation and its moving range are inconsistent with the working baseline.
Step 4: investigate the record before the machine
The team checks the source record, measurement device, unit conversion, database transfer and time stamp. This prevents a data-processing error from becoming an unnecessary mechanical intervention.
Step 5: investigate process changes
If the value is genuine, review relevant process events near the signal: setup, material, environment, maintenance, method and other plausible mechanisms. The list should be domain-specific rather than generic blame hunting.
Step 6: contain proportionately
If the measurement concerns a characteristic whose deviation creates meaningful risk, affected output may need containment while the investigation proceeds. The exact action depends on the real specification, consequence and authority; this fictional example does not prescribe a real disposition rule.
Step 7: repair the mechanism
If a cause is found, correct it and decide whether the monitoring plan itself should change. A recurring setup problem may require mistake-proofing or revised standard work, not just another reminder.
Step 8: govern the baseline
If the old baseline remains the intended process, keep it for future monitoring. If the process has been deliberately redesigned and stabilised at a new state, establish a new baseline with a documented effective point. Do not blur pre-change and post-change data into one average merely for convenience.
Step 9: assess capability only after stability
Once the process state is credible, compare its distribution with the legitimate requirements. Control and capability now answer complementary questions: is the process behaving consistently, and is that consistent behaviour good enough?
29. Practice problems
Exercise A: control limit or specification?
A customer requires a dimension between 19.90 and 20.10. A stable process has a control-chart centre line of 20.00 and control limits of 19.85 and 20.15. Which pair is calculated from process behaviour and which comes from the requirement?
Answer: 19.85 and 20.15 are the statistical control limits in the example. 19.90 and 20.10 are specification limits. The process may be stable while still producing output outside specification because its natural variation is wider than the allowed range.
Exercise B: stable but incapable
A process shows no special-cause signals for three months, but ten per cent of output falls outside a legitimate customer requirement. Is the process “good” because the chart is stable?
Answer: no. Stability says the current behaviour is sufficiently consistent under the monitoring model. It does not say that the behaviour meets the requirement. The common-cause system needs improvement.
Exercise C: changing sample size
A team records the number of nonconforming cases each day, but daily case volume ranges from 20 to 500. Why can a raw count chart be misleading?
Answer: the opportunity for a nonconforming case changes dramatically with daily volume. Depending on the data structure and assumptions, a proportion or rate-based chart may be more appropriate. The denominator needs to be part of the monitoring definition.
Exercise D: the favourable Cpk
A report shows Cpk = 1.8, but the time series contains a large shift halfway through the study period. What should you ask first?
Answer: whether the pooled dataset represents one stable process and whether the capability model is valid for it. A favourable index cannot repair a broken process-state assumption.
Exercise E: a point inside specification but beyond the control limit
A new measurement remains safely inside the engineering tolerance but falls beyond the established upper control limit. Should the signal be ignored?
Answer: not automatically. The output may still be acceptable, but the process may have changed. Investigating the signal can identify drift before the specification limit is reached.
Exercise F: many pattern rules
A team adds fifteen different signal rules because it wants to detect every possible change. What statistical and operational question should be asked?
Answer: what is the combined false-alarm behaviour and what response capacity exists for the resulting signals? Greater sensitivity is not free. The monitoring system needs operating characteristics and an action plan, not the maximum number of alarms.
30. An implementation checklist for a real SPC programme
- Define the decision: what change must the chart detect, and why does it matter?
- Define the process: where does it start and stop, and what conditions make observations comparable?
- Define the measure: unit, method, instrument, classification rule and data path.
- Check measurement adequacy: can the measurement system resolve the changes that matter?
- Preserve time order: do not begin by aggregating away the sequence.
- Choose subgrouping deliberately: group observations according to the process mechanism.
- Choose the chart family: match continuous, attribute, subgroup, rate or multivariate structure.
- Establish Phase I: understand historical behaviour and candidate special causes.
- Document exclusions and corrections: never erase inconvenient observations silently.
- Set signal rules: know the sensitivity and false-alarm implications of the monitoring scheme.
- Write the response plan: who investigates, contains, corrects and approves restart?
- Separate specifications: maintain the legitimate requirement independently of the statistical limits.
- Assess stability before capability: do not reduce multiple process states to one index without justification.
- Check distribution and dependence: examine skew, mixtures, autocorrelation and changing exposure.
- Monitor the measurement system: detect instrument, classification and data-pipeline changes.
- Govern baseline changes: freeze useful limits during monitoring and deliberately create new baselines after verified redesign.
- Review usefulness: a chart that nobody acts on, or that signals continuously without learning, needs redesign.
31. What SPC does not do
- It does not decide the engineering or customer specification.
- It does not prove the root cause of a signal.
- It does not make poor measurements reliable.
- It does not make incomparable cases comparable.
- It does not guarantee normality.
- It does not replace domain expertise.
- It does not prove that a stable process is acceptable.
- It does not prove that a capable process will remain capable indefinitely.
- It does not turn a dashboard into a causal explanation.
Those limitations are a strength when they are stated clearly. A good method knows which job it owns and which job must be handed to another part of the evidence system.
32. Continue through the eduKateSingapore Library
- How Scientific Measurement Works — measurands, traceability, uncertainty and comparable evidence.
- How Statistical Inference and Uncertainty Work — sampling, estimation, hypothesis tests and uncertainty.
- How Multiple Testing and Sequential Analysis Work — repeated looks, many signals and error control.
- Data Quality — accuracy, completeness, consistency, timeliness, validity and trust.
- Project Quality Management — how requirements, verification and acceptance protect delivered work.
- Research Collections Directory — routes into the wider evidence and methods estate.
33. Sources and evidence boundary
This article is an educational synthesis and worked guide. Its fictional datasets and calculations are constructed to teach reasoning. They are not production-control recommendations for a particular factory, laboratory, healthcare service or regulated process. A real control plan requires the relevant technical specifications, measurement evidence, process knowledge and competent authority.
- NIST/SEMATECH, e-Handbook of Statistical Methods: Process or Product Monitoring and Control. General SPC, control-chart, capability, acceptance-sampling and time-series reference.
- NIST/SEMATECH, What Are Control Charts?. Core anatomy and purpose of univariate and multivariate control charts.
- NIST/SEMATECH, What Are Variables Control Charts?. Shewhart-chart formulation and three-sigma convention.
- NIST/SEMATECH, What Is Process Capability?. Relationship between in-control process behaviour and specifications.
- NIST/SEMATECH, Assessing Process Stability and Assessing Process Capability. Stability and capability in production-process characterisation.
- ISO, ISO 7870-1:2019 — Control charts — Part 1: General guidelines. Current published general-guidelines edition at the time of this article; ISO reports it was reviewed and confirmed in 2025, with a future revision under development.
- ISO, ISO 7870-2:2023 — Control charts — Part 2: Shewhart control charts. Current Shewhart-control-chart standard referenced for the chart family.
- ISO, ISO 7870-6:2024 — Control charts — Part 6: EWMA control charts for the process mean. Current EWMA-specific standard referenced for small-to-moderate shift detection.
Source status and ISO publication pages were checked for this edition on 16 September 2026. Standards can be revised; a real implementation should verify the edition required by its contract, regulator, quality system or technical context.
Final compression: statistical process control works when time-ordered measurements are used to distinguish ordinary process variation from evidence of change, signals trigger disciplined investigation rather than reflex adjustment, capability is assessed only against a credible stable process, and every chart remains connected to the real mechanism, measurement system, requirement and response plan it was built to govern.
Continue through the World Systems Directory for connected systems, governance and evidence routes.
