EDUCATION SUBJECT ATLAS · DATA SCIENCE
What Is Data Science?
Data science is the integrated practice of turning data into useful evidence, explanations, predictions and decisions. It combines statistics, computer science, domain knowledge, data engineering, visualisation and model evaluation to answer questions that are too large, complex or uncertain for intuition alone.
Data science is not simply running machine-learning algorithms. A useful model depends on the quality of the question, the provenance of the data, the meaning of variables, the design of validation, the conditions of deployment and the decision that will follow. A technically impressive model can be useless if it optimises the wrong target.
Data science is the discipline of preserving meaning while data moves from messy reality into a model and back into a decision.
The data-science pipeline
- Define the problem and decision.
- Identify the population, entities and outcomes.
- Acquire data from trustworthy sources.
- Clean, join and validate records.
- Explore distributions, missingness and anomalies.
- Engineer features or representations.
- Choose statistical or machine-learning methods.
- Train models under controlled conditions.
- Validate against unseen data.
- Deploy with monitoring, governance and feedback.
Real projects are iterative. Discovery in step seven can reveal a flaw in step two, forcing the team back to redefine the problem.
Start with the decision, not the dataset
Organisations often begin with “we have this data; what can we do with it?” A stronger approach begins with the decision or question. What action will change if the analysis succeeds? What error is costly? What prediction horizon matters? Who will use the output?
A model that predicts churn tomorrow may be operationally useless if customer-support interventions require two weeks to organise. Decision timing belongs inside model design.
Data provenance
Provenance records where data came from, how it was collected, how it changed and who is responsible for it. Without provenance, a dataset may look authoritative while containing unknown transformations or stale definitions.
Data science needs lineage because models inherit every upstream assumption. A field called “active customer” is only useful if the definition is known and stable.
Data quality
Data quality includes completeness, accuracy, consistency, timeliness, uniqueness and validity. Quality is contextual: data can be accurate enough for monthly reporting but too delayed for real-time fraud detection.
Good data science defines quality thresholds according to the decision rather than pursuing abstract perfection.
Cleaning
Cleaning identifies malformed values, duplicates, inconsistent categories, impossible dates, unit mismatches and other defects. The difficult part is deciding whether a strange value is an error or a real rare event.
A good cleaning pipeline preserves an audit trail. Silent deletion destroys information about the data-generating process.
Joining data
Many data-science projects combine multiple datasets. Joining requires stable identifiers or probabilistic matching. Incorrect joins can create duplicate entities, missing relationships or false histories.
Entity resolution is therefore a modelling problem in its own right. “Same name” does not always mean “same person,” and different names can still refer to the same entity.
Missing data
Missing values carry information about the system. A field may be absent because a sensor failed, a customer declined to answer, a process changed or the value was never relevant.
Imputation should follow an understanding of the missingness mechanism. Replacing every missing value with an average can erase structure and understate uncertainty.
Exploratory data analysis
Exploratory data analysis uses summaries and visualisation to understand distributions, relationships, outliers and data-quality issues before formal modelling.
Exploration is not merely making charts. It is the process of discovering what the dataset actually contains rather than what the schema claims it contains.
Visualisation
Visualisation compresses complex patterns into forms humans can inspect. Scatterplots show relationships, histograms show distributions, line charts reveal temporal change and maps reveal spatial structure.
Visual design can also mislead through truncated axes, inappropriate scales or overplotting. The graph is part of the analysis, not decoration.
Feature engineering
Features are representations used by a model. A raw timestamp may become hour of day, day of week or elapsed time. A transaction history may become frequency, recency and average value.
Feature engineering encodes domain knowledge. It can improve models dramatically, but it can also introduce leakage if features contain information unavailable at prediction time.
Data leakage
Leakage occurs when information from the future or target outcome enters model training in a way that would not be available during real use. The model then appears highly accurate in testing but collapses in production.
Leakage is a systems error, not just a mathematical one. Validation must reproduce the information boundary of the real decision.
Supervised learning
Supervised learning trains models using examples with known outcomes. Regression predicts continuous values; classification predicts categories or probabilities.
Common methods include linear models, decision trees, random forests, gradient boosting and neural networks. Method choice depends on data structure, scale, interpretability, latency and risk.
Unsupervised learning
Unsupervised learning searches for structure without the same target labels. Clustering groups similar observations; dimensionality reduction compresses high-dimensional data; anomaly detection identifies unusual patterns.
Unsupervised results require careful interpretation because an algorithm can always partition data somehow. The question is whether the discovered structure is stable and useful.
Prediction versus inference
Prediction asks what will happen. Inference asks what relationship or mechanism generated the pattern. A model can predict extremely well using variables that are not causal.
Data science must therefore define whether the goal is forecasting, ranking, explanation, intervention or causal effect estimation.
Training, validation and test data
Training data fit the model. Validation data guide model choice and tuning. Test data provide a final estimate of performance on unseen examples.
Repeatedly inspecting the test set turns it into another validation set. A true final evaluation must remain protected from iterative tuning.
Cross-validation
Cross-validation repeatedly splits data into training and validation portions to estimate how performance varies across samples. It is useful when data are limited.
Splits must respect structure. Time-series forecasting should not train on future observations to predict the past, and grouped data may require entire groups to remain together.
Metrics
Model metrics should match the decision. Accuracy can be misleading when classes are imbalanced. Precision, recall, specificity, F-scores, calibration, ranking metrics and cost-sensitive measures capture different trade-offs.
A fraud system and a medical screening system may value false positives and false negatives differently. The metric must reflect consequence.
Calibration
A probability model is calibrated when events predicted with probability 0.7 occur roughly seventy percent of the time across comparable cases. Calibration matters when probabilities drive decisions rather than simple ranking.
A highly discriminative model can still be poorly calibrated.
Bias and variance
Simple models may underfit by missing real structure; highly flexible models may overfit noise. The bias-variance trade-off describes this tension.
Regularisation, validation and appropriate model complexity help control it.
Interpretability
Interpretability concerns whether humans can understand why a model produced an output or which factors drive predictions. Linear models can be easier to interpret than deep neural networks, though even simple coefficients can be misunderstood when variables interact.
Interpretability is especially important when decisions affect rights, safety, employment, credit or healthcare.
Explainability tools
Feature importance, partial dependence and local explanation methods can help inspect complex models. These tools explain model behaviour, not necessarily real-world causation.
“The model relied on this feature” is not the same as “this feature causes the outcome.”
Causal data science
Causal analysis asks what would change under an intervention. Randomised experiments are powerful when feasible; observational settings require stronger assumptions and designs such as natural experiments, matching or causal graphs.
A predictive model can identify who is likely to fail. A causal model asks who would improve because of a particular intervention. Those are different questions.
Data engineering
Data engineering builds the pipelines, storage systems and interfaces that make reliable analysis possible. It covers ingestion, transformation, orchestration, schemas, lineage, quality and access.
Most production data science depends on engineering. A brilliant notebook that cannot receive fresh trusted data is not an operational system.
Batch and streaming data
Batch systems process groups of records periodically. Streaming systems process events continuously or near-real-time. The right architecture depends on latency needs, cost and operational complexity.
Real-time processing should be justified by a real-time decision. Faster systems create additional failure modes.
Databases and warehouses
Operational databases support transactions; analytical warehouses organise historical data for analysis. Data lakes and lakehouse architectures support larger, more varied collections.
Architecture should follow access patterns, governance and reliability requirements rather than fashion.
Deployment
A deployed model becomes part of an operating system. Predictions must arrive with suitable latency, interfaces, authentication and fallback behaviour.
Deployment also changes accountability. A model that influences decisions needs owners, monitoring and defined intervention when performance degrades.
Model drift
Model drift occurs when relationships change after deployment. Customer behaviour, policy, technology or data collection can shift, making yesterday’s model less reliable.
Monitoring should track input distributions, performance, calibration and business outcomes rather than assuming a model remains valid forever.
MLOps
MLOps applies software-engineering and operational practices to machine learning. Versioning, testing, deployment, monitoring, reproducibility and rollback become part of the model lifecycle.
The goal is to transform a model from an experiment into a controlled production capability.
Reproducibility
A reproducible data-science project records data versions, code, dependencies, parameters and environment. Without this, a result may be impossible to regenerate months later.
Reproducibility is part of knowledge quality, not administrative housekeeping.
Privacy
Data science can expose sensitive information through raw data, linkage or model outputs. Privacy must be considered during collection, access, analysis and release.
Collecting less unnecessary personal data can reduce risk while simplifying governance.
Fairness
Models can produce unequal error rates across groups because historical data reflect unequal systems, sample coverage differs or the target itself encodes problematic assumptions.
Fairness cannot be reduced to one universal metric. Different definitions can conflict, so the appropriate choice depends on context and normative goals.
Data ethics
Ethical data science asks not only whether analysis is technically possible, but whether collection and use are appropriate, transparent and proportionate to benefit.
High predictive accuracy does not justify every deployment.
Generative AI and data science
Generative AI expands data-science workflows through coding assistance, unstructured-data processing, embeddings and natural-language interfaces. It also introduces new evaluation problems because outputs are probabilistic and may sound convincing when wrong.
Data science contributes the evaluation discipline needed to measure reliability, calibration, cost and failure modes.
Decision science
Prediction only becomes valuable when it changes a decision. Decision science connects model outputs with actions, costs, benefits and constraints.
The best model is not always the most accurate. A simpler model may create more value if it is faster, cheaper, more interpretable or easier to operate.
Experimentation
A/B testing and controlled experiments help evaluate interventions. Good experimentation defines outcomes before analysis, randomises correctly and protects against repeated peeking and selective reporting.
Experimentation closes the loop between model prediction and real-world effect.
A systems model of data science
- ENTITY: users, records, events, datasets, models and systems.
- STATE: feature values, model versions, schemas, distributions and deployment conditions.
- OCCURRENCE: collection, transformation, training, prediction, intervention and feedback.
- RELATIONSHIP: joins, causal links, dependencies, hierarchies and data lineage.
- INTENT: question, target, metric, decision and risk tolerance.
- OBSERVATION: source data, labels, logs, experiments and monitoring signals.
- ARTIFACT: datasets, pipelines, notebooks, models, APIs and dashboards.
- CLAIM: predictions, explanations, estimates and recommendations.
- VOID: missing data, unmeasured causes, distribution shift and unknown failure states.
structured systems analysis data science requires every output to retain its route back to source data, model version, evaluation evidence and operating context. The result is not released as knowledge until that chain is inspectable.
How to think like a data scientist
- Start from the decision.
- Define the target and population.
- Trace data provenance.
- Inspect quality and missingness.
- Separate prediction from causation.
- Protect validation from leakage.
- Choose metrics that match consequence.
- Test robustness across groups and time.
- Plan deployment and monitoring before launch.
- Document what the model does not know.
Common misconceptions
- “Data science is machine learning.” Machine learning is one component of a larger evidence and decision pipeline.
- “More data always improves a model.” More biased or irrelevant data can worsen outcomes.
- “A high accuracy score proves a useful model.” The metric may not match the decision.
- “Prediction tells us what intervention will work.” Causal effects require different evidence.
- “Deployment is an engineering afterthought.” Operating context determines whether a model creates value safely.
Mini case: predicting student difficulty
A model may predict which students are likely to struggle using prior scores, attendance and assignment patterns. But if the objective is to decide who benefits most from tutoring, prediction alone is insufficient. The highest-risk student may not be the student with the largest treatment response.
Data science distinguishes the prediction question from the intervention question before using the output operationally.
Mini case: a fraud model that looks excellent offline
A fraud model may score extremely well because a training feature was added only after an investigation concluded. In production that feature does not exist at decision time. This is leakage.
The correct validation reconstructs the information available at the precise moment the decision must be made.
Data science across the learning journey
Young learners can begin with tables, charts, simple coding and questions about evidence. Secondary learners can study statistics, spreadsheets, databases and introductory machine learning. Advanced study adds probability, optimisation, data engineering, causal inference, machine learning, distributed systems and responsible AI.
The progression is from reading data to designing complete systems that turn observations into decisions responsibly.
Why data science belongs inside education
Data science teaches learners to connect evidence, computation and consequence. It combines mathematical reasoning with practical system design and makes model limitations part of the answer.
As organisations increasingly automate analysis and decisions, data literacy must include not only how to build models but how to know when they should be trusted.
