Three learners review open books together at a classroom table, with stacks of textbooks, stationery and a whiteboard in the bright room.

What Is Data Science? | Data, Models, Evidence, Prediction and Decision Systems

EDUCATION SUBJECT ATLAS · DATA SCIENCE

What Is Data Science?

Data science is the integrated practice of turning data into useful evidence, explanations, predictions and decisions. It combines statistics, computer science, domain knowledge, data engineering, visualisation and model evaluation to answer questions that are too large, complex or uncertain for intuition alone.

Data science is not simply running machine-learning algorithms. A useful model depends on the quality of the question, the provenance of the data, the meaning of variables, the design of validation, the conditions of deployment and the decision that will follow. A technically impressive model can be useless if it optimises the wrong target.

Data science is the discipline of preserving meaning while data moves from messy reality into a model and back into a decision.

The data-science pipeline

  1. Define the problem and decision.
  2. Identify the population, entities and outcomes.
  3. Acquire data from trustworthy sources.
  4. Clean, join and validate records.
  5. Explore distributions, missingness and anomalies.
  6. Engineer features or representations.
  7. Choose statistical or machine-learning methods.
  8. Train models under controlled conditions.
  9. Validate against unseen data.
  10. Deploy with monitoring, governance and feedback.

Real projects are iterative. Discovery in step seven can reveal a flaw in step two, forcing the team back to redefine the problem.

Start with the decision, not the dataset

Organisations often begin with “we have this data; what can we do with it?” A stronger approach begins with the decision or question. What action will change if the analysis succeeds? What error is costly? What prediction horizon matters? Who will use the output?

A model that predicts churn tomorrow may be operationally useless if customer-support interventions require two weeks to organise. Decision timing belongs inside model design.

Data provenance

Provenance records where data came from, how it was collected, how it changed and who is responsible for it. Without provenance, a dataset may look authoritative while containing unknown transformations or stale definitions.

Data science needs lineage because models inherit every upstream assumption. A field called “active customer” is only useful if the definition is known and stable.

Data quality

Data quality includes completeness, accuracy, consistency, timeliness, uniqueness and validity. Quality is contextual: data can be accurate enough for monthly reporting but too delayed for real-time fraud detection.

Good data science defines quality thresholds according to the decision rather than pursuing abstract perfection.

Cleaning

Cleaning identifies malformed values, duplicates, inconsistent categories, impossible dates, unit mismatches and other defects. The difficult part is deciding whether a strange value is an error or a real rare event.

A good cleaning pipeline preserves an audit trail. Silent deletion destroys information about the data-generating process.

Joining data

Many data-science projects combine multiple datasets. Joining requires stable identifiers or probabilistic matching. Incorrect joins can create duplicate entities, missing relationships or false histories.

Entity resolution is therefore a modelling problem in its own right. “Same name” does not always mean “same person,” and different names can still refer to the same entity.

Missing data

Missing values carry information about the system. A field may be absent because a sensor failed, a customer declined to answer, a process changed or the value was never relevant.

Imputation should follow an understanding of the missingness mechanism. Replacing every missing value with an average can erase structure and understate uncertainty.

Exploratory data analysis

Exploratory data analysis uses summaries and visualisation to understand distributions, relationships, outliers and data-quality issues before formal modelling.

Exploration is not merely making charts. It is the process of discovering what the dataset actually contains rather than what the schema claims it contains.

Visualisation

Visualisation compresses complex patterns into forms humans can inspect. Scatterplots show relationships, histograms show distributions, line charts reveal temporal change and maps reveal spatial structure.

Visual design can also mislead through truncated axes, inappropriate scales or overplotting. The graph is part of the analysis, not decoration.

Feature engineering

Features are representations used by a model. A raw timestamp may become hour of day, day of week or elapsed time. A transaction history may become frequency, recency and average value.

Feature engineering encodes domain knowledge. It can improve models dramatically, but it can also introduce leakage if features contain information unavailable at prediction time.

Data leakage

Leakage occurs when information from the future or target outcome enters model training in a way that would not be available during real use. The model then appears highly accurate in testing but collapses in production.

Leakage is a systems error, not just a mathematical one. Validation must reproduce the information boundary of the real decision.

Supervised learning

Supervised learning trains models using examples with known outcomes. Regression predicts continuous values; classification predicts categories or probabilities.

Common methods include linear models, decision trees, random forests, gradient boosting and neural networks. Method choice depends on data structure, scale, interpretability, latency and risk.

Unsupervised learning

Unsupervised learning searches for structure without the same target labels. Clustering groups similar observations; dimensionality reduction compresses high-dimensional data; anomaly detection identifies unusual patterns.

Unsupervised results require careful interpretation because an algorithm can always partition data somehow. The question is whether the discovered structure is stable and useful.

Prediction versus inference

Prediction asks what will happen. Inference asks what relationship or mechanism generated the pattern. A model can predict extremely well using variables that are not causal.

Data science must therefore define whether the goal is forecasting, ranking, explanation, intervention or causal effect estimation.

Training, validation and test data

Training data fit the model. Validation data guide model choice and tuning. Test data provide a final estimate of performance on unseen examples.

Repeatedly inspecting the test set turns it into another validation set. A true final evaluation must remain protected from iterative tuning.

Cross-validation

Cross-validation repeatedly splits data into training and validation portions to estimate how performance varies across samples. It is useful when data are limited.

Splits must respect structure. Time-series forecasting should not train on future observations to predict the past, and grouped data may require entire groups to remain together.

Metrics

Model metrics should match the decision. Accuracy can be misleading when classes are imbalanced. Precision, recall, specificity, F-scores, calibration, ranking metrics and cost-sensitive measures capture different trade-offs.

A fraud system and a medical screening system may value false positives and false negatives differently. The metric must reflect consequence.

Calibration

A probability model is calibrated when events predicted with probability 0.7 occur roughly seventy percent of the time across comparable cases. Calibration matters when probabilities drive decisions rather than simple ranking.

A highly discriminative model can still be poorly calibrated.

Bias and variance

Simple models may underfit by missing real structure; highly flexible models may overfit noise. The bias-variance trade-off describes this tension.

Regularisation, validation and appropriate model complexity help control it.

Interpretability

Interpretability concerns whether humans can understand why a model produced an output or which factors drive predictions. Linear models can be easier to interpret than deep neural networks, though even simple coefficients can be misunderstood when variables interact.

Interpretability is especially important when decisions affect rights, safety, employment, credit or healthcare.

Explainability tools

Feature importance, partial dependence and local explanation methods can help inspect complex models. These tools explain model behaviour, not necessarily real-world causation.

“The model relied on this feature” is not the same as “this feature causes the outcome.”

Causal data science

Causal analysis asks what would change under an intervention. Randomised experiments are powerful when feasible; observational settings require stronger assumptions and designs such as natural experiments, matching or causal graphs.

A predictive model can identify who is likely to fail. A causal model asks who would improve because of a particular intervention. Those are different questions.

Data engineering

Data engineering builds the pipelines, storage systems and interfaces that make reliable analysis possible. It covers ingestion, transformation, orchestration, schemas, lineage, quality and access.

Most production data science depends on engineering. A brilliant notebook that cannot receive fresh trusted data is not an operational system.

Batch and streaming data

Batch systems process groups of records periodically. Streaming systems process events continuously or near-real-time. The right architecture depends on latency needs, cost and operational complexity.

Real-time processing should be justified by a real-time decision. Faster systems create additional failure modes.

Databases and warehouses

Operational databases support transactions; analytical warehouses organise historical data for analysis. Data lakes and lakehouse architectures support larger, more varied collections.

Architecture should follow access patterns, governance and reliability requirements rather than fashion.

Deployment

A deployed model becomes part of an operating system. Predictions must arrive with suitable latency, interfaces, authentication and fallback behaviour.

Deployment also changes accountability. A model that influences decisions needs owners, monitoring and defined intervention when performance degrades.

Model drift

Model drift occurs when relationships change after deployment. Customer behaviour, policy, technology or data collection can shift, making yesterday’s model less reliable.

Monitoring should track input distributions, performance, calibration and business outcomes rather than assuming a model remains valid forever.

MLOps

MLOps applies software-engineering and operational practices to machine learning. Versioning, testing, deployment, monitoring, reproducibility and rollback become part of the model lifecycle.

The goal is to transform a model from an experiment into a controlled production capability.

Reproducibility

A reproducible data-science project records data versions, code, dependencies, parameters and environment. Without this, a result may be impossible to regenerate months later.

Reproducibility is part of knowledge quality, not administrative housekeeping.

Privacy

Data science can expose sensitive information through raw data, linkage or model outputs. Privacy must be considered during collection, access, analysis and release.

Collecting less unnecessary personal data can reduce risk while simplifying governance.

Fairness

Models can produce unequal error rates across groups because historical data reflect unequal systems, sample coverage differs or the target itself encodes problematic assumptions.

Fairness cannot be reduced to one universal metric. Different definitions can conflict, so the appropriate choice depends on context and normative goals.

Data ethics

Ethical data science asks not only whether analysis is technically possible, but whether collection and use are appropriate, transparent and proportionate to benefit.

High predictive accuracy does not justify every deployment.

Generative AI and data science

Generative AI expands data-science workflows through coding assistance, unstructured-data processing, embeddings and natural-language interfaces. It also introduces new evaluation problems because outputs are probabilistic and may sound convincing when wrong.

Data science contributes the evaluation discipline needed to measure reliability, calibration, cost and failure modes.

Decision science

Prediction only becomes valuable when it changes a decision. Decision science connects model outputs with actions, costs, benefits and constraints.

The best model is not always the most accurate. A simpler model may create more value if it is faster, cheaper, more interpretable or easier to operate.

Experimentation

A/B testing and controlled experiments help evaluate interventions. Good experimentation defines outcomes before analysis, randomises correctly and protects against repeated peeking and selective reporting.

Experimentation closes the loop between model prediction and real-world effect.

A systems model of data science

structured systems analysis data science requires every output to retain its route back to source data, model version, evaluation evidence and operating context. The result is not released as knowledge until that chain is inspectable.

How to think like a data scientist

  1. Start from the decision.
  2. Define the target and population.
  3. Trace data provenance.
  4. Inspect quality and missingness.
  5. Separate prediction from causation.
  6. Protect validation from leakage.
  7. Choose metrics that match consequence.
  8. Test robustness across groups and time.
  9. Plan deployment and monitoring before launch.
  10. Document what the model does not know.

Common misconceptions

Mini case: predicting student difficulty

A model may predict which students are likely to struggle using prior scores, attendance and assignment patterns. But if the objective is to decide who benefits most from tutoring, prediction alone is insufficient. The highest-risk student may not be the student with the largest treatment response.

Data science distinguishes the prediction question from the intervention question before using the output operationally.

Mini case: a fraud model that looks excellent offline

A fraud model may score extremely well because a training feature was added only after an investigation concluded. In production that feature does not exist at decision time. This is leakage.

The correct validation reconstructs the information available at the precise moment the decision must be made.

Data science across the learning journey

Young learners can begin with tables, charts, simple coding and questions about evidence. Secondary learners can study statistics, spreadsheets, databases and introductory machine learning. Advanced study adds probability, optimisation, data engineering, causal inference, machine learning, distributed systems and responsible AI.

The progression is from reading data to designing complete systems that turn observations into decisions responsibly.

Why data science belongs inside education

Data science teaches learners to connect evidence, computation and consequence. It combines mathematical reasoning with practical system design and makes model limitations part of the answer.

As organisations increasingly automate analysis and decisions, data literacy must include not only how to build models but how to know when they should be trusted.

Continue through the subject atlas

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate SG

Subscribe now to keep reading and get access to the full archive.

Continue reading