Junior Codes
Model Detectives

Data Leakage in Machine Learning: Explained for Kids

When a clue secretly gives away the answer, the score cannot be trusted.

Start Wonderwild free with Pip

Try 5 starter activities free. Full access is ₹499 a year.

This mission is part of full access. The first missions are free to try.

Data leakage, explained simply

Sometimes a model gets a brilliant score for the wrong reason: one of its features accidentally contains the answer. This is called data leakage.

Imagine training a model to spot ripe fruit, but every ripe fruit photo was taken in the kitchen and every unripe one in the garden. The model may learn "kitchen means ripe" instead of anything about the fruit.

In the real world that shortcut disappears, and the model fails. Good model detectives ask of every feature: would we really know this at the moment we need to make the prediction?

Example: A suspiciously perfect score

  • A model predicts whether a library book will be returned late.
  • One feature is "late fee paid". It gets 100% accuracy.
  • But you only know a fee was paid after the book was late. The answer leaked into the input.

Try this

  • Look at a quiz where the question gives away the answer. How is that like leakage?
  • For a "will it rain tomorrow?" model, list features you would know today, and one that would be cheating.
  • When something scores 100%, practise asking "what shortcut could it be using?"

Practise it in Wonderwild

“The answer hidden in the input”

Discovery 11 in Model Detectives, one of 24 AI discoveries in Wonderwild.

Your child trains, tests and questions a small model in the browser with Pip. No coding needed. For ages 7+.

Start Wonderwild free with Pip

Try 5 starter activities free. Full access is ₹499 a year.

This mission is part of full access. The first missions are free to try.

Questions parents and kids ask

Why is a perfect score sometimes a warning sign?+

Real-world problems are messy, so a perfect score often means the model found a shortcut, such as a leaked feature, rather than the real pattern.

Is data leakage a real problem for AI companies?+

Yes. It is one of the most common reasons a model that looks excellent in testing performs badly once people start using it.

Prefer learning with a teacher?

Our live AI Explorers course teaches these ideas in small online classes led by a software engineer, with projects and feedback. Every live course also includes 12 months of full Wonderwild access for practice.

See the AI Explorers course