AI systems need information that matches the job you want them to do. For a machine-learning model, that information usually arrives as a dataset: a collection of examples used to learn and check a pattern. Understanding the examples is more useful than memorizing an algorithm name. A model trained on the wrong data can be precise on paper and still fail for real users.
This lesson explains examples, features, labels, and data quality through one simple task: predicting whether a library book will be returned late.
Start with the decision
Suppose a library wants to send a friendly reminder before a book is due. The task is to estimate whether a loan is likely to be late. First define the output: late or on time.
Each past loan becomes one example. A row may contain the loan length, whether the due date falls on a holiday, and the eventual return outcome. The data must describe the situation as it existed before the reminder would be sent. If you include the actual return date, the model sees the answer in advance and the exercise becomes meaningless.
| Past loan | Loan length | Holiday near due date? | Late? |
|---|---|---|---|
| A | 7 days | No | No |
| B | 14 days | Yes | Yes |
| C | 21 days | No | No |
The first three columns describe examples; the last column records the known outcome. This is only a small illustration, not enough data to train a reliable model.
What are features and labels?
A feature is an input value the model can use. In the table, loan length and the holiday indicator are candidate features. A label is the answer supplied during supervised training. Here, the label is whether the book was late. When the model receives a new loan, it sees the features but not the future label; it must predict that outcome.
Features can be numbers, categories, text, images, or signals derived from them. A category such as “Yes” or “No” often needs a numerical representation before a model can use it. That conversion must be learned or defined consistently and applied the same way when you make a new prediction. A label can also be a number: if the task were to estimate how many days late a book will be, the target would be numeric rather than a yes/no category.
Tip: Ask whether each proposed feature will truly be available when you need the prediction. A feature that exists only after the outcome is known cannot help a live system.
Check the data before training
Useful data is relevant, accurate, and representative. Check for duplicate loans, impossible dates, inconsistent holiday values, and missing outcomes. A blank value is not automatically zero. For example, “holiday unknown” differs from “no holiday.” You may correct a verified data-entry error, exclude an unusable example, or fill a missing value using a documented rule. Record what you did so later results remain understandable.
Look at who and what the data covers. If you train only on short loans from one branch, the model may work poorly on long loans or another branch. A dataset can also preserve historical unfairness. Avoid collecting personal details merely because they are available. Use the minimum information needed for the task, and review whether the reminder process treats readers appropriately.
A worked example of data leakage
Imagine adding a column named days_after_due_date. For past loans this value strongly reveals whether a return was late. A model might score almost perfectly during testing if that column is present. But on the day you want to send a reminder, the number of days after the due date does not yet exist. This is data leakage: information unavailable at prediction time reaches the model during development. Remove that column and repeat the evaluation.
Leakage can be less obvious. If the same loan appears in both training and testing data, the test no longer measures performance on a truly new example. If you calculate a data-cleaning statistic from the entire dataset before splitting it, the training process may indirectly learn about the test portion. Keep your test examples separate and apply transformations with care.
Prepare a small dataset step by step
- Define the task. Write the decision, prediction time, and intended user benefit in one sentence.
- Collect permitted examples. Include the outcome and only inputs available at prediction time.
- Inspect rows and columns. Count missing values, duplicates, inconsistent categories, and unusual values.
- Choose features. Keep relevant signals; remove answers in disguise and unnecessary personal data.
- Separate examples. Reserve unseen data for evaluation before tuning the model.
- Document changes. Record how you cleaned data and why you kept each feature.
Key points
A useful dataset has examples with features available at prediction time and trustworthy labels. Check data quality and keep test examples separate so evaluation reflects new cases.