Training, Validation, and Test Data

Artificial Intelligence Tutorials


A model can memorize its examples and still fail on new ones. To check whether it has learned a useful pattern, keep some data away from the training step. The usual roles are training data for fitting, validation data for choosing settings, and test data for a final check.

Training fits the model, validation helps choose settings, and test data checks the final choice.
Training fits the model, validation helps choose settings, and test data checks the final choice.

Why one dataset is not enough

Suppose a model learns from 100 past loans and correctly predicts all 100. That score tells you it handled loans it already saw. It does not tell you whether it can predict a new loan. If the model learned accidental details of the training rows, its apparent success may disappear on fresh data. This is overfitting.

Separate examples before you compare models. Training data changes the model. Validation data helps you choose features, settings, or a model design. Test data stays untouched until you have made those choices. If you repeatedly adjust the model after seeing test results, the test set becomes another validation set and loses its value as a final check.

Set What you use it for What to avoid
Training Fit patterns or parameters Reporting its score as proof of real-world quality
Validation Compare candidate settings during development Assuming repeated tuning cannot overfit this set
Test Check the chosen approach on unseen examples Using its answers to keep redesigning the model

Choose a split that matches the future

There is no universal percentage that works for every project. A common teaching example divides data into roughly 70% training, 15% validation, and 15% test, but the right sizes depend on how many examples you have and how reliable the estimate needs to be. Each set should contain enough relevant cases. If one class is rare, check that the evaluation sets still include it.

A random split is not always appropriate. If several rows belong to the same person, device, or transaction, keep related rows together so the model cannot recognize the same case on both sides of the split. If you will predict the future from the past, a time-based split may better reflect deployment: train on earlier events and evaluate on later ones. Remove duplicates and prevent any feature from revealing an outcome that was unavailable at prediction time.

Tip: Fit data-cleaning steps, such as a mean used to replace missing values, using training data only. Apply the resulting rule unchanged to validation, test, and future inputs.

Run a small Python example

The following program shuffles 15 toy loans with a fixed seed, makes three separate sets, and uses a simple nearest-neighbor rule. The rule looks at loans with similar lengths in the training set. Validation chooses whether to use one or three neighbors; the test set is checked only after that choice. You can run the example and change a few labels to see how much a tiny test score can move.

Example:

from random import Random

# Toy data: (days until due date, late return). This is for learning splits,
# not a reliable model of real library behavior.
loans = [
    (2, 1), (3, 1), (4, 1), (5, 1), (6, 1),
    (7, 0), (8, 0), (9, 0), (10, 0), (11, 0),
    (12, 0), (13, 0), (14, 0), (15, 0), (16, 0),
]

# Shuffle once so the original row order does not define the split.
Random(7).shuffle(loans)
train, validation, test = loans[:9], loans[9:12], loans[12:]


def predict(days, training, neighbors):
    # Find the nearest training examples by loan length.
    nearest = sorted(training, key=lambda row: abs(row[0] - days))[:neighbors]
    # More than half of the nearest examples must have a late label.
    return sum(late for _, late in nearest) > neighbors / 2


def accuracy(rows, training, neighbors):
    # Compare predictions with known labels in a separate set.
    correct = sum(predict(days, training, neighbors) == bool(late)
                  for days, late in rows)
    return correct / len(rows)


# Compare two settings on validation data, not on the final test data.
choices = (1, 3)
best = max(choices, key=lambda value: accuracy(validation, train, value))

print(f"Training examples: {len(train)}")
print(f"Validation examples: {len(validation)}")
print(f"Test examples: {len(test)}")
print(f"Chosen neighbors: {best}")
print(f"Validation accuracy: {accuracy(validation, train, best):.0%}")
# Use the test set only once after choosing the setting.
print(f"Test accuracy: {accuracy(test, train, best):.0%}")

Output:

Training examples: 9
Validation examples: 3
Test examples: 3
Chosen neighbors: 1
Validation accuracy: 100%
Test accuracy: 67%

The difference between 100% on validation and 67% on test is not a contradiction. With only three examples in each set, one mistake changes a score by about 33 percentage points. The demo teaches separation of roles, not a trustworthy accuracy estimate. A real project needs far more representative examples, uncertainty checks, and review of the types of mistakes that matter.

Read results without fooling yourself

Accuracy is only the fraction of correct predictions. It may hide important errors. In a late-return reminder, a false alarm may cause an unnecessary message, while a missed late return may mean no reminder. Examine those errors separately and decide what matters to readers. For an imbalanced dataset, a model can earn high accuracy by predicting the common class every time. Precision and recall can reveal more about its positive predictions and missed positives.

Compare the model with a simple baseline, such as always predicting the most common outcome. Inspect performance across relevant groups and time periods. Data can change after deployment; monitor whether inputs and results still resemble what you tested. If the model affects important decisions, add appropriate human review and safeguards.

A practical evaluation sequence

  1. Define the prediction time and target. Know what information the system can use and what outcome it should predict.
  2. Reserve a final test set. Keep related records together and use a time-based split where that matches the real task.
  3. Fit on training data. Learn model parameters and preprocessing rules without looking at validation or test outcomes.
  4. Choose with validation data. Compare candidate settings and record the decision.
  5. Evaluate once on the test set. Report relevant metrics, examples of errors, and important limitations.
  6. Keep watching after release. Reassess the system when users, data, or conditions change.

Key points

Train on one set, choose settings with another, and keep the test set for a final check. Choose a split that resembles future use, and inspect errors as well as scores.



Found This Page Useful? Share It!
Get the Latest Tutorials and Updates
Join us on Telegram