Overfitting and Underfitting in Machine Learning

Artificial Intelligence Tutorials


A machine-learning model should work on new examples, not just remember the examples used to build it. Underfitting and overfitting are two common reasons it fails. We can recognize both by comparing performance on training data with performance on separate validation data.

What is underfitting?

An underfit model has not learned enough of the useful pattern. It performs poorly on its training examples and on new examples. Imagine predicting delivery time using only the day of the week, while distance and traffic have a much larger effect. A model without the important inputs may be too simple for the task.

Underfitting can come from insufficient training, unhelpful features, or a model that cannot represent the relationship. Before making a model more complex, check that the data and labels are correct and that relevant inputs are available at prediction time.

What is overfitting?

An overfit model matches its training examples closely but does poorly on unfamiliar examples. It may learn accidental details rather than a rule that generalizes. For instance, if every late delivery in a tiny training set happens to have an order number ending in 7, a model might rely on that digit even though it has no real link to delivery time.

Overfitting is especially easy when a flexible model has few or unrepresentative examples. More training steps can also make a model fit noise. A low training error, by itself, is not proof of a useful model.

Compare training and validation results

The following percentages are an illustration, not results from a real dataset. Error is the share of wrong predictions, so lower is better.

Model Training error Validation error Likely interpretation
A 25% 27% Underfitting: both errors are high
B 8% 9% Better fit: both errors are lower and close
C 1% 18% Overfitting: training is excellent, validation is much worse

Do not judge the gap alone. A model with 49% training error and 50% validation error has a small gap but is still poor at a balanced two-class task.

As model complexity increases, training error falls while validation error eventually rises.
As model complexity increases, training error falls while validation error eventually rises.

In this schematic, each move to the right represents a more complex model, not more training time. Training error falls, while validation error first falls and then rises as complexity becomes excessive. The lowest validation error marks a useful range of complexity for these illustrative models.

How can you improve an underfit model?

  • Check the problem and labels. Incorrect labels make any model harder to learn from.
  • Add relevant features that are available when the prediction will be made.
  • Train long enough, if the model has not converged.
  • Try a model capable of representing a more useful relationship.

Change one thing at a time and measure the result on validation data. A more complex model is not automatically better.

How can you reduce overfitting?

  • Get more representative training examples when possible; simply copying rows does not add information.
  • Remove noisy or irrelevant features and consider a simpler model.
  • Use regularization to discourage unnecessarily large or complex patterns.
  • Use early stopping: stop training when validation performance stops improving.
  • Check for data leakage, such as an input that reveals the answer or duplicate records across splits.

These methods are not substitutes for a valid evaluation. Keep a separate test set untouched while choosing features, model settings, and stopping points. Use it once to estimate how your final choice performs on new data.

Practical check

Suppose a shop predicts whether an order will arrive late. Training error is 2% and validation error is 20%. First check that both sets represent the same intended task and that neither contains leaked answers. Then try a simpler model or stronger regularization. If training and validation errors are both around 25%, investigate the features, labels, and model capacity instead.

The goal is generalization: useful predictions for future orders. Track both training and validation results, explain why a feature should matter, and resist choosing a model only because it looks perfect on past data.



Found This Page Useful? Share It!
Get the Latest Tutorials and Updates
Join us on Telegram