Reducing an Overshooting SGD Learning Rate to Stop Oscillating Training and Validation Loss

Train and refine models.
Answer Correct answer: D — An oscillating loss curve means SGD overshoots the minimum, so decreasing the learning rate lets the model converge instead of bouncing around it.

An ML engineer has trained a neural network by using stochastic gradient descent (SGD). The neural network performs poorly on the test set. The values for training loss and validation loss remain high and show an oscillating pattern. The values decrease for a few epochs and then increase for a few epochs before repeating the same cycle. What should the ML engineer do to improve the training process?

  1. Introduce early stopping.
  2. Increase the size of the test set.
  3. Increase the learning rate.
  4. Decrease the learning rate. Correct Answer

Community Votes

D
100%

100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

An oscillating loss curve is the signature of a learning rate that is too high: each step overshoots the minimum and the optimizer bounces around it. Reducing the learning rate makes the updates smaller so the loss settles at the minimum.

A neural network trained with stochastic gradient descent performs poorly on the test set, and the training and validation loss stay high while oscillating, decreasing for a few epochs and then increasing before the cycle repeats. Loss that fails to settle and swings around indicates the optimizer is overshooting rather than converging.

Reaching for early stopping because the loss eventually rises, but early stopping halts training to prevent overfitting rather than fixing an optimizer that never converges in the first place. Increasing the learning rate would make the oscillation worse.

Community Discussion (5 comments)

motk123 👍 6 Selected: D
The oscillating pattern of training and validation loss indicates that the learning rate is too high. A high learning rate causes the model to overshoot the optimal point in the loss landscape, leading to oscillations instead of convergence. Reducing the learning rate allows the model to make smaller, more precise updates to the weights, improving convergence. A. Early stopping prevents overfitting by halting training when validation performance stops improving. However, it does not address the root cause of oscillating loss. B. The size of the test set does not affect the training dynamics or loss patterns. C. Increasing the learning rate would worsen the oscillations and prevent the model from converging.
ninomfr64 👍 4 Selected: D
A. No, early stopping is for preventing overfitting B. No, increasing test will not help with oscillating loss C. No, increasing learning rate will make things worsening D. Oscillating loss in training is a sign that the training is not converging, this can happen when learning rate is too high. Reducing learning rate will help here
gulf1324 👍 1 Selected: D
Oscillating patterns in a train/validation loss shows it's not converging to a minima. Low learning rate will make it converge.
feelgoodfactor 👍 2 Selected: D
The oscillating loss during training is a clear sign that the learning rate is too high. Reducing the learning rate will stabilize the optimization process, allowing the model to converge smoothly.
GiorgioGss 👍 2 Selected: D
oscillating = decrease the learning rate

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The defining symptom is oscillation: the loss decreases for a few epochs, then increases, and the cycle repeats instead of converging. That is the classic signature of a learning rate that is too high, because each SGD step overshoots the minimum and the optimizer oscillates around it rather than settling. Lowering the learning rate shrinks the step size so the updates become precise enough to descend into the minimum and stay there, which lowers both the training and validation loss and improves test performance. The vote was unanimous at 100 for D. motk123 identified overshooting of the optimal point in the loss landscape as the cause, ninomfr64 characterized the oscillating loss as non-convergence, and gulf1324 noted that a low learning rate lets the model converge.

Why the Other Options Are Wrong

Introducing early stopping (A) addresses overfitting by halting training when validation loss stops improving, but the model here never converged in the first place, so stopping early would leave it in a high-loss state rather than fix the optimization problem; ninomfr64 dismissed it on exactly this basis. Increasing the size of the test set (B) changes evaluation only and has no effect on the training dynamics, so the oscillating loss would persist unchanged. Increasing the learning rate (C) would push the optimizer further past the minimum and intensify the oscillation rather than dampening it, which is the opposite of what the curve calls for.

Community Comment Notes

The community was unanimous at 100 for D, and every substance comment independently identified the learning rate as the cause. motk123 gave the most complete answer, explaining that a high learning rate causes overshooting and that early stopping serves a different purpose. feelgoodfactor and GiorgioGss each stated the fix in a single line, and gulf1324 connected the oscillating pattern to failure to converge to a minimum, which is the same diagnosis expressed as optimizer behavior rather than as a hyperparameter value.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide