Reducing an Overshooting SGD Learning Rate to Stop Oscillating Training and Validation Loss
An ML engineer has trained a neural network by using stochastic gradient descent (SGD). The neural network performs poorly on the test set. The values for training loss and validation loss remain high and show an oscillating pattern. The values decrease for a few epochs and then increase for a few epochs before repeating the same cycle. What should the ML engineer do to improve the training process?
Community Votes
100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
An oscillating loss curve is the signature of a learning rate that is too high: each step overshoots the minimum and the optimizer bounces around it. Reducing the learning rate makes the updates smaller so the loss settles at the minimum.
A neural network trained with stochastic gradient descent performs poorly on the test set, and the training and validation loss stay high while oscillating, decreasing for a few epochs and then increasing before the cycle repeats. Loss that fails to settle and swings around indicates the optimizer is overshooting rather than converging.
Reaching for early stopping because the loss eventually rises, but early stopping halts training to prevent overfitting rather than fixing an optimizer that never converges in the first place. Increasing the learning rate would make the oscillation worse.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The defining symptom is oscillation: the loss decreases for a few epochs, then increases, and the cycle repeats instead of converging. That is the classic signature of a learning rate that is too high, because each SGD step overshoots the minimum and the optimizer oscillates around it rather than settling. Lowering the learning rate shrinks the step size so the updates become precise enough to descend into the minimum and stay there, which lowers both the training and validation loss and improves test performance. The vote was unanimous at 100 for D. motk123 identified overshooting of the optimal point in the loss landscape as the cause, ninomfr64 characterized the oscillating loss as non-convergence, and gulf1324 noted that a low learning rate lets the model converge.Why the Other Options Are Wrong
Introducing early stopping (A) addresses overfitting by halting training when validation loss stops improving, but the model here never converged in the first place, so stopping early would leave it in a high-loss state rather than fix the optimization problem; ninomfr64 dismissed it on exactly this basis. Increasing the size of the test set (B) changes evaluation only and has no effect on the training dynamics, so the oscillating loss would persist unchanged. Increasing the learning rate (C) would push the optimizer further past the minimum and intensify the oscillation rather than dampening it, which is the opposite of what the curve calls for.Community Comment Notes
The community was unanimous at 100 for D, and every substance comment independently identified the learning rate as the cause. motk123 gave the most complete answer, explaining that a high learning rate causes overshooting and that early stopping serves a different purpose. feelgoodfactor and GiorgioGss each stated the fix in a single line, and gulf1324 connected the oscillating pattern to failure to converge to a minimum, which is the same diagnosis expressed as optimizer behavior rather than as a hyperparameter value.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →