# Loss suddenly increases using Adam optimizer

**URL:** <https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338>\
**Category:** Uncategorized\
**Created:** [December 19, 2017, 12:02pm UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338 "2017-12-19T12:02:03Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![zhangboknight](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/zhangboknight/32/15722_2.png) [@zhangboknight](https://discuss.pytorch.org/u/zhangboknight)\
**Post date:** [December 19, 2017, 12:02pm UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/1 "2017-12-19T12:02:03Z")

</div>

Hi, I came across a problem when using Adam optimizer. At the start of the training, the loss decreases as expected. But after 3300 iterations, the loss suddenly explodes to a very large number(~1e3). I tried several times but the same problem occurs. How to solve this issue? Thanks!

![Capture](https://discuss.pytorch.org/uploads/default/original/2X/4/474e84a2ffd063c1edee109611a06b0a02b5f940.PNG)

---

<div class="post-metadata">

**Author:** ![munkiti](https://discuss.pytorch.org/letter_avatar_proxy/v4/letter/m/8e7dd6/32.png) [@munkiti](https://discuss.pytorch.org/u/munkiti)\
**Post date:** [December 19, 2017, 12:46pm UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/2 "2017-12-19T12:46:04Z")

</div>

Did you restart at 3300’th iteration? Or did you run it all along?. I think you need to give more info. what is the problem you are working with? What’s the step-size?

In anycase, the problem with Adam is that it uses moving average in the denominator term. So if the gradients get really small and the whole of denominator will be small. Since the gradients are already small, the denominator results in blowup thus pushing you very far away hence huge loss. You may have a look at [https://openreview.net/forum?id=ryQu7f-RZ](https://openreview.net/forum?id=ryQu7f-RZ) . I think there are many recent methods which avert this problem including AMSGrad (in the earlier mentioned paper), Hyper-gradient descent (for Adam) etc. Also look for comments in the openreview forum, there seems to be further discussion on this issue. Unfortunately i do not know of any pytorch implementation of these algorithms.

---

<div class="post-metadata">

**Author:** ![zhangboknight](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/zhangboknight/32/15722_2.png) [@zhangboknight](https://discuss.pytorch.org/u/zhangboknight)\
**Post date:** [December 19, 2017, 2:22pm UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/3 "2017-12-19T14:22:01Z")

</div>

Thanks a lot for your detailed reply, Munkiti. I run my training all along without any restart. The learning rate for Adam is 1e-3. The network is typical resnet structure.

I will check whether the problem comes from the small denominator with Adam. I will post it when I find a solution.

---

<div class="post-metadata">

**Author:** ![zhangboknight](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/zhangboknight/32/15722_2.png) [@zhangboknight](https://discuss.pytorch.org/u/zhangboknight)\
**Post date:** [December 20, 2017, 6:50am UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/4 "2017-12-20T06:50:09Z")

</div>

As suggestion, I replace the Adam optimizer with AMSGrad. The problem is solved^^ It indeed comes from the stabilization issue of the Adam itself.

In implementation, I reinstall my pytorch from source and in version 4.0, I can simply use AMSGrad with:  
`optimizer = optim.Adam(model.parameters(), lr=0.001, eps=1e-3, amsgrad=True)`

Thanks for your help very much!

---

<div class="post-metadata">

**Author:** ![Rakshit\_Kothari](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/rakshit_kothari/32/1974_2.png) [@Rakshit\_Kothari](https://discuss.pytorch.org/u/Rakshit_Kothari)\
**Post date:** [November 26, 2018, 10:28pm UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/5 "2018-11-26T22:28:10Z")

</div>

While AMSGrad really improves the training loss curve and it seems to progress for a longer number of epochs, but after certain number of epochs, even AMSGrad tends to increase training loss

 ![image](https://discuss.pytorch.org/uploads/default/original/2X/5/58d3996a673991003d5d3de86a76731fee82f2e8.png)

---

<div class="post-metadata">

**Author:** ![yj\_z](https://discuss.pytorch.org/letter_avatar_proxy/v4/letter/y/ce7236/32.png) [@yj\_z](https://discuss.pytorch.org/u/yj_z)\
**Post date:** [January 13, 2019, 7:10am UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/6 "2019-01-13T07:10:33Z")

</div>

Did you solve this problem?

---

<div class="post-metadata">

**Author:** ![Rakshit\_Kothari](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/rakshit_kothari/32/1974_2.png) [@Rakshit\_Kothari](https://discuss.pytorch.org/u/Rakshit_Kothari)\
**Post date:** [January 15, 2019, 5:26pm UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/7 "2019-01-15T17:26:17Z")

</div>

Well, kind of. As the training performance improves, I linearly reduce the learning rate (learning rate at perfect performance is 1/10th the original LR). This significantly combats this tendency to overshoot.

---

<div class="post-metadata">

**Author:** ![ryo](https://discuss.pytorch.org/letter_avatar_proxy/v4/letter/r/a587f6/32.png) [@ryo](https://discuss.pytorch.org/u/ryo)\
**Post date:** [January 27, 2019, 11:08am UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/8 "2019-01-27T11:08:23Z")

</div>

Have you tried just simply clipping the gradient?

---

<div class="post-metadata">

**Author:** ![Rakshit\_Kothari](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/rakshit_kothari/32/1974_2.png) [@Rakshit\_Kothari](https://discuss.pytorch.org/u/Rakshit_Kothari)\
**Post date:** [April 23, 2019, 8:54pm UTC](https://discuss.pytorch.org/t/loss-suddenly-increases-using-adam-optimizer/11338/9 "2019-04-23T20:54:40Z")

</div>

Yes, I use a combination of gradient clipping and batch normalization which has pretty much ensured that this never occurs again.
