# Does batch size have any effect on divergence of training alogorithm?

**URL:** <https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245>\
**Category:** Uncategorized\
**Created:** [April 25, 2017, 11:06am UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245 "2017-04-25T11:06:40Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![mderakhshani](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mderakhshani/32/1735_2.png) [@mderakhshani](https://discuss.pytorch.org/u/mderakhshani)\
**Post date:** [April 25, 2017, 11:06am UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245/1 "2017-04-25T11:06:40Z")

</div>

I want to implement the yolo (You look only once) in Pytorch. I wrote its code and set its batch size to 64. But when I ran the algorithm, its cost always increased. But when I set the batch size to 32, the cost decreased in a long term. Could you please tell me Is it logical or not?

---

<div class="post-metadata">

**Author:** ![Mika\_S](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mika_s/32/14253_2.png) [@Mika\_S](https://discuss.pytorch.org/u/Mika_S)\
**Post date:** [April 26, 2017, 6:55pm UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245/2 "2017-04-26T18:55:13Z")

</div>

Smaller batchsize gives the gradients sufficient noise to jump out of valleys. That being said 64 is not that big size. Are you seeing at the loss per image or the total loss (which would be higher for higher batch size).

---

<div class="post-metadata">

**Author:** ![mderakhshani](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mderakhshani/32/1735_2.png) [@mderakhshani](https://discuss.pytorch.org/u/mderakhshani)\
**Post date:** [April 27, 2017, 2:52pm UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245/3 "2017-04-27T14:52:10Z")

</div>

Thanks for your response @Mika_S! I have found that one of my code line had some problem which it was a `NAN` value. I would like to know, have you ever read the **yolo** paper?

---

<div class="post-metadata">

**Author:** ![Mika\_S](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mika_s/32/14253_2.png) [@Mika\_S](https://discuss.pytorch.org/u/Mika_S)\
**Post date:** [April 28, 2017, 8:58am UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245/4 "2017-04-28T08:58:24Z")

</div>

I have read yolo paper an year back. But shoot me questions and I can try to answer as best as I can :).

---

<div class="post-metadata">

**Author:** ![mderakhshani](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mderakhshani/32/1735_2.png) [@mderakhshani](https://discuss.pytorch.org/u/mderakhshani)\
**Post date:** [April 28, 2017, 4:15pm UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245/5 "2017-04-28T16:15:52Z")

</div>

I have a question about its cost function implementation. Because the source code was provided by C language and I am so newbie in c. I would like to know, have you got any implementation about the cost in python? I have implemented it but i think the yolo’s authors use some tricks to obtain the best answer which they are unknown. Is it possible to collaborate with each other to provide the python implementation of that?

I have started a post in google group of darknet (Yolo basic framework) [link](https://groups.google.com/forum/#!topic/darknet/lvwDatqfIU4), but did not get any answers.

---

<div class="post-metadata">

**Author:** ![Mika\_S](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mika_s/32/14253_2.png) [@Mika\_S](https://discuss.pytorch.org/u/Mika_S)\
**Post date:** [May 6, 2017, 12:06am UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245/6 "2017-05-06T00:06:20Z")

</div>

Sorry for the late reply. Unfortunately I do not have an implementation of the cost in python.

---

<div class="post-metadata">

**Author:** ![Mika\_S](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mika_s/32/14253_2.png) [@Mika\_S](https://discuss.pytorch.org/u/Mika_S)\
**Post date:** [May 6, 2017, 12:57am UTC](https://discuss.pytorch.org/t/does-batch-size-have-any-effect-on-divergence-of-training-alogorithm/2245/7 "2017-05-06T00:57:14Z")

</div>

Have you looked at this: [https://github.com/longcw/yolo2-pytorch/blob/master/darknet.py](https://github.com/longcw/yolo2-pytorch/blob/master/darknet.py)
