# Is there a benefit to cache images for data loading?

**URL:** <https://discuss.pytorch.org/t/is-there-a-benefit-to-cache-images-for-data-loading/174597>\
**Category:** data\
**Created:** [March 11, 2023, 9:11pm UTC](https://discuss.pytorch.org/t/is-there-a-benefit-to-cache-images-for-data-loading/174597 "2023-03-11T21:11:09Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Hui\_Liu](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/hui_liu/32/57030_2.png) [@Hui\_Liu](https://discuss.pytorch.org/u/Hui_Liu)\
**Post date:** [March 11, 2023, 9:11pm UTC](https://discuss.pytorch.org/t/is-there-a-benefit-to-cache-images-for-data-loading/174597/1 "2023-03-11T21:11:09Z")

</div>

Hi,

I see for yolov7 there is the default option to [cache all images](https://github.com/WongKinYiu/yolov7/blob/main/utils/datasets.py#L450) in memory for faster training.

However, based on [this thread](https://discuss.pytorch.org/t/multi-process-data-loading-and-prefetching/98997), `torch.utils.data.dataloader.DataLoader` will, by default, do prefetching? If so, the effect of caching images on training speed would be very minimal? Or I misunderstand how the prefetching works?

Thanks!!

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [March 11, 2023, 9:31pm UTC](https://discuss.pytorch.org/t/is-there-a-benefit-to-cache-images-for-data-loading/174597/2 "2023-03-11T21:31:27Z")

</div>

It would depend on the data loading speed and if you would see any bottlenecks or general wait times during the training if the `DataLoader` is used. If so, creating a cache might give you benefits assuming the cached images avoid the bottleneck in the first place.  
However, you would need to check if the cache uses your system RAM (and if you would have enough) or the disk, as the latter might not give you huge benefits.

---

<div class="post-metadata">

**Author:** ![Hui\_Liu](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/hui_liu/32/57030_2.png) [@Hui\_Liu](https://discuss.pytorch.org/u/Hui_Liu)\
**Post date:** [March 11, 2023, 10:29pm UTC](https://discuss.pytorch.org/t/is-there-a-benefit-to-cache-images-for-data-loading/174597/3 "2023-03-11T22:29:09Z")

</div>

I see. Thanks, Patrick! Yeah, I see slow-up without cache, but my whole dataset is too large to fit into the memory. So I was wondering if I could just prefetch and cache only the next batch. But then I see we already do that in dataloader?

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [March 12, 2023, 7:07am UTC](https://discuss.pytorch.org/t/is-there-a-benefit-to-cache-images-for-data-loading/174597/4 "2023-03-12T07:07:54Z")

</div>

Yes, the `DataLoader` will prefetch the next batches, but will not cache it.  
I don’t know if your Yolo code could cache only a subset of the dataset to avoid running out of memory, but a quick check of the linked code seems to show all samples would be added to the cache.
