# Pytorch 1.8.0 fasterrcnn\_resnet50\_fpn error

**URL:** <https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150>\
**Category:** autograd\
**Created:** [March 9, 2021, 3:45am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150 "2021-03-09T03:45:54Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kitsunetic](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/kitsunetic/32/35509_2.png) [@Kitsunetic](https://discuss.pytorch.org/u/Kitsunetic)\
**Post date:** [March 9, 2021, 3:45am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/1 "2021-03-09T03:45:54Z")

</div>

My environment is

- OS: Ubuntu 18.04
- GPU: RTX3090
- CUDA: CUDA11.2
- Pytorch 1.8.0\_with\_CUDA11.1 stable

```auto
from torchvision.models.detection import fasterrcnn_resnet50_fpn

box_model = fasterrcnn_resnet50_fpn(pretrained=True, progress=False).cuda()
xs = torch.rand(2, 3, 1080, 1920, dtype=torch.float32).cuda()
ys = [
  {
    "labels": torch.tensor([1], dtype=torch.int64).cuda(),
    "boxes": torch.tensor([[956.0000, 316.3117, 1134.0000, 838.8275]], 
                          dtype=torch.float32).cuda(),
  },
  {
    "labels": torch.tensor([1], dtype=torch.int64).cuda(),
    "boxes": torch.tensor([[956.0000, 316.3117, 1134.0000, 838.8275]], 
                          dtype=torch.float32).cuda(),
  },
]

box_model(xs, ys)

```

It occurs error like this.

```auto
---------------------------------------------------------------------------
RuntimeError Traceback (most recent call last)
<ipython-input-13-7f582a050256> in <module>
----> 1 box_model(xs, ys)

~/anaconda3/envs/torch/lib/python3.7/site-packages/torch/nn/modules/module.py in _call_impl(self, *input, **kwargs)
    887 result = self._slow_forward(*input, **kwargs)
    888 else:
--> 889 result = self.forward(*input, **kwargs)
    890 for hook in itertools.chain(
    891 _global_forward_hooks.values(),

~/anaconda3/envs/torch/lib/python3.7/site-packages/torchvision/models/detection/generalized_rcnn.py in forward(self, images, targets)
     95 if isinstance(features, torch.Tensor):
     96 features = OrderedDict([('0', features)])
---> 97 proposals, proposal_losses = self.rpn(images, features, targets)
     98 detections, detector_losses = self.roi_heads(features, proposals, images.image_sizes, targets)
     99 detections = self.transform.postprocess(detections, images.image_sizes, original_image_sizes)

~/anaconda3/envs/torch/lib/python3.7/site-packages/torch/nn/modules/module.py in _call_impl(self, *input, **kwargs)
    887 result = self._slow_forward(*input, **kwargs)
    888 else:
--> 889 result = self.forward(*input, **kwargs)
    890 for hook in itertools.chain(
    891 _global_forward_hooks.values(),

~/anaconda3/envs/torch/lib/python3.7/site-packages/torchvision/models/detection/rpn.py in forward(self, images, features, targets)
    363 regression_targets = self.box_coder.encode(matched_gt_boxes, anchors)
    364 loss_objectness, loss_rpn_box_reg = self.compute_loss(
--> 365 objectness, pred_bbox_deltas, labels, regression_targets)
    366 losses = {
    367 "loss_objectness": loss_objectness,

~/anaconda3/envs/torch/lib/python3.7/site-packages/torchvision/models/detection/rpn.py in compute_loss(self, objectness, pred_bbox_deltas, labels, regression_targets)
    294 """
    295 
--> 296 sampled_pos_inds, sampled_neg_inds = self.fg_bg_sampler(labels)
    297 sampled_pos_inds = torch.where(torch.cat(sampled_pos_inds, dim=0))[0]
    298 sampled_neg_inds = torch.where(torch.cat(sampled_neg_inds, dim=0))[0]

~/anaconda3/envs/torch/lib/python3.7/site-packages/torchvision/models/detection/_utils.py in __call__ (self, matched_idxs)
     55 # randomly select positive and negative examples
     56 perm1 = torch.randperm(positive.numel(), device=positive.device)[:num_pos]
---> 57 perm2 = torch.randperm(negative.numel(), device=negative.device)[:num_neg]
     58 
     59 pos_idx_per_image = positive[perm1]

RuntimeError: radix_sort: failed on 1st step: cudaErrorInvalidDevice: invalid device ordinal

```

When I try this code on CPU, it works fine.  
After that, I reinstalled `Pytorch 1.7.1_with_CUDA11.0 stable`, it works fine too.

---

<div class="post-metadata">

**Author:** ![albanD](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/alband/32/215_2.png) [@albanD](https://discuss.pytorch.org/u/albanD)\
**Post date:** [March 9, 2021, 5:01pm UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/2 "2021-03-09T17:01:12Z")

</div>

cc @ptrblck do you know where this could be coming from?

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [March 10, 2021, 12:16am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/3 "2021-03-10T00:16:58Z")

</div>

Haven’t seen this error, but let me reproduce it on a 3090.

EDIT: I was able to reproduce it with the 1.8.0+CUDA11.1 conda binaries and will debug it further.  
It’s not failing in a source build, so my first guess is to look into CUB/Thrust.

---

<div class="post-metadata">

**Author:** ![namirinz](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/namirinz/32/35696_2.png) [@namirinz](https://discuss.pytorch.org/u/namirinz)\
**Post date:** [March 13, 2021, 10:00am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/4 "2021-03-13T10:00:33Z")

</div>

Did you just solve the problem. I faces the same now.

```auto
~/torch_env/lib/python3.8/site-packages/torchvision/models/detection/rpn.py in forward(self, images, features, targets)
    362 labels, matched_gt_boxes = self.assign_targets_to_anchors(anchors, targets)
    363 regression_targets = self.box_coder.encode(matched_gt_boxes, anchors)
--> 364 loss_objectness, loss_rpn_box_reg = self.compute_loss(
    365 objectness, pred_bbox_deltas, labels, regression_targets)
    366 losses = {

~/torch_env/lib/python3.8/site-packages/torchvision/models/detection/rpn.py in compute_loss(self, objectness, pred_bbox_deltas, labels, regression_targets)
    294 """
    295 
--> 296 sampled_pos_inds, sampled_neg_inds = self.fg_bg_sampler(labels)
    297 sampled_pos_inds = torch.where(torch.cat(sampled_pos_inds, dim=0))[0]
    298 sampled_neg_inds = torch.where(torch.cat(sampled_neg_inds, dim=0))[0]

~/torch_env/lib/python3.8/site-packages/torchvision/models/detection/_utils.py in __call__ (self, matched_idxs)
     55 # randomly select positive and negative examples
     56 perm1 = torch.randperm(positive.numel(), device=positive.device)[:num_pos]
---> 57 perm2 = torch.randperm(negative.numel(), device=negative.device)[:num_neg]
     58 
     59 pos_idx_per_image = positive[perm1]

RuntimeError: radix_sort: failed on 1st step: cudaErrorInvalidDevice: invalid device ordinal

```

OS: Ubuntu 20.04  
GPU: RTX 3080  
package: pytorch-1.8.0 with CUDA 11.1

---

<div class="post-metadata">

**Author:** ![Kitsunetic](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/kitsunetic/32/35509_2.png) [@Kitsunetic](https://discuss.pytorch.org/u/Kitsunetic)\
**Post date:** [March 14, 2021, 10:24am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/5 "2021-03-14T10:24:25Z")

</div>

I’m just using Pytorch 1.7.1 now.

Thank you

---

<div class="post-metadata">

**Author:** ![Lucas\_Tom](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/lucas_tom/32/36193_2.png) [@Lucas\_Tom](https://discuss.pytorch.org/u/Lucas_Tom)\
**Post date:** [March 24, 2021, 4:33am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/6 "2021-03-24T04:33:14Z")

</div>

I got the same issue[Kitsunetic], what should I do?

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [March 24, 2021, 5:14am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/7 "2021-03-24T05:14:53Z")

</div>

It should be fixed already in the nightly conda binary and pip wheel.  
Could you update and check it, please?

CC @Kitsunetic @namirinz

---

<div class="post-metadata">

**Author:** ![Space\_Boy](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/space_boy/32/31639_2.png) [@Space\_Boy](https://discuss.pytorch.org/u/Space_Boy)\
**Post date:** [March 28, 2021, 2:11am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/8 "2021-03-28T02:11:24Z")

</div>

I have the same issue with

```auto
print(torch. __version__ ) 
1.8.0+cu111

```

I have a 3090 RTX and torch 1.8 with cuda 11.1 is the only one compatible with Detectron2. Any idea when it will be fixed @ptrblck ?  
Thank you

---

<div class="post-metadata">

**Author:** ![MCvin](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/mcvin/32/36580_2.png) [@MCvin](https://discuss.pytorch.org/u/MCvin)\
**Post date:** [April 1, 2021, 3:22pm UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/9 "2021-04-01T15:22:17Z")

</div>

I’m having the same issue with pytorch 1.8.1 and cuda 11.1.  
The error is different though:

```auto
~/anaconda3/envs/pytorch-1.8.1/lib/python3.8/site-packages/torchvision/models/detection/_utils.py in __call__ (self, matched_idxs)
     43 neg_idx = []
     44 for matched_idxs_per_image in matched_idxs:
---> 45 positive = torch.where(matched_idxs_per_image >= 1)[0]
     46 negative = torch.where(matched_idxs_per_image == 0)[0]
     47 

RuntimeError: CUDA error: device-side assert triggered

```

I works with:

- pytorch 1.8.1 + cuda 10.2
- pytorch 1.7.1 + cuda 11.0

Seems like cuda 11.1 is the problem here.  
Sorry I couldn’t help more.

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [April 2, 2021, 5:50am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/10 "2021-04-02T05:50:25Z")

</div>

The radix\_sort is already fixed in the nightly release and PyTorch `1.8.1`, so you would have to update to one of these versions.

@MCvin your error seems to be different.  
Could you rerun the code via `CUDA_LAUNCH_BLOCKING=1 python setup.py args` and post the complete stack trace (or create a new topic with your error and this information)?

---

<div class="post-metadata">

**Author:** ![gabridego](https://discuss.pytorch.org/letter_avatar_proxy/v4/letter/g/ee59a6/32.png) [@gabridego](https://discuss.pytorch.org/u/gabridego)\
**Post date:** [April 2, 2021, 3:45pm UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/11 "2021-04-02T15:45:57Z")

</div>

After updating to 1.8.1 I’m not getting the radix\_sort error anymore, but I’m getting the same error as @MCvin. I’m using the stable version with CUDA 11.1 on Ubuntu 18.04. Everything works when run on CPU.

```auto
  File "~/anaconda3/envs/neo/lib/python3.9/site-packages/pl_bolts/models/detection/faster_rcnn/faster_rcnn_module.py", line 112, in training_step
    loss_dict = self.model(images, targets)
  File "~/anaconda3/envs/neo/lib/python3.9/site-packages/torch/nn/modules/module.py", line 889, in _call_impl
    result = self.forward(*input, **kwargs)
  File "~/anaconda3/envs/neo/lib/python3.9/site-packages/torchvision/models/detection/generalized_rcnn.py", line 97, in forward
    proposals, proposal_losses = self.rpn(images, features, targets)
  File "~/anaconda3/envs/neo/lib/python3.9/site-packages/torch/nn/modules/module.py", line 889, in _call_impl
    result = self.forward(*input, **kwargs)
  File "~/anaconda3/envs/neo/lib/python3.9/site-packages/torchvision/models/detection/rpn.py", line 364, in forward
    loss_objectness, loss_rpn_box_reg = self.compute_loss(
  File "~/anaconda3/envs/neo/lib/python3.9/site-packages/torchvision/models/detection/rpn.py", line 296, in compute_loss
    sampled_pos_inds, sampled_neg_inds = self.fg_bg_sampler(labels)
  File "~/anaconda3/envs/neo/lib/python3.9/site-packages/torchvision/models/detection/_utils.py", line 46, in __call__
    positive = torch.where(matched_idxs_per_image >= 1)[0]
RuntimeError: CUDA error: device-side assert triggered

```

The error follows a long series of messages like:

```auto
/pytorch/aten/src/ATen/native/cuda/IndexKernel.cu:142: operator(): block: [0,0,0], thread: [0,0,0] Assertion `index >= -sizes[i] && index < sizes[i] && "index out of bounds"` failed.

```

I guess the problem is with the `torch.where` operation in `torchvision/models/detection/_utils.py`. I’ve been able to run my program moving the variable `matched_idxs_per_image` to CPU in the ` __call__ ` function of class `BalancedPositiveNegativeSampler`, like:

```auto
for matched_idxs_per_image in matched_idxs:
    matched_idxs_per_image = matched_idxs_per_image.cpu()
    positive = torch.where(matched_idxs_per_image >= 1)[0]
    negative = torch.where(matched_idxs_per_image == 0)[0]

```

from line 44.

Hope that helps.

---

<div class="post-metadata">

**Author:** ![ppwwyyxx](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ppwwyyxx/32/3031_2.png) [@ppwwyyxx](https://discuss.pytorch.org/u/ppwwyyxx)\
**Post date:** [April 2, 2021, 8:53pm UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/12 "2021-04-02T20:53:33Z")

</div>

detectron2 is also seeing an increasing number of CUDA error reports on CUDA\>=11.1 + pytorch 1.8.x + RTX30xx: [RuntimeError: CUDA error: device-side assert triggered · Issue #2837 · facebookresearch/detectron2 · GitHub](https://github.com/facebookresearch/detectron2/issues/2837)

Root cause seems to be still randperm: [CUDA error: device-side assert triggered(torch1.8.1+cuda11.1) · Issue #55027 · pytorch/pytorch · GitHub](https://github.com/pytorch/pytorch/issues/55027)

---

<div class="post-metadata">

**Author:** ![chrischoy](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/chrischoy/32/1612_2.png) [@chrischoy](https://discuss.pytorch.org/u/chrischoy)\
**Post date:** [April 8, 2021, 12:55am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/13 "2021-04-08T00:55:44Z")

</div>

I’m facing the same issue on MinokowskiEngine with pytorch 1.8.X + CUDA 11.X. [Cuda 11.1 - Coordinate manager · Issue #330 · NVIDIA/MinkowskiEngine (github.com)](https://github.com/NVIDIA/MinkowskiEngine/issues/330)

---

<div class="post-metadata">

**Author:** ![oo\_o](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/oo_o/32/37243_2.png) [@oo\_o](https://discuss.pytorch.org/u/oo_o)\
**Post date:** [April 17, 2021, 5:37pm UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/14 "2021-04-17T17:37:33Z")

</div>

Same here, seems to be related to randperm()

Reproducible code:

```auto
>>> import torch
>>> device = torch.device("cuda:0")
>>> torch.randperm(29999, device=device)
tensor([13324, 19251, 23333, ..., 18540, 14502, 26766], device='cuda:0')
>>> torch.randperm(30000, device=device)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  RuntimeError: radix_sort: failed on 1st step: cudaErrorInvalidDevice: invalid device ordinal

```

pytorch version: 1.8.0+cu111

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [April 19, 2021, 7:17am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/15 "2021-04-19T07:17:31Z")

</div>

Could you update PyTorch to `1.8.1` or the nightly as described in my previous post, please?

---

<div class="post-metadata">

**Author:** ![Kitsunetic](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/kitsunetic/32/35509_2.png) [@Kitsunetic](https://discuss.pytorch.org/u/Kitsunetic)\
**Post date:** [April 19, 2021, 10:07am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/16 "2021-04-19T10:07:06Z")

</div>

Tested on 1.8.1 and nightly with same environment.  
At 1.8.1, the error still happened but nightly works fine.  
Thank you

---

<div class="post-metadata">

**Author:** ![kehuantiantang](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/kehuantiantang/32/30001_2.png) [@kehuantiantang](https://discuss.pytorch.org/u/kehuantiantang)\
**Post date:** [April 29, 2021, 10:13am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/17 "2021-04-29T10:13:39Z")

</div>

Same problem happened when install nightly version by pip wheel, 1.9.0.dev20210428+cu111, python 3.8.8

```auto
  File "/opt/conda/envs/fpn/lib/python3.8/site-packages/torchvision/models/detection/rpn.py", line 363, in forward
    loss_objectness, loss_rpn_box_reg = self.compute_loss(
  File "/opt/conda/envs/fpn/lib/python3.8/site-packages/torchvision/models/detection/rpn.py", line 295, in compute_loss
    sampled_pos_inds, sampled_neg_inds = self.fg_bg_sampler(labels)
  File "/opt/conda/envs/fpn/lib/python3.8/site-packages/torchvision/models/detection/_utils.py", line 45, in __call__
    positive = torch.where(matched_idxs_per_image >= 1)[0]
RuntimeError: CUDA error: device-side assert triggered

```

```auto
/pytorch/aten/src/ATen/native/cuda/IndexKernel.cu:97: operator(): block: [0,0,0], thread: [2,0,0] Assertion `index >= -sizes[i] && index < sizes[i] && "index out of bounds"` failed.

```

Same code, and problem solved by **pytorch 1.7.1 + cuda 11.0** , I think this may be the problem of cuda 11.1

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [April 29, 2021, 5:28pm UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/18 "2021-04-29T17:28:44Z")

</div>

Did you check the indices, which create the error?  
If so, what are the min and max values of the indices and what is the shape of the indexed tensor?

---

<div class="post-metadata">

**Author:** ![pierlj](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/pierlj/32/32891_2.png) [@pierlj](https://discuss.pytorch.org/u/pierlj)\
**Post date:** [May 3, 2021, 9:59am UTC](https://discuss.pytorch.org/t/pytorch-1-8-0-fasterrcnn-resnet50-fpn-error/114150/19 "2021-05-03T09:59:47Z")

</div>

I had the same error message as @MCvin with pytorch 1.8.1 running faster r-cnn code. I tried a few things to reproduce it. It seems to work fine for small value of `n`. I tried this while running FRCNN code through vscode debugging. Yet, I was not able to reproduce it outside like @oo_o did.

![pytorch_issue](https://discuss.pytorch.org/uploads/default/original/3X/9/9/9941135cb514f581631d18b272337673af8b0875.png)

It was fixed by updating with nightly build ‘1.9.0.dev20210502’.

OS: Ubuntu 20.04  
GPU: RTX 3090  
Pytorch 1.8.1 / Cuda 11.1
