# RuntimeError: CUDA error: CUBLAS\_STATUS\_INVALID\_VALUE when calling \`cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)\`

**URL:** <https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544>\
**Category:** Uncategorized\
**Created:** [June 20, 2021, 5:25am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544 "2021-06-20T05:25:06Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![vanduong0504](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/vanduong0504/32/30358_2.png) [@vanduong0504](https://discuss.pytorch.org/u/vanduong0504)\
**Post date:** [June 20, 2021, 5:25am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/1 "2021-06-20T05:25:06Z")

</div>

I am training my models from `Google Collab` with `batch_size = 128` after 1 epoch it has this problem. I don’t know have to fix it with the same batch\_size (reduce batch\_size to 32 can avoid this problem). Here is Colab spec: `driver Version: 460.32.03 CUDA Version: 11.2 `  
You can find my notebook [here](https://colab.research.google.com/drive/1Ixz0zaWUCEc_-oDN03Rib4JUR0UsKG_z?usp=sharing).  
Thanks for your help.

---

<div class="post-metadata">

**Author:** ![tom](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/tom/32/3162_2.png) [@tom](https://discuss.pytorch.org/u/tom)\
**Post date:** [June 20, 2021, 11:47am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/2 "2021-06-20T11:47:22Z")

</div>

It seems that one of your operands is too large to fit in int32 (or negative, but that seems unlikely).

I thought that recent PyTorch will give a better error (but don’t work around it):

```python
import torch
LARGE = 2**31+1
for i, j, k in [(1, 1, LARGE), (1, LARGE, 1), (LARGE, 1, 1)]:
    inp = torch.randn(i, k, device="cuda", dtype=torch.half)
    weight = torch.randn(j, k, device="cuda", dtype=torch.half)
    try:
        torch.nn.functional.linear(inp, weight)
    except RuntimeError as e:
        print(e)
    del inp
    del weight

```

```auto
at::cuda::blas::gemm<float> argument k must be non-negative and less than 2147483647 but got 2147483649
at::cuda::blas::gemm<float> argument m must be non-negative and less than 2147483647 but got 2147483649
at::cuda::blas::gemm<float> argument n must be non-negative and less than 2147483647 but got 2147483649

```

But they don’t work around it. (It needs a lot of memory to trigger the bug…)

Maybe you can get a credible backtrace and record the input shapes to the operation that fails.

Best regards

Thomas

---

<div class="post-metadata">

**Author:** ![vanduong0504](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/vanduong0504/32/30358_2.png) [@vanduong0504](https://discuss.pytorch.org/u/vanduong0504)\
**Post date:** [June 20, 2021, 12:34pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/3 "2021-06-20T12:34:28Z")

</div>

So what can I do to solve this problem, I just know to change batch size to smaller.

---

<div class="post-metadata">

**Author:** ![tom](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/tom/32/3162_2.png) [@tom](https://discuss.pytorch.org/u/tom)\
**Post date:** [June 20, 2021, 5:10pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/4 "2021-06-20T17:10:36Z")

</div>

In order of difficulty:

- make batch size smaller,
- make a minimal reproducing example (i.e. just two or three inputs from torch.random and the call to the torch.nn.functional.linear) and file a bug,
- hot-patch torch.nn.functional.linear with a workaround (splitting the operation into multiple linear or matmul calls),
- submit a PR with a fix in PyTorch and discuss whether you can add a test or whether it’d take a prohibitive large amount of GPU memory to run (or hire someone to do so).

Best regards

Thomas

---

<div class="post-metadata">

**Author:** ![vanduong0504](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/vanduong0504/32/30358_2.png) [@vanduong0504](https://discuss.pytorch.org/u/vanduong0504)\
**Post date:** [June 20, 2021, 6:24pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/5 "2021-06-20T18:24:07Z")

</div>

Thank for your help.

---

<div class="post-metadata">

**Author:** ![Jeremy\_Cochoy](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/jeremy_cochoy/32/12318_2.png) [@Jeremy\_Cochoy](https://discuss.pytorch.org/u/Jeremy_Cochoy)\
**Post date:** [July 17, 2021, 11:32am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/6 "2021-07-17T11:32:23Z")

</div>

For the peoples getting this error and ending up on this post, please know that it can also be caused if you have a mismatch between the dimension of your input tensor and the dimensions of your nn.Linear module. (ex. x.shape = (a, b) and nn.Linear(c, c, bias=False) with c not matching)

It is a bit sad that pytorch don’t give a more explicit error messages.

---

<div class="post-metadata">

**Author:** ![Madhavan\_Seshadri](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/madhavan_seshadri/32/40551_2.png) [@Madhavan\_Seshadri](https://discuss.pytorch.org/u/Madhavan_Seshadri)\
**Post date:** [July 22, 2021, 2:53pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/7 "2021-07-22T14:53:17Z")

</div>

@Jeremy_Cochoy This was really helpful. Solved my issue.

---

<div class="post-metadata">

**Author:** ![Hugo](https://discuss.pytorch.org/letter_avatar_proxy/v4/letter/h/b38774/32.png) [@Hugo](https://discuss.pytorch.org/u/Hugo)\
**Post date:** [July 29, 2021, 4:12am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/8 "2021-07-29T04:12:24Z")

</div>

@Jeremy_Cochoy Thanks for your comments!

---

<div class="post-metadata">

**Author:** ![martinoywa](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/martinoywa/32/27832_2.png) [@martinoywa](https://discuss.pytorch.org/u/martinoywa)\
**Post date:** [July 29, 2021, 6:55am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/9 "2021-07-29T06:55:43Z")

</div>

@Jeremy_Cochoy Thanks!

---

<div class="post-metadata">

**Author:** ![SNair](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/snair/32/29925_2.png) [@SNair](https://discuss.pytorch.org/u/SNair)\
**Post date:** [July 29, 2021, 6:25pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/10 "2021-07-29T18:25:43Z")

</div>

Hello @Jeremy_Cochoy  
I have added an nn.Linear(512,10) layer to my model and the shape of the input that goes into this layer is torch.Size([32,512,1,1]). I have tried reducing the batch size from 128 to 64 and now to 32, but each of these gives me the same error.  
Any idea what could be going wrong?

---

<div class="post-metadata">

**Author:** ![Jeremy\_Cochoy](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/jeremy_cochoy/32/12318_2.png) [@Jeremy\_Cochoy](https://discuss.pytorch.org/u/Jeremy_Cochoy)\
**Post date:** [July 29, 2021, 6:52pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/11 "2021-07-29T18:52:56Z")

</div>

I think you want to transpose the dimensions of your input tensor before and after ([Linear — PyTorch 1.9.0 documentation](https://pytorch.org/docs/stable/generated/torch.nn.Linear.html) say it expect a Nx\*xC\_in tensor and you give him a 32x…x1 tensor)

Something like `linear(x.transpose(1,3)).transpose(1,3)` ?

---

<div class="post-metadata">

**Author:** ![NicoHambauer](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/nicohambauer/32/39241_2.png) [@NicoHambauer](https://discuss.pytorch.org/u/NicoHambauer)\
**Post date:** [September 17, 2021, 10:51am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/13 "2021-09-17T10:51:26Z")

</div>

Thanks a lot also solved my Issue!

---

<div class="post-metadata">

**Author:** ![markstein](https://discuss.pytorch.org/letter_avatar_proxy/v4/letter/m/ecb155/32.png) [@markstein](https://discuss.pytorch.org/u/markstein)\
**Post date:** [September 30, 2021, 8:10am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/14 "2021-09-30T08:10:48Z")

</div>

I got the same error because of a mismatch of the input dimensions in the first layer.

Thanks for the hint!

---

<div class="post-metadata">

**Author:** ![Parshin\_Shojaee](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/parshin_shojaee/32/42727_2.png) [@Parshin\_Shojaee](https://discuss.pytorch.org/u/Parshin_Shojaee)\
**Post date:** [October 3, 2021, 5:12pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/15 "2021-10-03T17:12:38Z")

</div>

helped me! thank you!

---

<div class="post-metadata">

**Author:** ![black0017](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/black0017/32/24100_2.png) [@black0017](https://discuss.pytorch.org/u/black0017)\
**Post date:** [October 18, 2021, 1:15pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/16 "2021-10-18T13:15:00Z")

</div>

hello all,

I had this problem while I was using a smaller batch size (=4) for testing some code changes, while my initial batch size was 64. I checked the shapes for `nn.Linear` and they matched

After 1 hour I found that the only change was the batch size. By **increasing batch size back to 64** everything worked perfectly. Pytorch version 1.8.1. Not sure why this error is caused.

Hope it helps!

---

<div class="post-metadata">

**Author:** ![Chuyang](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/chuyang/32/34516_2.png) [@Chuyang](https://discuss.pytorch.org/u/Chuyang)\
**Post date:** [January 22, 2022, 5:52am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/17 "2022-01-22T05:52:40Z")

</div>

Thanks a lot! Solved my issue.

---

<div class="post-metadata">

**Author:** ![wtliao](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/wtliao/32/12750_2.png) [@wtliao](https://discuss.pytorch.org/u/wtliao)\
**Post date:** [December 2, 2022, 1:13pm UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/18 "2022-12-02T13:13:47Z")

</div>

Hello all,

I am using pytorch ‘1.13.0+cu117’, my env is NVIDIA-SMI 450.80.02 Driver Version: 450.80.02 CUDA Version: 11.0

in the terminal of python, I tried the very simple example:

```auto
>>> import torch
>>> x=torch.ones(2,2,1).to('cuda')
>>> y=torch.ones(2,1,2).to('cuda')
>>> x
tensor([[[1.],
         [1.]],

        [[1.],
         [1.]]], device='cuda:0')
>>> y
tensor([[[1., 1.]],

        [[1., 1.]]], device='cuda:0')
>>> y@x
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
RuntimeError: CUDA error: CUBLAS_STATUS_NOT_SUPPORTED when calling `cublasSgemmStridedBatched( handle, opa, opb, m, n, k, &alpha, a, lda, stridea, b, ldb, strideb, &beta, c, ldc, stridec, num_batches)`
>>> torch.bmm(y,x)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
RuntimeError: CUDA error: CUBLAS_STATUS_NOT_SUPPORTED when calling `cublasSgemmStridedBatched( handle, opa, opb, m, n, k, &alpha, a, lda, stridea, b, ldb, strideb, &beta, c, ldc, stridec, num_batches)`
>>> torch.matmul(y,x)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
RuntimeError: CUDA error: CUBLAS_STATUS_NOT_SUPPORTED when calling `cublasSgemmStridedBatched( handle, opa, opb, m, n, k, &alpha, a, lda, stridea, b, ldb, strideb, &beta, c, ldc, stridec, num_batches)`
>>> x=torch.ones(2,1).to('cuda')
>>> y=torch.ones(1,2).to('cuda')
>>> y@x
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
RuntimeError: CUDA error: CUBLAS_STATUS_NOT_SUPPORTED when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)`
>>> torch.mm(y,x)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
RuntimeError: CUDA error: CUBLAS_STATUS_NOT_SUPPORTED when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)`
>>> torch.mm(x,y)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
RuntimeError: CUDA error: CUBLAS_STATUS_NOT_SUPPORTED when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)`
>>>

```

The issues are obviously not caused by the mismatch size. Anyone has any idea? thanks!

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [December 3, 2022, 12:48am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/19 "2022-12-03T00:48:53Z")

</div>

Could you post the output of `python -m torch.utils.collect_env`, please, as I cannot reproduce the error in `1.13.0+cu117` on a 3090.

---

<div class="post-metadata">

**Author:** ![pt.megamozg80](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/pt.megamozg80/32/55632_2.png) [@pt.megamozg80](https://discuss.pytorch.org/u/pt.megamozg80)\
**Post date:** [December 12, 2022, 10:21am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/20 "2022-12-12T10:21:04Z")

</div>

Hello! I had the same issue as [wtliao](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/18)

This os my post of `python -m torch.utils.collect_env`:

PyTorch version: 1.13.0+cu117  
Is debug build: False  
CUDA used to build PyTorch: 11.7  
ROCM used to build PyTorch: N/A

OS: CentOS Linux 7 (Core) (x86\_64)  
GCC version: (GCC) 4.8.5 20150623 (Red Hat 4.8.5-44)  
Clang version: Could not collect  
CMake version: Could not collect  
Libc version: glibc-2.17

Python version: 3.8.11 (default, Sep 1 2021, 12:33:46) [GCC 9.3.1 20200408 (Red Hat 9.3.1-2)] (64-bit runtime)  
Python platform: Linux-3.10.0-1160.42.2.el7.x86\_64-x86\_64-with-glibc2.2.5  
Is CUDA available: True  
CUDA runtime version: 11.4.120  
CUDA\_MODULE\_LOADING set to: LAZY  
GPU models and configuration: GPU 0: Tesla T4  
Nvidia driver version: 470.57.02  
cuDNN version: Probably one of the following:  
/usr/lib64/libcudnn.so.8.2.4  
/usr/lib64/libcudnn\_adv\_infer.so.8.2.4  
/usr/lib64/libcudnn\_adv\_train.so.8.2.4  
/usr/lib64/libcudnn\_cnn\_infer.so.8.2.4  
/usr/lib64/libcudnn\_cnn\_train.so.8.2.4  
/usr/lib64/libcudnn\_ops\_infer.so.8.2.4  
/usr/lib64/libcudnn\_ops\_train.so.8.2.4  
HIP runtime version: N/A  
MIOpen runtime version: N/A  
Is XNNPACK available: True

Versions of relevant libraries:  
[pip3] numpy==1.23.4  
[pip3] torch==1.13.0  
[pip3] torchaudio==0.13.0  
[pip3] torchcam==0.3.2  
[pip3] torchvision==0.14.0  
[conda] Could not collect

Also this error occurs while running

```auto
import torch.nn.functional as F
import torch
a = torch.rand((1, 2, 3)).to('cuda')
b = torch.rand((1, 3, 24, 94)).to('cuda')
grid = F.affine_grid(a, b.size())

```

File ~/.venv/default/lib64/python3.8/site-packages/torch/nn/functional.py:4332, in affine\_grid(theta, size, align\_corners)  
4329 elif min(size) \<= 0:  
4330 raise ValueError(“Expected non-zero, positive output size. Got {}”.format(size))  
 → 4332 return torch.affine\_grid\_generator(theta, size, align\_corners)

RuntimeError: CUDA error: CUBLAS\_STATUS\_INVALID\_VALUE when calling `cublasSgemmStridedBatched( handle, opa, opb, m, n, k, &alpha, a, lda, stridea, b, ldb, strideb, &beta, c, ldc, stridec, num_batches)`

---

<div class="post-metadata">

**Author:** ![ptrblck](https://discuss.pytorch.org/user_avatar/discuss.pytorch.org/ptrblck/32/1823_2.png) [@ptrblck](https://discuss.pytorch.org/u/ptrblck)\
**Post date:** [December 13, 2022, 8:04am UTC](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544/21 "2022-12-13T08:04:05Z")

</div>

I cannot reproduce the issue on a T4 with `torch==1.13.0+cu117`:

```python
import torch
torch.cuda.get_device_name(0)
# 'Tesla T4'
torch. __version__
# '1.13.0+cu117'

import torch.nn.functional as F
import torch
a = torch.rand((1, 2, 3)).to('cuda')
b = torch.rand((1, 3, 24, 94)).to('cuda')
grid = F.affine_grid(a, b.size())

print(grid)
tensor([[[[0.4507, 0.2959],
          [0.4582, 0.3064],
          [0.4656, 0.3169],
          ...,
          [1.1288, 1.2493],
          [1.1363, 1.2597],
          [1.1437, 1.2702]],

         [[0.4635, 0.3081],
          [0.4710, 0.3186],
          [0.4784, 0.3291],
          ...,
          [1.1417, 1.2615],
          [1.1491, 1.2719],
          [1.1566, 1.2824]],

         [[0.4763, 0.3203],
          [0.4838, 0.3308],
          [0.4913, 0.3413],
          ...,
          [1.1545, 1.2736],
          [1.1619, 1.2841],
          [1.1694, 1.2946]],
...

```

[Next page](https://discuss.pytorch.org/t/runtimeerror-cuda-error-cublas-status-invalid-value-when-calling-cublassgemm-handle-opa-opb-m-n-k-alpha-a-lda-b-ldb-beta-c-ldc/124544.md?page=2)
