# Latest

**URL:** https://discuss.pytorch.org/latest.md

[Latest](https://discuss.pytorch.org/latest.md) · [Categories](https://discuss.pytorch.org/categories.md)

---

## [CrashLens: local GPU log triage for OOM, NCCL and Xid errors](https://discuss.pytorch.org/t/crashlens-local-gpu-log-triage-for-oom-nccl-and-xid-errors/225493)

<div class="topic-metadata">

**Author:** [@Vivaan](https://discuss.pytorch.org/u/Vivaan)\
**Replies:** 2\
**Last updated:** [October 10, 2026, 12:41pm UTC](https://discuss.pytorch.org/t/crashlens-local-gpu-log-triage-for-oom-nccl-and-xid-errors/225493 "2026-10-10T12:41:01Z")

</div>

Hi everyone — I’m Vivaan, building ThermaCompute AI to help teams understand GPU failures and wasted compute. Our current work is passive diagnostics: turning logs and telemetry into evidence engineers can investigate, w…

---

## [Multi label classification accuracy calculation](https://discuss.pytorch.org/t/multi-label-classification-accuracy-calculation/225513)

<div class="topic-metadata">

**Author:** [@Darhen](https://discuss.pytorch.org/u/Darhen)\
**Replies:** 2\
**Last updated:** [October 10, 2026, 10:10am UTC](https://discuss.pytorch.org/t/multi-label-classification-accuracy-calculation/225513 "2026-10-10T10:10:03Z")

</div>

Hello, I have a dataset with images and a tensor of multi labels in one hot encoding. During the training I am not able to calculate the accuracy, when I print the predictions for the batch what I am getting is a single …

---

## [Trying to reproduce a small For-Loop in LibTorch JIT](https://discuss.pytorch.org/t/trying-to-reproduce-a-small-for-loop-in-libtorch-jit/225515)

<div class="topic-metadata">

**Author:** [@dbdb](https://discuss.pytorch.org/u/dbdb)\
**Replies:** 1\
**Last updated:** [October 10, 2026, 12:10am UTC](https://discuss.pytorch.org/t/trying-to-reproduce-a-small-for-loop-in-libtorch-jit/225515 "2026-10-10T00:10:24Z")

</div>

Hello there, I was able to successfully represent the following Python code def test\_if\_else(condition: bool, input\_val: Tensor) -\> Tuple\[int, Tensor\]: status = 0 value = input\_val \* 2 …

---

## [KernelLens: Deploying Triton Kernels to C++ Inference Engines (TensorRT & ONNX)](https://discuss.pytorch.org/t/kernellens-deploying-triton-kernels-to-c-inference-engines-tensorrt-onnx/225514)

<div class="topic-metadata">

**Author:** [@Just1truc](https://discuss.pytorch.org/u/Just1truc)\
**Replies:** 0\
**Last updated:** [October 8, 2026, 6:35pm UTC](https://discuss.pytorch.org/t/kernellens-deploying-triton-kernels-to-c-inference-engines-tensorrt-onnx/225514 "2026-10-08T18:35:36Z")

</div>

KernelLens: Deploying PyTorch @triton.jit Kernels to ONNX Runtime & TensorRT C++ Plugins Category: Deployment / C++ / Compilers Hi PyTorch Community! When writing custom deep learning operators in PyTorch using OpenAI …

---

## [Running Windows C++ executables (LibTorch + CUDA) on a Linux GPU server](https://discuss.pytorch.org/t/running-windows-c-executables-libtorch-cuda-on-a-linux-gpu-server/225491)

<div class="topic-metadata">

**Author:** [@AMANI](https://discuss.pytorch.org/u/AMANI)\
**Replies:** 1\
**Last updated:** [October 5, 2026, 1:37pm UTC](https://discuss.pytorch.org/t/running-windows-c-executables-libtorch-cuda-on-a-linux-gpu-server/225491 "2026-10-05T13:37:29Z")

</div>

I have a few Windows C++ executables, some of which use LibTorch with CUDA to run TorchScript models. I’d like to run them on a Linux server with an NVIDIA GPU, with exactly the same results as on Windows. Is recompiling…

---

## [\[ROCm\]\[MI200\] linear\_cross\_entropy recomputed logits use backward-only FP16 alternate GEMM path](https://discuss.pytorch.org/t/rocm-mi200-linear-cross-entropy-recomputed-logits-use-backward-only-fp16-alternate-gemm-path/225415)

<div class="topic-metadata">

**Author:** [@dhuang3-amd](https://discuss.pytorch.org/u/dhuang3-amd)\
**Replies:** 1\
**Last updated:** [October 5, 2026, 1:33pm UTC](https://discuss.pytorch.org/t/rocm-mi200-linear-cross-entropy-recomputed-logits-use-backward-only-fp16-alternate-gemm-path/225415 "2026-10-05T13:33:52Z")

</div>

Description I investigated three fp16 linear\_cross\_entropy test failures on gfx90a / MI200 originally reported in ROCm/TheRock#6993 and was able to reproduce them and trace the numerical behavior further. The failures a…

---

## [Ease development by running computations on remote GPU](https://discuss.pytorch.org/t/ease-development-by-running-computations-on-remote-gpu/121002)

<div class="topic-metadata">

**Author:** [@meandmymind](https://discuss.pytorch.org/u/meandmymind)\
**Replies:** 11\
**Last updated:** [October 4, 2026, 9:08pm UTC](https://discuss.pytorch.org/t/ease-development-by-running-computations-on-remote-gpu/121002 "2026-10-04T21:08:54Z")

</div>

Hello! For development I use local machine with no GPU and have a remote machine with GPU. I like to debug my code via IDE tools but also want to have access to gpu. Using something a-la vscode over ssh is kinda slow …

---

## [Rgpu: a PyTorch device whose tensors live on a remote GPU](https://discuss.pytorch.org/t/rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu/225481)

<div class="topic-metadata">

**Author:** [@yanmi](https://discuss.pytorch.org/u/yanmi)\
**Replies:** 1\
**Last updated:** [October 3, 2026, 6:11am UTC](https://discuss.pytorch.org/t/rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu/225481 "2026-10-03T06:11:26Z")

</div>

Hi all. I’ve been building rgpu, which lets you write device="rgpu" on a laptop while the tensors and kernels live on a GPU host you reach over SSH. pip install rgpu import torch, rgpu x = torch.ones(4, device="rgpu") …

---

## [TraceML: always-on runtime visibility for PyTorch training](https://discuss.pytorch.org/t/traceml-always-on-runtime-visibility-for-pytorch-training/224679)

<div class="topic-metadata">

**Author:** [@abhinavsriva](https://discuss.pytorch.org/u/abhinavsriva)\
**Replies:** 1\
**Last updated:** [October 1, 2026, 9:37am UTC](https://discuss.pytorch.org/t/traceml-always-on-runtime-visibility-for-pytorch-training/224679 "2026-10-01T09:37:00Z")

</div>

Hi everyone, I am building TraceML, an open-source tool for always-on step-level training visibility in PyTorch. It shows where time and memory go inside each training step while training is still running, making it ea…

---

## [\[Project Announcement\] numgraph: A High-Performance C++ Numerical Core for Learnable Graphs (Looking for Maintainers!)](https://discuss.pytorch.org/t/project-announcement-numgraph-a-high-performance-c-numerical-core-for-learnable-graphs-looking-for-maintainers/225463)

<div class="topic-metadata">

**Author:** [@sundaram](https://discuss.pytorch.org/u/sundaram)\
**Replies:** 2\
**Last updated:** [September 30, 2026, 3:51pm UTC](https://discuss.pytorch.org/t/project-announcement-numgraph-a-high-performance-c-numerical-core-for-learnable-graphs-looking-for-maintainers/225463 "2026-09-30T15:51:05Z")

</div>

Hi everyone, I am building numgraph, a lightweight and high-performance numerical core for learnable graphs, designed with a C++ backend and clean Python bindings to handle complex graph-based tensor operations and deep…

---

## [Possible resources leak in void ProcessGroupNCCL::shutdown()](https://discuss.pytorch.org/t/possible-resources-leak-in-void-processgroupnccl-shutdown/225317)

<div class="topic-metadata">

**Author:** [@zk\_z](https://discuss.pytorch.org/u/zk_z)\
**Replies:** 1\
**Last updated:** [September 30, 2026, 9:24am UTC](https://discuss.pytorch.org/t/possible-resources-leak-in-void-processgroupnccl-shutdown/225317 "2026-09-30T09:24:02Z")

</div>

While reviewing the NCCL symmetric-memory cleanup path, I noticed a possible cleanup and lifetime issue in ProcessGroupNCCL::shutdown(). This is currently a code-review question; ProcessGroupNCCL::shutdown() may destroy…

---

## [LTX-2.5 on 16GB Mac Mini M4: My ComfyUI Solution Explained 🔥 How It Actually Works](https://discuss.pytorch.org/t/ltx-2-5-on-16gb-mac-mini-m4-my-comfyui-solution-explained-how-it-actually-works/225462)

<div class="topic-metadata">

**Author:** [@Dev\_Raj](https://discuss.pytorch.org/u/Dev_Raj)\
**Replies:** 0\
**Last updated:** [September 29, 2026, 5:22am UTC](https://discuss.pytorch.org/t/ltx-2-5-on-16gb-mac-mini-m4-my-comfyui-solution-explained-how-it-actually-works/225462 "2026-09-29T05:22:57Z")

</div>

For visual reference and step-by-step node execution showing peak tensor handling and VAE decoding on 16GB UMA (mps⁠, ⁠memory⁠, ⁠diffusion⁠, and ⁠mac⁠) on Mac mini m4 16gb Ram with extended VRAM. export PYTORCH\_MPS\_HIGH…

---

## [DataLoader shared-memory unlink race on Python 3.14 only after midnight (Ubuntu 22.04, PyTorch 2.12.0+cu130)](https://discuss.pytorch.org/t/dataloader-shared-memory-unlink-race-on-python-3-14-only-after-midnight-ubuntu-22-04-pytorch-2-12-0-cu130/225460)

<div class="topic-metadata">

**Author:** [@evenrose](https://discuss.pytorch.org/u/evenrose)\
**Replies:** 0\
**Last updated:** [September 28, 2026, 1:08pm UTC](https://discuss.pytorch.org/t/dataloader-shared-memory-unlink-race-on-python-3-14-only-after-midnight-ubuntu-22-04-pytorch-2-12-0-cu130/225460 "2026-09-28T13:08:04Z")

</div>

Environment: PyTorch 2.12.0+cu130, Python 3.14, Ubuntu 22.04, CUDA 13.0. In a standard training loop using DataLoader with num\_workers \> 0, training occasionally crashes with RuntimeError: could not unlink the shared mem…

---

## [PyTorch MPS Discussions](https://discuss.pytorch.org/t/pytorch-mps-discussions/225459)

<div class="topic-metadata">

**Author:** [@Dev\_Raj](https://discuss.pytorch.org/u/Dev_Raj)\
**Replies:** 0\
**Last updated:** [September 28, 2026, 12:10pm UTC](https://discuss.pytorch.org/t/pytorch-mps-discussions/225459 "2026-09-28T12:10:35Z")

</div>

Quick clarification on the post above: The benchmarks and watermark bypass (⁠PYTORCH\_MPS\_HIGH\_WATERMARK\_RATIO⁠) reflect the current limitations we hit when running heavy diffusion loops under PyTorch’s MPS allocator on 1…

---

## [torch.distributions.kumaraswamy.Kumaraswamy generates samples outside its support](https://discuss.pytorch.org/t/torch-distributions-kumaraswamy-kumaraswamy-generates-samples-outside-its-support/173404)

<div class="topic-metadata">

**Author:** [@kiranchari](https://discuss.pytorch.org/u/kiranchari)\
**Replies:** 12\
**Last updated:** [September 24, 2026, 2:13pm UTC](https://discuss.pytorch.org/t/torch-distributions-kumaraswamy-kumaraswamy-generates-samples-outside-its-support/173404 "2026-09-24T14:13:57Z")

</div>

The Kumaraswamy distribution is defined over the open interval (0,1) but torch.distributions.kumaraswamy.Kumaraswamy can generate samples that equal 1. This causes “nan” outputs when computing the log\_prob of such sample…

---

## [Active-speaker-driven video cropping: what the inference actually cost](https://discuss.pytorch.org/t/active-speaker-driven-video-cropping-what-the-inference-actually-cost/225447)

<div class="topic-metadata">

**Author:** [@ColinGPT9](https://discuss.pytorch.org/u/ColinGPT9)\
**Replies:** 0\
**Last updated:** [September 23, 2026, 12:29am UTC](https://discuss.pytorch.org/t/active-speaker-driven-video-cropping-what-the-inference-actually-cost/225447 "2026-09-23T00:29:51Z")

</div>

Sharing a local video pipeline and, more usefully, the two inference lessons that cost me the most time. The task: crop a 16:9 stream to 9:16 following whoever is speaking. YOLOv8 for people, OpenCV for the crop path,…

---

## [Comparing raw MRI vs FreeSurfer-derived features for multimodal fusion with a small dataset](https://discuss.pytorch.org/t/comparing-raw-mri-vs-freesurfer-derived-features-for-multimodal-fusion-with-a-small-dataset/225429)

<div class="topic-metadata">

**Author:** [@Reynard-the-Fox](https://discuss.pytorch.org/u/Reynard-the-Fox)\
**Replies:** 0\
**Last updated:** [September 17, 2026, 5:57pm UTC](https://discuss.pytorch.org/t/comparing-raw-mri-vs-freesurfer-derived-features-for-multimodal-fusion-with-a-small-dataset/225429 "2026-09-17T17:57:07Z")

</div>

Hi, I’m an MSc student and pretty new to medical imaging/deep learning. I’m working on an Alzheimer’s classification project using the ANMerge dataset. It has around 1,700 participants overall, but only around 450 have M…

---

## [\[RFC-0036\] Zero-GC 64-Byte Cache-Aligned Flat Arena for Speculative Decoding Verification](https://discuss.pytorch.org/t/rfc-0036-zero-gc-64-byte-cache-aligned-flat-arena-for-speculative-decoding-verification/225427)

<div class="topic-metadata">

**Author:** [@Mark\_Gilbert](https://discuss.pytorch.org/u/Mark_Gilbert)\
**Replies:** 0\
**Last updated:** [September 17, 2026, 5:56pm UTC](https://discuss.pytorch.org/t/rfc-0036-zero-gc-64-byte-cache-aligned-flat-arena-for-speculative-decoding-verification/225427 "2026-09-17T17:56:56Z")

</div>

Hi PyTorch community, We recently submitted PyTorch Core RFC-0036 proposing an optional zero-runtime-allocation, 64-byte cache-aligned (alignas(64)) C++20 flat arena memory pattern: :backhand\_index\_pointing\_right: RFC …

---

## [Official wheel downloads return HTTP 403](https://discuss.pytorch.org/t/official-wheel-downloads-return-http-403/225404)

<div class="topic-metadata">

**Author:** [@Jesse\_Yang](https://discuss.pytorch.org/u/Jesse_Yang)\
**Replies:** 3\
**Last updated:** [September 16, 2026, 7:59pm UTC](https://discuss.pytorch.org/t/official-wheel-downloads-return-http-403/225404 "2026-09-16T19:59:14Z")

</div>

I’m downloading—but not installing—Linux CPython 3.12 wheels for a separate machine, using Python 3.14 on Windows 11. These four packages all return HTTP 403 from download-r2.pytorch.org: torch 2.9.1+cu130 torchaudio …

---

## [\[Showcase\] pyforker: A python framework for \[core functionality\]](https://discuss.pytorch.org/t/showcase-pyforker-a-python-framework-for-core-functionality/225412)

<div class="topic-metadata">

**Author:** [@sundaram](https://discuss.pytorch.org/u/sundaram)\
**Replies:** 1\
**Last updated:** [September 14, 2026, 1:04pm UTC](https://discuss.pytorch.org/t/showcase-pyforker-a-python-framework-for-core-functionality/225412 "2026-09-14T13:04:11Z")

</div>

Hi everyone! 👋 I recently published \`pyforker\` on PyPI (https://pypi.org/project/pyforker/) to help streamline \[explain core problem: e.g., running parallel process forks / task isolation\] in Python workflows. ### 🌟 Ke…

---

## [ShrinkAI: knowledge distillation and compression package](https://discuss.pytorch.org/t/shrinkai-knowledge-distillation-and-compression-package/225411)

<div class="topic-metadata">

**Author:** [@elouanzer](https://discuss.pytorch.org/u/elouanzer)\
**Replies:** 0\
**Last updated:** [September 14, 2026, 9:54am UTC](https://discuss.pytorch.org/t/shrinkai-knowledge-distillation-and-compression-package/225411 "2026-09-14T09:54:28Z")

</div>

I just released shrinkai, a python package for reducing neural network size and inference latency (distillation + pruning + quantization), that could be useful for edge/mobile deployment. Built to feel native if you alre…

---

## [Clifra: a differentiable Clifford algebra computation layer for PyTorch](https://discuss.pytorch.org/t/clifra-a-differentiable-clifford-algebra-computation-layer-for-pytorch/225410)

<div class="topic-metadata">

**Author:** [@Concode](https://discuss.pytorch.org/u/Concode)\
**Replies:** 0\
**Last updated:** [September 14, 2026, 1:58am UTC](https://discuss.pytorch.org/t/clifra-a-differentiable-clifford-algebra-computation-layer-for-pytorch/225410 "2026-09-14T01:58:49Z")

</div>

Hi, I’ve been working on clifra, a differentiable Clifford algebra computation layer for PyTorch. The main design goal is to keep Clifford algebra inside the normal PyTorch programming model. Values are ordinary torch.…

---

## [Missing debug version binaries](https://discuss.pytorch.org/t/missing-debug-version-binaries/225403)

<div class="topic-metadata">

**Author:** [@mtsb](https://discuss.pytorch.org/u/mtsb)\
**Replies:** 1\
**Last updated:** [September 11, 2026, 12:56pm UTC](https://discuss.pytorch.org/t/missing-debug-version-binaries/225403 "2026-09-11T12:56:57Z")

</div>

Hello everyone, it seems that the debug version of pre-compiled binaries are missing from the download page? (https://pytorch.org/get-started/locally/). I am upgrading the libtorch version on my project and I am pretty s…

---

## [WebGPU visual PyTorch debugger](https://discuss.pytorch.org/t/webgpu-visual-pytorch-debugger/225211)

<div class="topic-metadata">

**Author:** [@Piotr\_Gryko](https://discuss.pytorch.org/u/Piotr_Gryko)\
**Replies:** 2\
**Last updated:** [September 11, 2026, 3:46am UTC](https://discuss.pytorch.org/t/webgpu-visual-pytorch-debugger/225211 "2026-09-11T03:46:28Z")

</div>

Hi everyone, I have built a 3d real time PyTorch debugger. It works directly from python and runs in your browser. I reverse engineered most of the basic PyTorch modules. All of them are rendered in 3d with real time int…

---

## [Anker Innovations — Embodied AI / Agent Engineer — 2027 graduates, Shenzhen](https://discuss.pytorch.org/t/anker-innovations-embodied-ai-agent-engineer-2027-graduates-shenzhen/225399)

<div class="topic-metadata">

**Author:** [@Tab\_WU](https://discuss.pytorch.org/u/Tab_WU)\
**Replies:** 0\
**Last updated:** [September 10, 2026, 11:19am UTC](https://discuss.pytorch.org/t/anker-innovations-embodied-ai-agent-engineer-2027-graduates-shenzhen/225399 "2026-09-10T11:19:44Z")

</div>

Hi everyone, I work at Anker Innovations and wanted to share a graduate opening that may interest people working on multimodal models and embodied AI. This is a full-time role based in Shenzhen, China, for 2027 graduate…

---

## [PyTorch Symmetric Memory backend NCCL/NVSHMEM vs PyTorch Distributed backend NCCL](https://discuss.pytorch.org/t/pytorch-symmetric-memory-backend-nccl-nvshmem-vs-pytorch-distributed-backend-nccl/225396)

<div class="topic-metadata">

**Author:** [@VasudevaK](https://discuss.pytorch.org/u/VasudevaK)\
**Replies:** 0\
**Last updated:** [September 10, 2026, 2:41am UTC](https://discuss.pytorch.org/t/pytorch-symmetric-memory-backend-nccl-nvshmem-vs-pytorch-distributed-backend-nccl/225396 "2026-09-10T02:41:51Z")

</div>

Hi, We are trying to understand the LLM (\>70b; two variants: dense, MOE) pre-training communication. To better leverage the communication paradigms available in PyTorch, I explored relatively newer additions, i.e., Symm…

---

## [\[Question\] What is the relationship between PyTorch Symmetric Memory and Torchcomms?](https://discuss.pytorch.org/t/question-what-is-the-relationship-between-pytorch-symmetric-memory-and-torchcomms/225394)

<div class="topic-metadata">

**Author:** [@guitang001](https://discuss.pytorch.org/u/guitang001)\
**Replies:** 0\
**Last updated:** [September 9, 2026, 11:56am UTC](https://discuss.pytorch.org/t/question-what-is-the-relationship-between-pytorch-symmetric-memory-and-torchcomms/225394 "2026-09-09T11:56:52Z")

</div>

Are they redundant efforts?As mentioned in the title, I’m exploring the memory management techniques in Torchcomms. However, I noticed that PyTorch already has a Symmetric Memory mechanism, so I’m unclear about the rela…

---

## [\[Maxpooling\] what should the indices be when there're multiple max values?](https://discuss.pytorch.org/t/maxpooling-what-should-the-indices-be-when-therere-multiple-max-values/162390)

<div class="topic-metadata">

**Author:** [@ryanbty](https://discuss.pytorch.org/u/ryanbty)\
**Replies:** 4\
**Last updated:** [September 9, 2026, 7:29am UTC](https://discuss.pytorch.org/t/maxpooling-what-should-the-indices-be-when-therere-multiple-max-values/162390 "2026-09-09T07:29:05Z")

</div>

What to do when there are multiple values in the kernel that equals to the max? For example, for these values when we use torch.argmax : torch.tensor(\[\[1., 1.\], \[1., 1.\]\]) The max is simply 1. What should max in…

---

## [C10::bit\_cast rejects c10::complex\<float\>/c10::complex\<double\>, stricter than std::bit\_cast](https://discuss.pytorch.org/t/c10-bit-cast-rejects-c10-complex-float-c10-complex-double-stricter-than-std-bit-cast/225238)

<div class="topic-metadata">

**Author:** [@dkozel](https://discuss.pytorch.org/u/dkozel)\
**Replies:** 1\
**Last updated:** [September 8, 2026, 4:35pm UTC](https://discuss.pytorch.org/t/c10-bit-cast-rejects-c10-complex-float-c10-complex-double-stricter-than-std-bit-cast/225238 "2026-09-08T16:35:33Z")

</div>

I’m looking into an issue in the handling of complex number types, specifically in the vendored implementation of bit\_cast which does not work with c10::complex or c10::complex torch/headeronly/util/bit\_cast.h:L34-37 s…

---

## [1.43× Faster MoE Training: LoongForge Redefines EP Expert Load Balancing with Topology-Aware Optimal Transport](https://discuss.pytorch.org/t/1-43x-faster-moe-training-loongforge-redefines-ep-expert-load-balancing-with-topology-aware-optimal-transport/225390)

<div class="topic-metadata">

**Author:** [@nullnonenilNULL](https://discuss.pytorch.org/u/nullnonenilNULL)\
**Replies:** 0\
**Last updated:** [September 8, 2026, 3:30pm UTC](https://discuss.pytorch.org/t/1-43x-faster-moe-training-loongforge-redefines-ep-expert-load-balancing-with-topology-aware-optimal-transport/225390 "2026-09-08T15:30:08Z")

</div>

MoE is now the default architecture for frontier models, and the direction of the next generation is clear: more experts, sparser activation. DeepSeek-V4-Pro raised the number of routed experts per layer from 256 in V3 t…

[Next page](https://discuss.pytorch.org/latest.md?page=1)
