Since you are not freeing any memory, torch.cuda.max_memory_allocated will still return the currently used memory, as it’s still the peak.
Have a look at this code snippet:
# Check for empty
torch.cuda.max_memory_allocated(0)
# Create one tensor
x = torch.rand(100, 100, 100, device='cuda:0')
# should yield the same value
torch.cuda.memory_allocated(0)
torch.cuda.max_memory_allocated(0)
# Create other tensor
y = torch.rand(100, 100, 100, device='cuda:0')
# Should be same, but higher
torch.cuda.memory_allocated(0)
torch.cuda.max_memory_allocated(0)
# Delete one tensor
del y
# max_memory_allocated should keep it's old value
torch.cuda.memory_allocated(0)
torch.cuda.max_memory_allocated(0)
# Reset to track new peak memory usage
torch.cuda.reset_max_memory_allocated(0)
torch.cuda.memory_allocated(0)
torch.cuda.max_memory_allocated(0)