@ptrblck thank you for your response and apologies for the delay in the reply. I have added a couple of code snippets that are mentioned in the error log. Please let me know if there is anything missing below that might be helpful to identify the cause of the issue.
The error pops up when I run the training script:
# -*- coding: utf-8 -*-
# Copyright (c) Facebook, Inc. and its affiliates.
import logging
import numpy as np
import time
import weakref
from typing import List, Mapping, Optional
import torch
from torch.nn.parallel import DataParallel, DistributedDataParallel
import detectron2.utils.comm as comm
from detectron2.utils.events import EventStorage, get_event_storage
from detectron2.utils.logger import _log_api_usage
__all__ = ["HookBase", "TrainerBase", "SimpleTrainer", "AMPTrainer"]
class HookBase:
Base class for hooks that can be registered with :class:`TrainerBase`.
Each hook can implement 4 methods. The way they are called is demonstrated
in the following snippet:
for iter in range(start_iter, max_iter):
iter += 1
1. In the hook method, users can access ``self.trainer`` to access more
properties about the context (e.g., model, current iteration, or config
if using :class:`DefaultTrainer`).
2. A hook that does something in :meth:`before_step` can often be
implemented equivalently in :meth:`after_step`.
If the hook takes non-trivial time, it is strongly recommended to
implement the hook in :meth:`after_step` instead of :meth:`before_step`.
The convention is that :meth:`before_step` should only take negligible time.
Following this convention will allow hooks that do care about the difference
between :meth:`before_step` and :meth:`after_step` (e.g., timer) to
function properly.
trainer: "TrainerBase" = None
A weak reference to the trainer object. Set by the trainer when the hook is registered.
def before_train(self):
Called before the first iteration.
def after_train(self):
Called after the last iteration.
def before_step(self):
Called before each iteration.
def after_step(self):
Called after each iteration.
def state_dict(self):
Hooks are stateless by default, but can be made checkpointable by
implementing `state_dict` and `load_state_dict`.
return {}
class TrainerBase:
Base class for iterative trainer with hooks.
The only assumption we made here is: the training runs in a loop.
A subclass can implement what the loop is.
We made no assumptions about the existence of dataloader, optimizer, model, etc.
iter(int): the current iteration.
start_iter(int): The iteration to start with.
By convention the minimum possible value is 0.
max_iter(int): The iteration to end training.
storage(EventStorage): An EventStorage that's opened during the course of training.
def __init__(self) -> None:
self._hooks: List[HookBase] = []
self.iter: int = 0
self.start_iter: int = 0
self.max_iter: int
self.storage: EventStorage
_log_api_usage("trainer." + self.__class__.__name__)
def register_hooks(self, hooks: List[Optional[HookBase]]) -> None:
Register hooks to the trainer. The hooks are executed in the order
they are registered.
hooks (list[Optional[HookBase]]): list of hooks
hooks = [h for h in hooks if h is not None]
for h in hooks:
assert isinstance(h, HookBase)
# To avoid circular reference, hooks and trainer cannot own each other.
# This normally does not matter, but will cause memory leak if the
# involved objects contain __del__:
# See http://engineering.hearsaysocial.com/2013/06/16/circular-references-in-python/
h.trainer = weakref.proxy(self)
def train(self, start_iter: int, max_iter: int):
start_iter, max_iter (int): See docs above
logger = logging.getLogger(__name__)
logger.info("Starting training from iteration {}".format(start_iter))
self.iter = self.start_iter = start_iter
self.max_iter = max_iter
with EventStorage(start_iter) as self.storage:
for self.iter in range(start_iter, max_iter):
# self.iter == max_iter can be used by `after_train` to
# tell whether the training successfully finished or failed
# due to exceptions.
self.iter += 1
except Exception:
logger.exception("Exception during training:")
def before_train(self):
for h in self._hooks:
def after_train(self):
self.storage.iter = self.iter
for h in self._hooks:
def before_step(self):
# Maintain the invariant that storage.iter == trainer.iter
# for the entire execution of each step
self.storage.iter = self.iter
for h in self._hooks:
def after_step(self):
for h in self._hooks:
def run_step(self):
raise NotImplementedError
def state_dict(self):
ret = {"iteration": self.iter}
hooks_state = {}
for h in self._hooks:
sd = h.state_dict()
if sd:
name = type(h).__qualname__
if name in hooks_state:
# TODO handle repetitive stateful hooks
hooks_state[name] = sd
if hooks_state:
ret["hooks"] = hooks_state
return ret
def load_state_dict(self, state_dict):
logger = logging.getLogger(__name__)
self.iter = state_dict["iteration"]
for key, value in state_dict.get("hooks", {}).items():
for h in self._hooks:
name = type(h).__qualname__
except AttributeError:
if name == key:
logger.warning(f"Cannot find the hook '{key}', its state_dict is ignored.")
class SimpleTrainer(TrainerBase):
A simple trainer for the most common type of task:
single-cost single-optimizer single-data-source iterative optimization,
optionally using data-parallelism.
It assumes that every step, you:
1. Compute the loss with a data from the data_loader.
2. Compute the gradients with the above loss.
3. Update the model with the optimizer.
All other tasks during training (checkpointing, logging, evaluation, LR schedule)
are maintained by hooks, which can be registered by :meth:`TrainerBase.register_hooks`.
If you want to do anything fancier than this,
either subclass TrainerBase and implement your own `run_step`,
or write your own training loop.
def __init__(self, model, data_loader, optimizer):
model: a torch Module. Takes a data from data_loader and returns a
dict of losses.
data_loader: an iterable. Contains data to be used to call model.
optimizer: a torch optimizer.
We set the model to training mode in the trainer.
However it's valid to train a model that's in eval mode.
If you want your model (or a submodule of it) to behave
like evaluation during training, you can overwrite its train() method.
self.model = model
self.data_loader = data_loader
self._data_loader_iter = iter(data_loader)
self.optimizer = optimizer
def run_step(self):
Implement the standard training logic described above.
assert self.model.training, "[SimpleTrainer] model was changed to eval mode!"
start = time.perf_counter()
If you want to do something with the data, you can wrap the dataloader.
data = next(self._data_loader_iter)
data_time = time.perf_counter() - start
If you want to do something with the losses, you can wrap the model.
loss_dict = self.model(data)
if isinstance(loss_dict, torch.Tensor):
losses = loss_dict
loss_dict = {"total_loss": loss_dict}
losses = sum(loss_dict.values())
If you need to accumulate gradients or do something similar, you can
wrap the optimizer with your custom `zero_grad()` method.
self._write_metrics(loss_dict, data_time)
If you need gradient clipping/scaling or other processing, you can
wrap the optimizer with your custom `step()` method. But it is
suboptimal as explained in https://arxiv.org/abs/2006.15704 Sec 3.2.4
def _write_metrics(
loss_dict: Mapping[str, torch.Tensor],
data_time: float,
prefix: str = "",
) -> None:
SimpleTrainer.write_metrics(loss_dict, data_time, prefix)
def write_metrics(
loss_dict: Mapping[str, torch.Tensor],
data_time: float,
prefix: str = "",
) -> None:
loss_dict (dict): dict of scalar losses
data_time (float): time taken by the dataloader iteration
prefix (str): prefix for logging keys
metrics_dict = {k: v.detach().cpu().item() for k, v in loss_dict.items()}
metrics_dict["data_time"] = data_time
# Gather metrics among all workers for logging
# This assumes we do DDP-style training, which is currently the only
# supported method in detectron2.
all_metrics_dict = comm.gather(metrics_dict)
if comm.is_main_process():
storage = get_event_storage()
# data_time among workers can have high variance. The actual latency
# caused by data_time is the maximum among workers.
data_time = np.max([x.pop("data_time") for x in all_metrics_dict])
storage.put_scalar("data_time", data_time)
# average the rest metrics
metrics_dict = {
k: np.mean([x[k] for x in all_metrics_dict]) for k in all_metrics_dict[0].keys()
total_losses_reduced = sum(metrics_dict.values())
if not np.isfinite(total_losses_reduced):
raise FloatingPointError(
f"Loss became infinite or NaN at iteration={storage.iter}!\n"
f"loss_dict = {metrics_dict}"
storage.put_scalar("{}total_loss".format(prefix), total_losses_reduced)
if len(metrics_dict) > 1:
def state_dict(self):
ret = super().state_dict()
ret["optimizer"] = self.optimizer.state_dict()
return ret
def load_state_dict(self, state_dict):
class AMPTrainer(SimpleTrainer):
Like :class:`SimpleTrainer`, but uses PyTorch's native automatic mixed precision
in the training loop.
def __init__(self, model, data_loader, optimizer, grad_scaler=None):
model, data_loader, optimizer: same as in :class:`SimpleTrainer`.
grad_scaler: torch GradScaler to automatically scale gradients.
unsupported = "AMPTrainer does not support single-process multi-device training!"
if isinstance(model, DistributedDataParallel):
assert not (model.device_ids and len(model.device_ids) > 1), unsupported
assert not isinstance(model, DataParallel), unsupported
super().__init__(model, data_loader, optimizer)
if grad_scaler is None:
from torch.cuda.amp import GradScaler
grad_scaler = GradScaler()
self.grad_scaler = grad_scaler
def run_step(self):
Implement the AMP training logic.
assert self.model.training, "[AMPTrainer] model was changed to eval mode!"
assert torch.cuda.is_available(), "[AMPTrainer] CUDA is required for AMP training!"
from torch.cuda.amp import autocast
start = time.perf_counter()
data = next(self._data_loader_iter)
data_time = time.perf_counter() - start
with autocast():
loss_dict = self.model(data)
if isinstance(loss_dict, torch.Tensor):
losses = loss_dict
loss_dict = {"total_loss": loss_dict}
losses = sum(loss_dict.values())
self._write_metrics(loss_dict, data_time)
def state_dict(self):
ret = super().state_dict()
ret["grad_scaler"] = self.grad_scaler.state_dict()
return ret
def load_state_dict(self, state_dict):
And this is the file where we are mapping the IDs to the images and the error pops up:
# Copyright (c) Facebook, Inc. and its affiliates. All Rights Reserved
import copy
import logging
import numpy as np
import torch
from fvcore.common.file_io import PathManager
from PIL import Image
from detectron2.data import detection_utils as utils
from detectron2.data import transforms as T
import pandas as pd
from detectron2.data.catalog import MetadataCatalog
This file contains the default mapping that's applied to "dataset dicts".
__all__ = ["DatasetMapperWithSupport"]
class DatasetMapperWithSupport:
A callable which takes a dataset dict in Detectron2 Dataset format,
and map it into a format used by the model.
This is the default callable to be used to map your dataset dict into training data.
You may need to follow it to implement your own one for customized logic,
such as a different way to read or transform images.
See :doc:`/tutorials/data_loading` for details.
The callable currently does the following:
1. Read the image from "file_name"
2. Applies cropping/geometric transforms to the image and annotations
3. Prepare data and annotations to Tensor and :class:`Instances`
def __init__(self, cfg, is_train=True):
if cfg.INPUT.CROP.ENABLED and is_train:
self.crop_gen = T.RandomCrop(cfg.INPUT.CROP.TYPE, cfg.INPUT.CROP.SIZE)
logging.getLogger(__name__).info("CropGen used in training: " + str(self.crop_gen))
self.crop_gen = None
self.tfm_gens = utils.build_transform_gen(cfg, is_train)
# fmt: off
self.img_format = cfg.INPUT.FORMAT
self.mask_on = cfg.MODEL.MASK_ON
self.mask_format = cfg.INPUT.MASK_FORMAT
self.keypoint_on = cfg.MODEL.KEYPOINT_ON
self.load_proposals = cfg.MODEL.LOAD_PROPOSALS
self.few_shot = cfg.INPUT.FS.FEW_SHOT
self.support_way = cfg.INPUT.FS.SUPPORT_WAY
self.support_shot = cfg.INPUT.FS.SUPPORT_SHOT
# fmt: on
if self.keypoint_on and is_train:
# Flip only makes sense in training
self.keypoint_hflip_indices = utils.create_keypoint_hflip_indices(cfg.DATASETS.TRAIN)
self.keypoint_hflip_indices = None
if self.load_proposals:
self.proposal_min_box_size = cfg.MODEL.PROPOSAL_GENERATOR.MIN_SIZE
self.proposal_topk = (
if is_train
self.is_train = is_train
if self.is_train:
# support_df
self.support_on = True
if self.few_shot:
self.support_df = pd.read_pickle("./datasets/voc/10_shot_support_df.pkl")
self.support_df = pd.read_pickle("./datasets/voc/train_support_df.pkl")
metadata = MetadataCatalog.get('voc_2012_training')
# unmap the category mapping ids for COCO
reverse_id_mapper = lambda dataset_id: metadata.thing_dataset_id_to_contiguous_id[dataset_id] # noqa
self.support_df['category_id'] = self.support_df['category_id'].map(reverse_id_mapper)
def __call__(self, dataset_dict):
dataset_dict (dict): Metadata of one image, in Detectron2 Dataset format.
dict: a format that builtin models in detectron2 accept
dataset_dict = copy.deepcopy(dataset_dict) # it will be modified by code below
# USER: Write your own image loading if it's not from a file
image = utils.read_image(dataset_dict["file_name"], format=self.img_format)
utils.check_image_size(dataset_dict, image)
if self.is_train:
# support
if self.support_on:
if "annotations" in dataset_dict:
# USER: Modify this if you want to keep them for some reason.
for anno in dataset_dict["annotations"]:
if not self.mask_on:
anno.pop("segmentation", None)
if not self.keypoint_on:
anno.pop("keypoints", None)
support_images, support_bboxes, support_cls = self.generate_support(dataset_dict)
dataset_dict['support_images'] = torch.as_tensor(np.ascontiguousarray(support_images))
dataset_dict['support_bboxes'] = support_bboxes
dataset_dict['support_cls'] = support_cls
if "annotations" not in dataset_dict:
image, transforms = T.apply_transform_gens(
([self.crop_gen] if self.crop_gen else []) + self.tfm_gens, image
# Crop around an instance if there are instances in the image.
# USER: Remove if you don't use cropping
if self.crop_gen:
crop_tfm = utils.gen_crop_transform_with_instance(
image = crop_tfm.apply_image(image)
image, transforms = T.apply_transform_gens(self.tfm_gens, image)
if self.crop_gen:
transforms = crop_tfm + transforms
image_shape = image.shape[:2] # h, w
# Pytorch's dataloader is efficient on torch.Tensor due to shared-memory,
# but not efficient on large generic data structures due to the use of pickle & mp.Queue.
# Therefore it's important to use torch.Tensor.
dataset_dict["image"] = torch.as_tensor(np.ascontiguousarray(image.transpose(2, 0, 1)))
# USER: Remove if you don't use pre-computed proposals.
# Most users would not need this feature.
if self.load_proposals:
if not self.is_train:
# USER: Modify this if you want to keep them for some reason.
dataset_dict.pop("annotations", None)
dataset_dict.pop("sem_seg_file_name", None)
return dataset_dict
if "annotations" in dataset_dict:
# USER: Modify this if you want to keep them for some reason.
for anno in dataset_dict["annotations"]:
if not self.mask_on:
anno.pop("segmentation", None)
if not self.keypoint_on:
anno.pop("keypoints", None)
# USER: Implement additional transformations if you have other types of data
annos = [
obj, transforms, image_shape, keypoint_hflip_indices=self.keypoint_hflip_indices
for obj in dataset_dict.pop("annotations")
if obj.get("iscrowd", 0) == 0
instances = utils.annotations_to_instances(
annos, image_shape, mask_format=self.mask_format
# Create a tight bounding box from masks, useful when image is cropped
if self.crop_gen and instances.has("gt_masks"):
instances.gt_boxes = instances.gt_masks.get_bounding_boxes()
dataset_dict["instances"] = utils.filter_empty_instances(instances)
# USER: Remove if you don't do semantic/panoptic segmentation.
if "sem_seg_file_name" in dataset_dict:
with PathManager.open(dataset_dict.pop("sem_seg_file_name"), "rb") as f:
sem_seg_gt = Image.open(f)
sem_seg_gt = np.asarray(sem_seg_gt, dtype="uint8")
sem_seg_gt = transforms.apply_segmentation(sem_seg_gt)
sem_seg_gt = torch.as_tensor(sem_seg_gt.astype("long"))
dataset_dict["sem_seg"] = sem_seg_gt
return dataset_dict
def generate_support(self, dataset_dict):
support_way = self.support_way #2
support_shot = self.support_shot #5
id = dataset_dict['annotations'][0]['id']
query_cls = self.support_df.loc[self.support_df['id']==id, 'category_id'].tolist()[0] # they share the same category_id and image_id
query_img = self.support_df.loc[self.support_df['id']==id, 'image_id'].tolist()[0]
all_cls = self.support_df.loc[self.support_df['image_id']==query_img, 'category_id'].tolist()
# Crop support data and get new support box in the support data
support_data_all = np.zeros((support_way * support_shot, 3, 320, 320), dtype = np.float32)
support_box_all = np.zeros((support_way * support_shot, 4), dtype = np.float32)
used_image_id = [query_img]
used_id_ls = []
for item in dataset_dict['annotations']:
#used_category_id = [query_cls]
used_category_id = list(set(all_cls))
support_category_id = []
mixup_i = 0
for shot in range(support_shot):
# Support image and box
support_id = self.support_df.loc[(self.support_df['category_id'] == query_cls) & (~self.support_df['image_id'].isin(used_image_id)) & (~self.support_df['id'].isin(used_id_ls)), 'id'].sample(random_state=id).tolist()[0]
support_cls = self.support_df.loc[self.support_df['id'] == support_id, 'category_id'].tolist()[0]
support_img = self.support_df.loc[self.support_df['id'] == support_id, 'image_id'].tolist()[0]
support_db = self.support_df.loc[self.support_df['id'] == support_id, :]
assert support_db['id'].values[0] == support_id
support_data = utils.read_image('./datasets/voc/' + support_db["file_path"].tolist()[0], format=self.img_format)
support_data = torch.as_tensor(np.ascontiguousarray(support_data.transpose(2, 0, 1)))
support_box = support_db['support_box'].tolist()[0]
support_data_all[mixup_i] = support_data
support_box_all[mixup_i] = support_box
support_category_id.append(0) #support_cls)
mixup_i += 1
if support_way == 1:
for way in range(support_way-1):
other_cls = self.support_df.loc[(~self.support_df['category_id'].isin(used_category_id)), 'category_id'].drop_duplicates().sample(random_state=id).tolist()[0]
for shot in range(support_shot):
# Support image and box
support_id = self.support_df.loc[(self.support_df['category_id'] == other_cls) & (~self.support_df['image_id'].isin(used_image_id)) & (~self.support_df['id'].isin(used_id_ls)), 'id'].sample(random_state=id).tolist()[0]
support_cls = self.support_df.loc[self.support_df['id'] == support_id, 'category_id'].tolist()[0]
support_img = self.support_df.loc[self.support_df['id'] == support_id, 'image_id'].tolist()[0]
support_db = self.support_df.loc[self.support_df['id'] == support_id, :]
assert support_db['id'].values[0] == support_id
support_data = utils.read_image('./datasets/voc/' + support_db["file_path"].tolist()[0], format=self.img_format)
support_data = torch.as_tensor(np.ascontiguousarray(support_data.transpose(2, 0, 1)))
support_box = support_db['support_box'].tolist()[0]
support_data_all[mixup_i] = support_data
support_box_all[mixup_i] = support_box
support_category_id.append(1) #support_cls)
mixup_i += 1
return support_data_all, support_box_all, support_category_id
Since the files are interdependent on multiple files, I am not sure if these two main code blocks would be enough to debug. Kindly let me know if there’s any other code snippet from the library that could help in debugging the issue. Thank you!