You are trying to call backward on a created computation graph (coming from policy_proba_dists) with already updated trainable parameters. This post describes the issue trying to use stale forward activations in more details.
In your case, the weight of the conv layer was already updated and thus cannot be used in the backward call in the second iteration to compute the gradients.