In Dynamo+AOTAutograd, why run Faketensor through the code multiple times?

I posted a series of questions on the forum today. I’ve listed them in hope that putting them together would better shed light on what I do know and don’t know.

When tracing a FX Graph through Dynamo & Lowering it to ATen IR through AOTAutograd, for the forward pass, it looks like the FakeTensor inputs are run through the whole code three times:

  1. Within Dynamo, when interpreting the Bytecode CALL, to get the FakeTensor output of the result of this CALL, to push into the stack as input to the next functions.: wrap_fx_proxy → … → get_fake_value → .. run_node
  2. Within AOTAutograd, within create_aot_state, to find out about mutations, etc: aot_autograd.py:L575
  3. Within AOTAutograd, within aot_stage1_graph_capture, to 1) Obtain the ATen IR level graph and 2) To record operations in the Autograd tape, to so that the backward pass can be traced through: aot_dispatch_autograd_graph → … make_fx → …Tracer.trace

Question: Why do we have to run FakeTensor through the function multiple times?
This DevDiscussion seems to suggest that AOTAutograd is needed to reach into C++ world where the Automatic Differentiation engine sits (quote below).

However, even from Dynamo, we do go into C++ land, and the Dispatcher is able to throw things back to Python land with __torch_dispatch__. Couldn’t we just throw in some mode/contexts and, when Dynamo is dispatching the functions, make Autograd record the tape? (i.e. Dispatch the Autograd feature)

(Quote from above link)

Training adds challenges because the PyTorch Automatic Differentiation engine sits below the PyTorch dispatcher in C++. Therefore, the operators running in the backward pass are not directly visible to TorchDynamo at the Python level.