Enable Transformer layer in CUDA Capture

Is what this simple patch does - please merge:

Remove torch::equal from multi_head_attention_forward by AnFunctionArray · Pull Request #177660 · pytorch/pytorch

In my opinion. Or do review it at least. It bottlenecks my workflow - I need to keep custom version in my tree.