About bidirectional gru with seq2seq example and some modifications

does this line change to result = Variable(torch.zeros(2, 1, self.hidden_size)) or not?

edit –
also, how do you pass the hidden state from a bidirectional encoder to a decoder? Let’s say that the hidden dimension for the encoder is 256. Then you’d get output of 512, but would not the hidden state be 256 still? You might pass the output (dim of 512) to an encoder that has a hidden dim of 512, but then what do you do about using the hidden state? What is reccomended? Is this not an issue? Can you pass half the output to a smaller encoder or do you pass twice the hidden state to a larger encoder? Am I seeing this wrong?