Initialization of first hidden state in LSTM and truncated BPTT

  1. Yes, zero initial hiddenstate is standard so much so that it is the default in nn.LSTM if you don’t pass in a hidden state (rather than, e.g. throwing an error). Random initialization could also be used if zeros don’t work. Two basic ideas here:

    • If your hidden state evolution is “ergodic”, the state will move closer to some “steady distribution” anyways, so it doesn’t matter as much.
    • You want the initial hidden state handling to be somewhat consistent between training and inference.
    • The fancy Bayesian way would be to sample from said steady state, but deep learning is too wild to resort to fancy when it isn’t necessary.
  2. For BPTT (aka “feeding in a long sequence bit by bit”), you could keep the last hidden state and use that (detached) as the new initial hidden state if you think that state should be carried between batches. Language model training on Wikipedia (as a common example) will do things like that.

  3. In theory this is really true. In practice you would run out of memory instead. I can’t speak about other tutorials, but those that I have seen do the detaching (or don’t keep state from previous batches) and it would seem necessary to do so.

Best regards

Thomas

P.S.: For the forum: lines with triple backticks ```python at the beginning of your codeand ``` at the end will make it look nice.