Self-Guiding Multimodal LSTM - when we do not have a perfect training dataset for image captioning | Read Paper on Bytez