From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning | Read Paper on Bytez