Improving speaker turn embedding by crossmodal transfer learning from face embedding | Read Paper on Bytez