MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance | Read Paper on Bytez