Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning | Read Paper on Bytez