MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization | Read Paper on Bytez