b
Discover
Models
Search
About
Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven Approach
6 months ago
·
arXiv