Sparse matrix-vector multiplication on GPGPU clusters: A new storage format and a scalable implementation | Read Paper on Bytez