Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers | Read Paper on Bytez