Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models | Read Paper on Bytez