RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades | Read Paper on Bytez