Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs
Wisedocs processed medical claims across ten different code repositories that no one wanted to touch.
Denys Linkov's team spent six months consolidating them into a single repository and benchmarked whether they should have waited for better AI models instead. A refactoring task took three hours with the o3 model and resulted in ten major errors, while newer models like Sonnet 4.6 and Opus 4.8 completed it faster with fewer mistakes. However, when GPT 5.5 was given the entire job at once, it wrote 2,000 lines of code in 10 minutes that were mostly shell without actual functionality—something the model itself acknowledged. Linkov concludes that doing the work directly was worthwhile: code repositories were opened for commits, work that took months now takes a week, and developers from other parts of the company now volunteer to help.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.