Aider polyglot benchmark
225 of the hardest Exercism exercises across C++, Go, Java, JavaScript, Python and Rust, scored on correctness after two attempts, with cost per run reported.
Aider's benchmark measures a model's ability to edit code across languages in a realistic tool loop: the model gets the exercise, writes a solution, sees the test failures, and gets one retry. Reporting cost alongside score makes it useful for procurement: GPT-5 (high) 88% at $29, DeepSeek-V3.2 74% at $1.30.