The Fault in our Stars: Quality Assessment of Code Generation Benchmarks

Published in 24th IEEE International Conference on Source Code Analysis and Manipulation (SCAM 2024), 2024

We conduct the first large-scale study of prompt quality in code generation benchmarks, analyzing 3,566 prompts drawn from 9 widely used benchmarks. We find that these benchmarks are heavily skewed toward Python coding exercises with limited contextual dependencies, and that many prompts contain quality issues such as spelling and grammatical errors and inconsistent documentation styles. Fixing these issues improves Python code generation performance, though gains are smaller for Java. Our analysis also surfaces potential data contamination in GPT-3.5-Turbo and CodeGen-2.5, raising questions about the trustworthiness of current benchmark-based evaluations.

Read the paper on IEEE Xplore

Recommended citation: Mohammed Latif Siddiq, Simantika Dristi, Joy Saha, Joanna C. S. Santos. "The Fault in our Stars: Quality Assessment of Code Generation Benchmarks." SCAM, 2024.
Download Paper