Repository navigation
Add paper: Evaluating Code Slop in Long-Horizon Coding Agents - #368
Open
shubhamrgandhi wants to merge 1 commit into
Open
shubhamrgandhi wants to merge 1 commit into
shubhamrgandhi wants to merge 1 commit into
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds our paper to the Surveys {{SECTION}} Empirical Studies section.
Evaluating Code Slop in Long-Horizon Coding Agents
Shubham Gandhi, Shyam Agarwal, Nachiket Kotalwar, Atharva Naik (Carnegie Mellon University)
COLM 2026 Workshop on Agent Behavior. Paper: https://openreview.net/forum?id=VLgFkLRUfV
Coding-agent evaluations usually check whether a task was solved, not what code was left behind for humans to review and maintain. We call unnecessary code volume, control-flow complexity, and maintainability debt beyond task requirements "code slop" and measure it as trajectory-induced deltas in three static-analysis metrics. Across 2,268 runs (42 SWE-EVO tasks, three coding models, three seeds, six cleanup settings), aggregate functional utility stays within a narrow band (70–76%) while code-volume deltas range from −14.3 to −79.3 across cleanup policies. A single final cleanup pass yields little slop reduction, periodic in-trajectory cleanup helps but is costly, and a monitor-triggered policy gives the largest reduction on all three slop metrics at about 25% lower cost than the strongest periodic schedule.
I am one of the authors. The entry follows the format of the surrounding entries; happy to adjust placement or wording.
Placed at the top of Surveys & Empirical Studies as the newest entry and tagged Empirical Study. The paper has no arXiv version, so the Paper badge links to OpenReview and drops the arXiv logo.