Settings

Theme

Agent Harness Evolution Shapes Coding Agent Quality

arxiv.org

2 points by wek · 1 comment

Reader

1 thread
wekOP

From their abstract: "Our findings reveal that despite continuous development activity and growing codebase complexity of the agent harnesses, there is no statistically significant improvement in SWE-bench benchmark score (i.e., resolve rates of bugs) across releases for a given fixed LLM version. Worse, later agent harness versions consume nearly double the computational tokens and tool calls without corresponding quality gains. We explain this paradox from two angles: at the project level, we identify the development patterns (e.g., feature additions, fix-heavy releases, scattered small changes) correlating with quality fluctuations, while at the architecture level, we localize regressions to specific high-risk architectural components. Our findings call for a new practice of quality assurance in the development of agent harnesses."

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection