GitHub - Proximal-Labs/frontier-swe: FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research

1 min read Original article ↗

Skip to content

Navigation Menu

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Sign up

Appearance settings

FrontierSWE

FrontierSWE is an effort to test coding agents on the hardest ultra-long horizon technical challenges. Together with partners from academia and industry, we have collected real-world problems from domains including performance engineering, computational science, and ML research, and evaluated how well frontier models can perform on them.

See the leaderboard and blog for results and analysis. FrontierSWE is also available as a Prime Intellect Environment.

About

FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research

Resources

Readme

Activity

Custom properties

Stars

187 stars

Watchers

2 watching

Forks

18 forks