This new benchmark version introduces tests for instruction following and complex project management to align with current software development workflows. All tested models received lower marks due to increased difficulty and longer horizon requirements that demand more consistency over time.

Cursor designed the update as a living instance to keep pace with the rapid evolution of AI capabilities. The version 4.0 release specifically targets how models handle challenging projects over extended periods rather than short, isolated tasks.

Sign in to suggest edits

Key sources

  1. SOURCE@stringchaos“It includes new tasks for how well models follow instructions, work on challenging projects over time, and is more difficult than before (so all models score lower)”x.com
Markdown