---
format: "aidr-story-markdown/v1"
id: "9c531c254f4ca75a7e502c3329d5406070effd136db2f724f4d8e7f592dcf1db"
canonical_url: "https://aidr.today/9c531c25?lang=en"
title: "Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-12T20:25:48.000Z"
category: "Research"
topics: ["benchmark","llm","swe-bench","enterprise","open-source","harness"]
source_urls: ["https://withspecific.com/benchmarks/real-swe","https://news.ycombinator.com/item?id=49676820"]
summary: "Real-SWE benchmarks frontier AI models on private production codebases licensed from real companies. Eight model and harness configurations, ten tasks, 640 scored rollouts."
---

# Real\-SWE: Benchmarking AI models on private, real\-world, enterprise codebases

> [Open the canonical story](<https://aidr.today/9c531c25?lang=en>)

**Published:** 2026-09-12T20:25:48.000Z
**Category:** Research
**Topics:** benchmark, llm, swe\-bench, enterprise, open\-source, harness

## Summary

Real\-SWE benchmarks frontier AI models on private production codebases licensed from real companies\. Eight model and harness configurations, ten tasks, 640 scored rollouts\.

## Sources

- [Story source](<https://withspecific.com/benchmarks/real-swe>)
- [Discussion](<https://news.ycombinator.com/item?id=49676820>)

