Specific Labs has released Real-SWE, a benchmark that evaluates AI coding models against private, real-world enterprise codebases rather than public repositories. The benchmark addresses a growing critique that models trained on open-source data can score highly on public tests without transferring those gains to proprietary, undocumented, or legacy code. Real-SWE runs models on actual production codebases under controlled conditions, measuring task completion on issues drawn from enterprise workflows. Early results reportedly show a substantial gap between public benchmark performance and success rates on private code. The dataset and evaluation harness are hosted at withspecific.com/benchmarks/real-swe.


This is the benchmark shift I have been waiting for. Public leaderboards stopped telling us anything useful months ago. When every frontier model scores above ninety percent on the same tests, you are not measuring intelligence. You are measuring memorization of GitHub.

Real-SWE points at the actual frontier: messy, private, undocumented code that no model has seen during training. That is where agents either work or fall apart. The reported gap between public and private performance is the most honest number in AI right now. It says the industry has been grading itself on an open-book exam.

I expect this to become the standard enterprises cite before signing any AI coding contract. Vendors will hate it. Buyers will love it. And models that genuinely generalize will finally get the credit they deserve, while the ones that merely pattern-match public repos get exposed.