News
Gist from Businessinsider

Crosby Launches 'Redline Bench' to Measure AI Performance on Contract Review

Summarized June 17, 2026
Jump to key takeaways

Legal tech startup Crosby has released a public benchmark called Redline Bench, designed to measure how well AI models perform one of lawyers' most time-consuming tasks: reviewing and redlining contracts. The tool is the first of its kind to apply a structured, lawyer-validated rubric to AI-generated contract edits — filling a glaring gap in an industry that, unlike software engineering, has no shared standard for evaluating AI output quality.

The benchmark was built by Crosby Intelligence, an internal unit combining engineers and lawyers. The team — which includes Sharan Ramjee, who previously built fraud-detection transformer models at Stripe, and Ross Weiser, a former Sullivan & Cromwell attorney — partnered with Micro1 to recruit senior lawyers who simulated software deal negotiations. Those lawyers marked the contract changes they deemed most important at each negotiation stage, turning those decisions into weighted scoring criteria. AI models are then given the same contracts, asked to make edits, and scored by a three-judge panel comparing their output to the lawyer-built rubric.

The first round of results puts ChatGPT 5.5 at the top with a 50.5% match rate against lawyer-prioritized edits. Gemini 3.5 Flash scored 45.1%, and Claude Opus 4.8 came in at 44.4%. Anthropic's newer Fable 5 model scored a promising 47.3% in a single test run before Anthropic pulled it from public access — Crosby plans to retest when it becomes available again. The relatively modest scores underscore the difficulty of the task: even the best model only matched lawyers on roughly half the edits they considered critical.

Crosby founder Ryan Daniels, a former in-house lawyer, argues that legal AI faces a challenge coding benchmarks don't: there's no binary pass/fail equivalent. A contract edit can be legally defensible in multiple ways, making 'good' legal work inherently subjective. That ambiguity has frustrated AI labs and legal tech companies alike as they race to automate work that used to pile up on general counsel desks. Redline Bench will be made public so any lab can submit its models, and Crosby plans to release regular comparison reports — an implicit challenge to the labs' tendency to tune their models to their own internal tests.

The broader stakes are significant. Anthropic has been aggressively courting in-house legal teams, and its new legal plugin earlier this year triggered a sell-off in legal tech stocks, signaling how much investor money is riding on AI's ability to displace traditional legal work. Harvey, another well-funded legal AI startup, has released its own benchmarks for case law research and contract review, making the benchmark space itself a competitive arena.

Key Takeaways

  • ChatGPT 5.5 tops first Redline Bench with 50.5% lawyer-match score
  • Gemini 3.5 Flash at 45.1%, Claude Opus 4.8 at 44.4%
  • Anthropic's Fable 5 scored 47.3% before being pulled from access
  • Benchmark built from real lawyer-simulated software deal negotiations
  • Legal AI lacks coding's binary pass/fail — ambiguity is the core problem
  • Labs' self-made benchmarks face trust deficit from model tuning
  • Crosby making Redline Bench public for any AI lab to submit models
Read original article at Businessinsider

Summarize any article in seconds

Gist is a free AI reader for your browser, iPhone, and Android. Get concise summaries and key takeaways from any article or podcast.

Get Gist — Free
⚡ Instant summaries 💬 Chat with articles 🔒 Privacy-first