An Older Claude Model Outperformed GPT-5.6. Here’s Why That Matters.
An older Anthropic model outperformed OpenAI’s latest model inside commercial legal AI products. Our third ScaffoldBench suggests the model was not the deciding factor. The scaffold was.
An older Anthropic model outperformed OpenAI’s latest model inside commercial legal AI products. At first glance, this sounds surprising. GPT-5.6 is the newer generation model, and conventional wisdom suggests the latest frontier model should consistently deliver better legal performance.
Our third ScaffoldBench suggests something different. The model wasn’t the deciding factor. The scaffold was.
If the scaffold materially influences legal performance, evaluating raw models alone cannot predict how commercial legal AI systems will perform.
The problem with today’s legal benchmarks
Over the past year we’ve seen an increasing number of legal AI benchmarks comparing frontier models on legal tasks. These benchmarks are incredibly valuable for AI researchers: they help identify capability gaps, guide model development and measure progress over time.
But they answer a different question from the one legal buyers actually have. Corporate legal departments, boutique firms and solo practitioners rarely interact with raw models. They use products like Claude Cowork, GPT Work, Harvey and other legal AI platforms. In practice, legal work flows through a complete AI system.
Most benchmarks measure only the last component. For practitioners, that provides half the picture.
The challenge
Testing this hypothesis turned out to be much harder than expected. An ideal experiment would connect the same model to multiple commercial scaffolds. Commercial products don’t work that way.
Only exposes Claude models.
Only exposes GPT models.
Because the model and the scaffold are tightly coupled, comparing commercial legal AI systems becomes an apples-to-oranges comparison.
Our approach
To overcome this limitation, we introduced a common baseline. We used MikeOSS, an open-source legal scaffold, as the controlled reference point, then measured how much additional performance each commercial scaffold extracted from its underlying model.
The scaffold reversed the leaderboard
On the controlled MikeOSS scaffold, GPT-5.6 was the stronger raw model. Inside their native commercial products, the ranking changed. Switch the stage to see the reversal.
Claude Cowork extracted substantially more performance from Opus 4.8 than GPT Work extracted from GPT-5.6.
A complete reversal of the leaderboard: an older Anthropic model finished ahead of OpenAI’s latest model inside their respective commercial legal products.
How the gap closed, then crossed
The most interesting observation wasn’t the final score. It was how the performance gap evolved. MikeOSS isolated intrinsic capability. Commercial scaffolds already reduced GPT’s advantage. Adding identical Privacy Skills pushed Claude Cowork ahead.
This matters because both products received exactly the same domain-specific legal instructions. The Skill itself cannot explain the difference. The evidence suggests that scaffold architecture — and the model’s ability to execute those instructions — plays a decisive role.
The scaffold didn’t merely improve performance. It changed the ranking between frontier models.
The same pattern in every dimension
Looking at the individual evaluation dimensions reveals the same trend. Evidence extraction was near-identical on the controlled scaffold. Legal soundness started in GPT’s favour. Legal reasoning showed the largest shift of all.
The winning system was also the most expensive
Performance was not the only thing the scaffold changed. Average spend per call diverged sharply once each model ran inside its native commercial product.
The top-scoring configuration cost roughly four times more per call than the GPT Work setup it beat by 3.5 points. Buyers are therefore choosing a point on a quality-cost curve, not a single best model.
What this means
Our first ScaffoldBench demonstrated that scaffolds materially influence legal AI performance. Our second showed that domain-specific Skills explain an important part of that improvement, particularly for legal reasoning. This third benchmark adds another piece: GPT-5.6 is intrinsically the stronger raw model on the controlled MikeOSS scaffold, but Claude Cowork extracts substantially more value from Opus 4.8 than GPT Work extracts from GPT-5.6.
Evaluating foundation models in isolation is no longer sufficient.
Legal professionals don’t purchase models. They purchase complete legal AI systems that combine models with planning, retrieval, verification, orchestration, context engineering and domain-specific workflows. Those architectural decisions can amplify — or suppress — the capabilities of the underlying model.
Which system, at which cost per matter, performs best on our recurring work?
How much of our performance comes from the scaffold we control, not the model?
Do our Skills transfer across scaffolds, or only inside one product?
Can our model execute domain instructions as reliably as it reasons?
The question buyers should ask is changing
An older Anthropic model outperformed OpenAI’s latest model inside their respective commercial legal products — not because it was necessarily the stronger raw model, but because its surrounding legal scaffold extracted more of the model’s capabilities.
“Which model performs best?”
“Which legal AI system delivers the best outcomes for my legal work?”
Because in production, lawyers never use models in isolation. They use legal AI systems.
Legal AI needs benchmarks that evaluate whole systems, not just models.
If you are building legal AI systems, legal-agent workflows, or applied AI evaluations for legal teams, we would be happy to compare notes.
Book a Discovery Call