Research/Benchmarks
Research · Legal AI Evals

An Older Claude Model Outperformed GPT-5.6. Here’s Why That Matters.

An older Anthropic model outperformed OpenAI’s latest model inside commercial legal AI products. Our third ScaffoldBench suggests the model was not the deciding factor. The scaffold was.

LN Labs Research·Published July 2026·7 min read

An older Anthropic model outperformed OpenAI’s latest model inside commercial legal AI products. At first glance, this sounds surprising. GPT-5.6 is the newer generation model, and conventional wisdom suggests the latest frontier model should consistently deliver better legal performance.

Our third ScaffoldBench suggests something different. The model wasn’t the deciding factor. The scaffold was.

If the scaffold materially influences legal performance, evaluating raw models alone cannot predict how commercial legal AI systems will perform.

01 · Context

The problem with today’s legal benchmarks

Over the past year we’ve seen an increasing number of legal AI benchmarks comparing frontier models on legal tasks. These benchmarks are incredibly valuable for AI researchers: they help identify capability gaps, guide model development and measure progress over time.

But they answer a different question from the one legal buyers actually have. Corporate legal departments, boutique firms and solo practitioners rarely interact with raw models. They use products like Claude Cowork, GPT Work, Harvey and other legal AI platforms. In practice, legal work flows through a complete AI system.

The unit of work
Legal taskLegal scaffoldFoundation model

Most benchmarks measure only the last component. For practitioners, that provides half the picture.

02 · Constraint

The challenge

Testing this hypothesis turned out to be much harder than expected. An ideal experiment would connect the same model to multiple commercial scaffolds. Commercial products don’t work that way.

Locked · 01
Claude Cowork

Only exposes Claude models.

Locked · 02
GPT Work

Only exposes GPT models.

Because the model and the scaffold are tightly coupled, comparing commercial legal AI systems becomes an apples-to-oranges comparison.

03 · Method

Our approach

To overcome this limitation, we introduced a common baseline. We used MikeOSS, an open-source legal scaffold, as the controlled reference point, then measured how much additional performance each commercial scaffold extracted from its underlying model.

Dataset
20 tasks
Real-world privacy compliance
Skill
Privacy Skill
Claude for Legal repository
LLM Judge
Google Flash 3.6
Rubric-based scoring
Baseline
MikeOSS
Open-source legal scaffold
Four system configurations
Config 01
MikeOSS + GPT-5.6
Controlled baseline
Config 02
MikeOSS + Opus 4.8
Controlled baseline
Config 03
GPT Work + GPT-5.6
Commercial scaffold
Config 04
Claude Cowork + Opus 4.8
Commercial scaffold
Figure 1 · Methodology
One baseline, two commercial lanes
Controlled → Production
Stage 1
MikeOSS + GPT-5.6
Controlled baseline
Stage 2
GPT Work + GPT-5.6
Commercial scaffold
Stage 3
+ Privacy Skills
Identical instructions
Stage 1
MikeOSS + Opus 4.8
Controlled baseline
Stage 2
Claude Cowork + Opus 4.8
Commercial scaffold
Stage 3
+ Privacy Skills
Identical instructions

MikeOSS served as a controlled baseline, allowing us to compare the relative value added by two commercial legal AI scaffolds.

04 · Results

The scaffold reversed the leaderboard

On the controlled MikeOSS scaffold, GPT-5.6 was the stronger raw model. Inside their native commercial products, the ranking changed. Switch the stage to see the reversal.

Overall leaderboard
Pass rate · 20 privacy tasks
1
Claude Cowork + Privacy SkillsTop+6.8
Commercial scaffold + Skills
81.6
2
GPT Work + Privacy Skills+1.3
Commercial scaffold + Skills
78.1
Identical Privacy Skills added to both products
Gap: Claude Cowork ahead by 3.50

Claude Cowork extracted substantially more performance from Opus 4.8 than GPT Work extracted from GPT-5.6.

A complete reversal of the leaderboard: an older Anthropic model finished ahead of OpenAI’s latest model inside their respective commercial legal products.

05 · Progression

How the gap closed, then crossed

The most interesting observation wasn’t the final score. It was how the performance gap evolved. MikeOSS isolated intrinsic capability. Commercial scaffolds already reduced GPT’s advantage. Adding identical Privacy Skills pushed Claude Cowork ahead.

Performance gap by stage
Scores GPT / Claude · delta = GPT minus Claude
MikeOSS
74.0 / 69.8
+4.25
Commercial scaffold
76.8 / 74.8
+2.00
+ Privacy SkillsReversal
78.1 / 81.6
−3.50

This matters because both products received exactly the same domain-specific legal instructions. The Skill itself cannot explain the difference. The evidence suggests that scaffold architecture — and the model’s ability to execute those instructions — plays a decisive role.

The scaffold didn’t merely improve performance. It changed the ranking between frontier models.

06 · Category analysis

The same pattern in every dimension

Looking at the individual evaluation dimensions reveals the same trend. Evidence extraction was near-identical on the controlled scaffold. Legal soundness started in GPT’s favour. Legal reasoning showed the largest shift of all.

Figure 4 · Category breakdown
Evidence extraction
Pass rate · GPT vs Claude
MikeOSS+1.5
70.8
69.3
Commercial−0.5
72.5
73.0
+ Skills−4.0
74.3
78.3
Near-identical on the baseline. Claude Cowork gained far more from the same Skill.
Legal soundness
Pass rate · GPT vs Claude
MikeOSS+3.75
77.5
73.8
Commercial+2.25
80.3
78.0
+ Skills−2.75
81.5
84.3
GPT-5.6 started stronger on legal correctness. Claude Cowork overtook it after Skills.
Legal reasoningBiggest swing
Pass rate · GPT vs Claude
MikeOSS+7.5
73.8
66.3
Commercial+4.25
77.8
73.5
+ Skills−3.75
78.5
82.3
An 11.3-point swing: Claude Opus finished ahead where it had trailed by 7.5.
GPT-5.6Claude Opus 4.8Claude leads
07 · Cost

The winning system was also the most expensive

Performance was not the only thing the scaffold changed. Average spend per call diverged sharply once each model ran inside its native commercial product.

Cost per call
USD · averaged over 20 privacy tasks
Lower is cheaper
MikeOSS + GPT-5.6$0.49
GPT Work + Privacy Skills$0.60
MikeOSS + Opus 4.8$0.86
Claude Cowork + Privacy Skills$2.42

The top-scoring configuration cost roughly four times more per call than the GPT Work setup it beat by 3.5 points. Buyers are therefore choosing a point on a quality-cost curve, not a single best model.

08 · Implications

What this means

Our first ScaffoldBench demonstrated that scaffolds materially influence legal AI performance. Our second showed that domain-specific Skills explain an important part of that improvement, particularly for legal reasoning. This third benchmark adds another piece: GPT-5.6 is intrinsically the stronger raw model on the controlled MikeOSS scaffold, but Claude Cowork extracts substantially more value from Opus 4.8 than GPT Work extracts from GPT-5.6.

Evaluating foundation models in isolation is no longer sufficient.

Legal professionals don’t purchase models. They purchase complete legal AI systems that combine models with planning, retrieval, verification, orchestration, context engineering and domain-specific workflows. Those architectural decisions can amplify — or suppress — the capabilities of the underlying model.

In-house legal teams

Which system, at which cost per matter, performs best on our recurring work?

BigLaw AI teams

How much of our performance comes from the scaffold we control, not the model?

Legal data vendors

Do our Skills transfer across scaffolds, or only inside one product?

Frontier labs

Can our model execute domain instructions as reliably as it reasons?

09 · Conclusion

The question buyers should ask is changing

An older Anthropic model outperformed OpenAI’s latest model inside their respective commercial legal products — not because it was necessarily the stronger raw model, but because its surrounding legal scaffold extracted more of the model’s capabilities.

Instead of asking

“Which model performs best?”

We should ask

“Which legal AI system delivers the best outcomes for my legal work?”

Because in production, lawyers never use models in isolation. They use legal AI systems.

Work with us

Legal AI needs benchmarks that evaluate whole systems, not just models.

If you are building legal AI systems, legal-agent workflows, or applied AI evaluations for legal teams, we would be happy to compare notes.

Book a Discovery Call