Skip to content
    HAQQ
    • Prețuri
    Începeți gratuit
    Începeți gratuitRezervați o demonstrație
    Autentificare
    1. Acasă
    2. Blog
    3. Claude Opus 5.5 for legal work: we ran our own benchmark and could not pick a winner
    AI și Legal Tech

    Claude Opus 5.5 for legal work: we ran our own benchmark and could not pick a winner

    We put Opus 5.5, Opus 5 and Fable 5.1 through ten lawyer-grade prompts and 120 blind verdicts. They finished four verdicts apart, in a loop. Only price separated them.

    September 23, 2026
    11 min de citit
    |
    HAQQ Team
    Claude Opus 5.5 for legal work: we ran our own benchmark and could not pick a winner

    In short: we ran Claude Opus 5.5, Opus 5 and Fable 5.1 against ten lawyer-grade prompts and collected 120 blind verdicts from two non-Anthropic judges. The models finished within four verdicts of each other, and the results form a loop rather than a ranking. Price was the one thing that separated cleanly, with Opus 5.5 coming in at half Fable 5.1's cost per answer. We are publishing the parts that went wrong too, because one of them nearly handed us a fake result.

    Everyone is quoting the same three numbers

    Search for a comparison of Anthropic's current models and you will find a page of articles that agree with each other. Opus 5.5 matches Fable 5.1 on most tasks. It is far cheaper. It is faster.

    Those numbers are accurate, and we checked them at the source rather than taking the coverage's word for it. Anthropic's own page states that input and output tokens are "$4 and $20 per million, 20% less than Opus 5", and that "Opus 5.5 also generates output more than 30% faster than Opus 5". Against Fable 5.1's $10 and $50, that is 60% less per token.

    What almost nobody has done is run the models against each other independently. The coverage is real reporting, but it is reporting on a vendor's self-assessment.

    The one independent legal benchmark we could verify at source broadly agrees with the vendor. Vals AI's Legal Research Bench puts Opus 5.5 at 50.48%, ranked 4th of 68 models. That is a good result, and it is also a result about legal research, which is one task. It does not tell an in-house lawyer whether the model drafts a better earn-out clause than the model that costs twice as much.

    We went looking for that answer and could not find it, so we ran it ourselves.

    A note on what we did not publish. Two widely repeated figures about Opus 5.5 on a legal agent benchmark did not survive a check against the primary source: the benchmark's own results page does not list the model at all, and the vendor page does not carry the score attributed to it. We have left both out. If you see those numbers quoted elsewhere, ask where they came from.

    What we ran

    Ten prompts, written as tasks a practising lawyer would recognise rather than exam questions:

    • Redlining a liability cap in a SaaS agreement
    • An IRAC memo on non-compete enforceability
    • Citation-heavy research on personal jurisdiction
    • A cross-border UAE to France data transfer opinion
    • Drafting an earn-out clause
    • EU AI Act triage for a legal-tech product
    • An arbitration versus litigation strategy memo
    • Acqui-hire diligence red flags
    • A client letter delivering an adverse judgment
    • DIFC versus onshore UAE employment gratuity

    Two of those are MENA matters, because that is the law our own product is built for and because it is where the thin training data lives.

    Each model answered all ten. Then every answer met every other answer in a blind pairwise comparison, scored by two judges from outside Anthropic: GPT-5.5 and Gemini 3.1 Pro. Each judge saw each pair twice, once in each order, which cancels the well-documented tendency of a language model to favour whichever answer it reads first. Three pairings, ten prompts, two judges, two orders. 120 verdicts.

    Judges scored four dimensions on a 1 to 10 scale: accuracy, completeness, practicality and drafting quality. They were told that verbosity is not quality, and to judge like a partner reviewing an associate's work.

    Two settings matter more than they look. We pinned reasoning effort to high on all three models, so the comparison measures capability at a fixed setting rather than three different vendor defaults. And we omitted temperature entirely, because Fable 5.1 does not accept the parameter through OpenRouter. Setting it would have handed the two Opus models a determinism Fable could not have.

    The whole run cost $15.26. All model prices in this post are OpenRouter list prices read from the API on 23 September 2026, and all per-answer costs are computed from the tokens our own run actually consumed.

    The result does not form a ranking

    Three head-to-heads, 40 blind verdicts each

    Ten lawyer-grade prompts, two non-Anthropic judges, both presentation orders. Ties shown separately.

    Opus 5.5 v Opus 5

    Opus 5.5
    16
    Opus 5
    22
    Tied
    2

    Opus 5.5 v Fable 5.1

    Opus 5.5
    19
    Fable 5.1
    16
    Tied
    5

    Opus 5 v Fable 5.1

    Opus 5
    17
    Fable 5.1
    20
    Tied
    3

    Read the pattern rather than any single row. Opus 5.5 beats Fable 5.1, Fable 5.1 beats Opus 5, and Opus 5 beats Opus 5.5. No ordering of the three models is consistent with all three results.

    Opus 5.5 beats Fable 5.1. Fable 5.1 beats Opus 5. Opus 5 beats Opus 5.5.

    That is a loop, not a leaderboard. There is no ordering of these three models consistent with all three results, which is the clearest possible signal that the differences are smaller than the measurement.

    Look at it from the other direction and the same thing shows up. Each model appears in two pairings, so each carries 80 verdicts.

    Verdicts won, out of 80 per model

    Every model met both others across all ten prompts, so each carries 80 of the 240 judgements.

    Opus 5
    39 of 80
    Fable 5.1
    36 of 80
    Opus 5.5
    35 of 80

    Four verdicts of spread. On this task set, at this sample size, the three models are interchangeable on quality.

    We could have written the headline the other way. "Opus 5.5 beats Fable 5.1 on legal work" is true, it is what our own numbers say, and it would have travelled further than this post will. It would also have been a three-verdict margin over forty, reported as a finding. We do not think that survives a re-run, and we would rather say so than find out in public.

    What a four-verdict spread actually means

    A benchmark answers the question you can afford to ask, not the question you wanted to ask. Ten prompts is what $15 buys at high reasoning effort. It is enough to detect a large difference and not enough to resolve a small one, and what we found is a small one.

    The honest reading is that on this task set, at this sample size, these three models are interchangeable on quality. Anyone publishing a confident ordering of them from a benchmark this size is reporting noise with a decimal point on it.

    A model that wins 35 to 39 has not lost. A model that wins 85 to 15 has won.

    That is not a criticism of benchmarks. It is an argument for reading them for effect size rather than rank.

    The judges disagreed, and that is load-bearing

    We used two judges specifically so that neither could carry the result alone. They did not agree, and their disagreement is systematic rather than random.

    JudgeOpus 5.5Opus 5Fable 5.1TiesWhat it rewards
    GPT-5.52116230caution, clean authorities
    Gemini 3.1 Pro14231310coverage, completeness

    GPT-5.5 put Fable 5.1 first and never once called a tie. Gemini 3.1 Pro put Opus 5 first by a wide margin and called ten.

    Read the rationales and the reason is visible. GPT-5.5 rewards caution, clean authorities and staying inside the assignment. Gemini 3.1 Pro rewards coverage, and consistently preferred the longer answer that mapped more obligations and flagged more contingencies.

    Neither is wrong. They are two defensible views of what good associate work looks like, and they produce different winners on the same answers. If you take one thing from this post, take that: when a benchmark reports a single score, ask who the judge was and what it rewards, because that choice is doing more work than it appears to.

    One result survives both judges. Opus 5.5 places second with each of them, so neither judge ranks it last.

    Where the scores do separate

    The pairwise verdicts are close. The dimension scores are more revealing, because they show the models failing in different directions.

    Mean judge score by dimension, 1 to 10

    Averaged across every verdict in which each model appeared. Bars are drawn on the full 0 to 10 scale.

    Accuracy

    Opus 5.5
    8.55
    Opus 5
    8.18
    Fable 5.1
    8.55

    Completeness

    Opus 5.5
    8.93
    Opus 5
    9.44
    Fable 5.1
    9.22

    Practicality

    Opus 5.5
    8.72
    Opus 5
    9.03
    Fable 5.1
    8.84

    Drafting

    Opus 5.5
    8.56
    Opus 5
    8.64
    Fable 5.1
    8.70

    The bars look alike because they are alike: the entire spread across three models and four dimensions is 1.26 points. Read the labels, not the lengths. That flatness is the finding, not a rendering problem.

    Accuracy is a tie at the top. Opus 5.5 and Fable 5.1 both average 8.55. Opus 5 sits below both at 8.18.

    Completeness runs the other way. Opus 5 leads at 9.44, Fable 5.1 follows at 9.22, and Opus 5.5 trails at 8.93.

    That inversion is the most useful thing in the run. Opus 5 wrote the longest, most thorough answers in the set and scored highest on completeness and practicality, while scoring lowest on accuracy. Opus 5.5 wrote shorter answers that judges rated more precise. Fable 5.1 sat closest to the top on both.

    For a lot of work, completeness is what you want. For legal work, the ordering between those two dimensions is a choice you should make deliberately rather than inherit from a leaderboard, because an answer that covers more ground and is less accurate is not obviously the better answer to send to a client.

    Încearcă HAQQ AI gratuit

    Experimentează redactarea și cercetarea juridică cu inteligență artificială

    The one thing that is not close

    Cost per answer, measured on real tokens

    Ten legal prompts per model at high reasoning effort, billed at OpenRouter list prices on 23 September 2026.

    Opus 5.5
    $0.3647
    Opus 5
    $0.4362
    Fable 5.1
    $0.7249

    Reasoning tokens bill as output. Opus 5.5 thought the longest of the three, 10,031 reasoning tokens per answer against Fable 5.1's 6,953, and still came out cheapest.

    Measured on real tokens rather than list price, Opus 5.5 answered these ten prompts at $0.3647 per answer. Opus 5 came in at $0.4362. Fable 5.1 came in at $0.7249, which is very nearly double Opus 5.5 for a quality difference our own data cannot resolve.

    Speed pointed the same way. Opus 5.5 averaged 184 seconds per answer, Fable 5.1 193, and Opus 5 238.

    What each model charges to think is worth a paragraph of its own, because it surprised us. Opus 5.5 was the cheapest model and also the one that thought the longest. Fable 5.1 wrote the shortest visible answers of the three and still cost the most, because its list price is 2.5 times Opus 5.5's. Being cheap per token bought Opus 5.5 the room to think harder and still come out ahead on the bill.

    The benchmark nearly lied to us

    This is the part most benchmark write-ups leave out, and it is the part we would most want to read.

    Our first pass capped every answer at 32,000 tokens. We had evidence for that number: an earlier run of the same prompt set proved 32k was truncation-free, with the longest answer coming in under 11,000 tokens.

    That evidence did not transfer, and the reason is a detail worth knowing. Reasoning tokens bill as output and count against the same token budget as the answer. Pinning reasoning effort to high added roughly 8,000 reasoning tokens per answer, and the budget that had been comfortable became tight.

    On the earn-out drafting prompt, both Opus models hit the ceiling and were cut off mid-answer. Fable 5.1 finished normally. So on that prompt we were not comparing three models. We were comparing one model against two models with their endings removed, and the resulting verdicts pointed straight at the models under test.

    Re-run at 64,000 tokens, Opus 5.5 used 51,358 tokens on that single prompt. The 32k cap had been cutting it roughly in half.

    There was a second failure behind the first. We had interrupted a run partway through judging, and it had already cached verdicts computed against the truncated answers. Those cached verdicts survived the answer re-run, because a truncated answer is not an error and looked complete to the cache. We had to find and discard every verdict for the affected prompts and regenerate them.

    It mattered. On the clean but incomplete eight-prompt subset, Fable 5.1 led Opus 5.5. With the two repaired prompts restored, the result reversed. A cap we had good reason to trust, plus a cache we forgot we had, would have produced a confident and wrong headline.

    What we changed

    • The harness now aborts the run if any answer comes back with a finish reason other than a clean stop. Previously it printed a warning and carried on. A warning in a log is not a gate.
    • A truncated cell is no longer treated as cached. The staleness check tests the finish reason, not just the presence of an error field, because truncation is not an error.
    • Every cell writes to disk as it completes, so an interrupted run resumes without re-paying, and every verdict records which answer version it judged.

    We mention this because the failure mode generalises. If you are benchmarking reasoning models against each other and you set a token cap that made sense for non-reasoning models, you are very likely measuring the cap.

    HAQQ's take

    We build legal AI, so we are not neutral about model choice, and we will say plainly what this run changed for us.

    We moved routine legal drafting to Opus 5.5. Not because our data crowns it, which it cannot, but because it is indistinguishable on quality from the two models that cost more, while running faster. When two options are measurably the same and one is half the price, the decision is not close even though the benchmark is.

    We kept Fable 5.1 for the final quality-judged passes, the work where an answer gets read by someone outside the company. That is a hedge rather than a finding. Our numbers do not justify it, and a larger prompt set may well retire it.

    For a legal team choosing a model, the useful conclusion is not which model won. It is that the quality gap between current frontier models on legal tasks is now small enough that price, latency and how a model fails matter more than which one tops a leaderboard. Pick the one whose failure mode you can live with, then test it on your own matters, because your matters are not our ten prompts.

    Limitations

    • Ten prompts is a small sample. It is enough to detect a large difference and it did not find one. Treat every pairwise margin here as provisional.
    • Two judges, both language models. We used non-Anthropic judges and swapped positions to control the obvious biases, but an LLM judge is not a lawyer, and we have written elsewhere about how badly LLM judges can fail on exactly this kind of task.
    • Judge disagreement is unresolved. Our two judges produced different winners. We report both rather than averaging them into a single number that hides the split.
    • Single run, no repeats. We did not run each prompt multiple times, so we cannot separate model variance from model capability.
    • Our prompt set is ours. It leans toward commercial, cross-border and MENA work because that is what we build for.

    Key Takeaways

    • Three frontier Anthropic models finished within four verdicts of each other across the 240 blind judgements we collected on lawyer-grade legal tasks. There is no reliable ordering between them at this sample size.
    • The results are non-transitive: Opus 5.5 beats Fable 5.1, Fable 5.1 beats Opus 5, Opus 5 beats Opus 5.5. That pattern is itself evidence that the differences are smaller than the measurement.
    • Two judges produced two different winners from the same answers. When a benchmark reports one number, ask what its judge rewards.
    • Price was the one clean separation. Opus 5.5 answered at roughly half Fable 5.1's cost per answer, and no model in the set was faster.
    • Accuracy and completeness moved in opposite directions. The most thorough model in the set scored lowest on accuracy.
    • If you pin reasoning effort, re-check your token cap. Reasoning tokens count against the same budget, and a cap that was safe before will silently truncate.

    Sources and further reading

    • Anthropic, Introducing Claude Opus 5.5
    • Vals AI, Claude Opus 5.5 model card
    • Vals AI, Legal Research Bench
    • When an LLM judge rewards fabrication
    • Legal AI benchmark 2026
    • AI legal hallucination audit
    H

    HAQQ Team

    Editorial

    Resurse conexe

    When an LLM judge rewards fabricationLegal AI benchmark 2026AI legal hallucination audit

    Postări similare

    Cea Mai Bună AI pentru Muncă Juridică în 2026? Am Notat 3.000 de Răspunsuri

    Cea Mai Bună AI pentru Muncă Juridică în 2026? Am Notat 3.000 de Răspunsuri

    Am testat GPT-6 Astra pe 41 de întrebări juridice reale

    Am testat GPT-6 Astra pe 41 de întrebări juridice reale

    Legal AI Benchmark 2026: Cum a performat HAQQ într-o evaluare independentă

    Legal AI Benchmark 2026: Cum a performat HAQQ într-o evaluare independentă

    Întrebări frecvente

    Is Claude Opus 5.5 good for legal work?

    On our ten-prompt benchmark it performed indistinguishably from Claude Fable 5.1, Anthropic's most expensive model, and the two tied on accuracy at 8.55 out of 10. Vals AI's Legal Research Bench independently places Opus 5.5 at 50.48%, 4th of 68 models. It is a reasonable default for routine legal drafting, but you should test it on your own matters before relying on it.

    Is Opus 5.5 better than Fable 5.1 for legal research?

    Our run gave Opus 5.5 a narrow win, 19 verdicts to 16 with 5 ties. That margin is too small to be reliable at ten prompts, and the results across all three models formed a loop rather than a ranking. The defensible statement is that we could not separate them on quality, while Opus 5.5 cost about half as much per answer.

    How much cheaper is Claude Opus 5.5 than Fable 5.1?

    On OpenRouter list pricing read on 23 September 2026, Opus 5.5 is $4 per million input tokens and $20 per million output, against Fable 5.1 at $10 and $50. Measured on real answers to our ten legal prompts, Opus 5.5 cost $0.3647 per answer and Fable 5.1 cost $0.7249.

    Which Claude model should a law firm use?

    It depends on which failure you can tolerate. In our run Opus 5 wrote the most complete and practical answers but scored lowest on accuracy. Opus 5.5 wrote shorter answers that judges rated more precise. Fable 5.1 sat near the top on both and costs the most. Test the shortlist on your own matters rather than inheriting a leaderboard.

    How do you benchmark an AI model for legal work?

    Use tasks a lawyer would actually be given rather than exam questions, have answers compared blind and pairwise rather than scored in isolation, swap the presentation order to cancel position bias, and use judges from a different vendor than the models under test. Pin reasoning effort so you are comparing capability rather than defaults, and check that no answer was truncated before you trust a single verdict.

    Why did your benchmark not produce a winner?

    Because the models are close enough that ten prompts cannot separate them. Across the 240 judgements in this run the three models finished four verdicts apart, and the pairwise results were non-transitive. We published that rather than picking whichever ordering our numbers happened to favour.

    Ce urmează?

    Încearcă HAQQ AI gratuit

    Experimentează redactarea și cercetarea juridică cu inteligență artificială

    Calculează-ți ROI-ul

    Vezi cât timp și bani economisește HAQQ firmei tale

    Răsfoiește peste 380 de prompturi juridice

    Prompturi gata de utilizare pentru fiecare sarcină juridică

    Înapoi la blog

    Articolul precedent

    Responsabilitatea operatorului de IA: ce impune noul cod de conduită Microsoft cabinetului tău

    Următorul articol

    Jev într-o arhitectură de IA juridică: 8 tipare de decizie și limita pe care am măsurat-o

    Cuprins

    11 min de citit

    Share this

    Pune asta la treabă

    Adresează-i lui HAQQ întrebarea pe care ți-a ridicat-o acest articol.

    HAQQ across all devices
    HAQQ Legal AI Platform Logo

    Sistemul dvs. Legal AI Twin & Practice Management pentru redactare, facturare și succes.

    Download on theApp StoreGet it onGoogle Play

    Documentații

    • Documentație se deschide într-o filă nouă
    • Noțiuni introductive se deschide într-o filă nouă
    • Presă se deschide într-o filă nouă
    • Actualizări de produs se deschide într-o filă nouă
    • Stare se deschide într-o filă nouă
    • Securitate
    • Întrebări frecvente se deschide într-o filă nouă
    • Comunitate se deschide într-o filă nouă
    • Asistență se deschide într-o filă nouă

    Academie

    • Partener se deschide într-o filă nouă
    • Curs se deschide într-o filă nouă
    • Știri Juridice se deschide într-o filă nouă
    • Competențe se deschide într-o filă nouă
    • Clauză se deschide într-o filă nouă
    • Bibliotecă de prompturi se deschide într-o filă nouă
    • Instrumente se deschide într-o filă nouă
    • Centru de Cercetare se deschide într-o filă nouă
    • Documente se deschide într-o filă nouă

    Site web

    • eFirm
    • Legal AI Chat
    • Aplicație mobilă
    • Justinian AI Engine
    • HAQQ eBar
    • HAQQ eWallet
    • Prețuri
    • Comparați-ne
    • Soluții
    • Blog
    • Faceți cunoștință cu echipa
    • Alăturați-vă nouă se deschide într-o filă nouă
    Deschideți aplicația
    • Localeenarfresitdeptrohi
    • Contactinfo@haqq.ai
    • Stareoperațional·întrerupt
    • Termeni și condiții
    • Politica de confidențialitate
    • Politica privind cookie-urile
    • Prelucrarea datelor
    • Oameni se deschide într-o filă nouăAvocați se deschide într-o filă nouăSecuritate se deschide într-o filă nouă
    © 2026 HAQQ Inc. Toate drepturile rezervate.Produs dezvoltat intern de HAQQ. Site web construit cu instrumente web moderne.