Skip to content
    HAQQ
    • Preise
    Kostenlos starten
    Kostenlos startenDemo buchen
    Einloggen
    Ressourcen-Hub

    Durchsuchen

    • Rechts-KI Skills338
    • Blog168
      • Änderungen der KI-Verordnung 2026: Was die Verordnung (EU) 2026/1744 geändert hat
      • Rechts-KI-Benchmark 2026: Wie HAQQ in einer unabhängigen Bewertung abschnitt
      • Rechts-KI für Unternehmen: Was große Organisationen prüfen, bevor sie ihr vertrauen
      • Die beste Rechts-KI für Einwanderungsanwälte
      • Rechtsrecherche-KI mit belegten Quellen: Gesetze und Urteile über Jurisdiktionen hinweg
      • Kanzlei 3.0: Warum Kanzleien ein juristisches Betriebssystem brauchen
      • Kann KI Anwälte ersetzen? Nein – und warum das die falsche Frage ist
      • Ist KI-Rechtsberatung sicher und zuverlässig? Was Sie prüfen sollten, bevor Sie einer Legal AI vertrauen
      • Legal-AI-Halluzination: Das gefälschte Zitat, das jede Prüfung besteht
      • KI für Scheidungsanwälte: Der 8-Phasen-Plan (2026)
      • Spellbook-Alternativen: Die besten KI-Tools zur Vertragserstellung 2026
      • Legal-AI-Start-ups in San Francisco: Die Landkarte 2026 (Harvey, Eve, GC AI, Ivo)
      • HAQQ gegen Onit: Legal AI gegenüber Enterprise Legal Management (2026)
      • CoCounsel im Test 2026: Preise, Benchmark & Alternativen
      • Harvey vs Legora vs CoCounsel: Eine 50-Punkte-Rubrik
      • Alle anzeigen62
    • Kostenlose Tools20
    • HAQQ nutzen27
    • KI-Anwalt-Zertifizierung29
    • Lösungen174
    • Whitepapers2
    • Recherche29
    1. Startseite
    2. Blog
    3. Rechts-KI-Benchmark 2026: Wie HAQQ in einer unabhängigen Bewertung abschnitt
    Zurück zum BlogKI & Rechtstechnologie

    Rechts-KI-Benchmark 2026: Wie HAQQ in einer unabhängigen Bewertung abschnitt

    Ein unabhängiger Prüfer testete HAQQ an schwierigen juristischen Aufgaben. Es übertraf die Durchschnittswerte der Rechts-KI und der Branche, führte beim Recht des Nahen Ostens und halluzinierte nicht.

    July 23, 2026
    11 Min. Lesezeit
    |
    HAQQ Research
    Rechts-KI-Benchmark 2026: Wie HAQQ in einer unabhängigen Bewertung abschnitt

    TL;DR — We handed HAQQ to an independent third party and asked them to test it the way a skeptical buyer would: 61 real legal tasks across 21 jurisdictions, every task run twice, no coaching, no special access.

    HAQQ passed 27.0% of tasks overall, above both the legal-AI average and the industry average. On The High Bar — the subset of tasks almost nothing in the field can finish — HAQQ passed 5.9% against 1.1% for the average legal AI: more than five times the field. It was the strongest tool tested on Omani-law drafting, scored 65.2% on scanned and low-quality documents, and hit 100% on the tasks engineered to bait it into inventing content that wasn't there.

    We're publishing our own numbers below. Competing products appear only as an aggregate average — that's the evaluator's disclosure rule, not our choice.

    Why an independent benchmark

    Most legal AI benchmarks are run by the companies selling the product. Ours included, and ours is biased by definition, because it's ours. So we did the opposite: we gave HAQQ to an independent evaluator with no stake in the result and asked them to grade it against the field.

    When a third party with nothing to gain runs the test, the number means something different. That is where the market is heading. Buyers, and increasingly the LLMs that recommend tools, look for trusted independent sources over vendor marketing. We would rather earn the score than claim it.

    How the benchmark works

    A score is only worth as much as the method behind it, so here is the method in full. None of this is our design — it is the evaluator's published methodology, applied identically to every application in the cohort.

    • 61 real tasks, contributed by practising lawyers, spanning contract drafting and information extraction across 21 jurisdictions including the United States, the United Kingdom, India, Singapore, the Netherlands and Oman.
    • Documents arrive as they do in practice — native Word files, PDFs, scans, photographs, and files with tracked changes. Nothing is pre-converted into clean text, so reading the original file is itself part of what's being measured.
    • Binary pass/fail criteria, written by lawyers, fixed before testing. A task passes only when every applicable criterion is met. Where several answers would be professionally defensible, the criteria allow for that — while still failing specific omissions and errors.
    • Two axes, never blended: substance and form. Substance asks whether the work is correct and complete; form asks whether a lawyer could use it as delivered. A polished answer that is legally wrong still fails.
    • Every task is run twice, independently, and the score is the average of the two runs — so a single lucky generation can't carry a result.
    • Two independent LLM judges grade every output, and any disagreement that could change whether a task passes is escalated to a qualified lawyer who reviews the output blind, without knowing which product produced it. The human verdict is final and is preserved for audit.
    • The High Bar is the subset that no more than 2 of the 10 evaluated applications could complete: 34 tasks, scored across 68 attempts. It exists specifically to separate exceptional performance from ordinary competence.

    Two consequences worth internalising before reading any number below. First, pass rates across this benchmark look brutally low compared to the marketing numbers you see elsewhere — that is the all-criteria-or-nothing rule doing its job, and it applies to everyone equally. Second, we never see our competitors' individual scores. The evaluator only ever shows us an aggregate legal-AI average, which is why this post compares HAQQ to averages rather than to named products.

    Key facts

    • Overall pass rate: 27.0%, above both the legal-AI average and the industry average. A task passes only when every lawyer-authored criterion is met, which is why absolute numbers are low across the entire field.
    • The High Bar: 5.9% vs 1.1%. On the hardest subset, HAQQ passed more than five times as often as the average legal AI, and nearly three times the general-purpose average.
    • Strongest tool tested on Omani-law drafting (50.0%) — not competitive, the strongest.
    • 100% on missing-reference handling. On tasks engineered to make a model invent absent schedules and references, HAQQ preserved the gap every single time.
    • 2.53/3 on form — clarity 2.73, structure 2.65 — above both averages. Correct and usable as delivered.
    • The task set was deliberately hard; the whole field scores low, because most tools still fail on genuinely difficult legal work.

    What the benchmark found

    Here are the head-to-head numbers. "Legal AI avg" is the aggregate of the other legal AI applications evaluated; "general-purpose avg" is the frontier chat assistants. Form is scored 1 to 3, where 3 is best.

    MeasureHAQQLegal AI avgGeneral-purpose avg
    Contract drafting — pass rate24.2%20.6%23.9%
    Contract drafting — form (1-3)2.562.442.44
    Data extraction — pass rate30.0%29.6%29.3%
    The High Bar — pass rate5.9%1.1%2.2%

    And the measures where the evaluator reported HAQQ on its own:

    MeasureHAQQ
    Overall pass rate27.0%
    Overall form (1-3)2.53
    Contract drafting — criteria met78.6%
    OCR and scan handling — criteria met65.2%
    Omani-law drafting50.0%
    Missing-reference handling100.0%
    Repeatable success across both runs21.3%

    Contract drafting

    The evaluator scored drafting on two axes: substance (is the answer accurate and complete?) and form (is the deliverable polished and ready to put in front of a client?).

    HAQQ beat both the general-purpose tools and the legal-AI average on each one: 24.2% pass rate against 23.9% and 20.6%, and 2.56/3 on form against 2.44 for both fields. Across individual criteria rather than whole tasks, HAQQ satisfied 78.6% of what the lawyers asked for. Not just technically correct. Correct and client-ready.

    The High Bar: the hardest tasks

    Inside the benchmark sat a smaller set of the hardest problems, formally called The High Bar: the 34 tasks that no more than 2 of the 10 evaluated applications could complete. Multi-step work, multiple reference documents, complex fact patterns. The kind of thing where most AI simply breaks.

    HAQQ passed 5.9% of those 68 attempts. The average legal AI passed 1.1%; the general-purpose average was 2.2%. That is more than five times the legal-AI field and nearly three times the frontier chat assistants — the widest margin HAQQ posted anywhere in the evaluation.

    Read those numbers honestly: 5.9% is a low number in absolute terms, and it should be. These are tasks specifically selected because the entire industry fails them. The gap is the signal here, not the level. This is the part we care about most, because it looks like real legal work rather than a demo.

    Three things that stood out

    1. Middle Eastern legal work

    On drafting governed by Omani law, HAQQ produced the strongest results in the benchmark, at 50.0%. Not competitive. The strongest. That is the thing we set out to build, so it was not a surprise to us, but it is good to see it confirmed by someone with no reason to flatter us.

    2. Scanned and low-resolution documents

    Legal work runs on bad PDFs: faxed contracts, scanned filings, photographed pages. On scanned, photographed, redacted and otherwise imperfect source material, HAQQ met 65.2% of criteria and passed 33.3% of tasks outright — above average on both. It more often extracted legible content while flagging what it couldn't read, because it does not lean on a model's built-in OCR. It runs its own.

    3. It didn't make things up

    The benchmark included tasks engineered to bait a model into inventing what isn't there: missing schedules, absent references. HAQQ scored 100% on the missing-reference checkpoint. When key schedules were missing, it consistently preserved the gap instead of inventing their contents.

    One honest caveat, stated precisely: that 100% is on one specific measure — recognising when a referenced source or requested fact is simply unavailable. On the benchmark's broader hallucination-resistance tasks, HAQQ scored 28.0% pass and 61.7% of criteria met: above average, not perfect. No AI is universally hallucination-free. Inventing absent content happens to be the failure mode that ends careers in law, which is why it is the one we obsess over.

    Why hallucination is the whole game in legal

    In vibe-coding, a hallucination is a bug you catch. In law, a hallucination is a fake case cited to a judge. You have seen the headlines. For a contract or a court filing, a single invented clause or a cross-reference to a law that doesn't exist isn't a rough edge. It is disqualifying.

    So "mostly right" is not a product. The real bar is whether a lawyer can trust the output enough to build on it. Independent testing on absent-content tasks is one of the few honest ways to measure that, and it is exactly where HAQQ was built to be strong.

    The architecture, not the model

    Here is the shift the whole industry is living through: it is no longer the model that wins. Everyone can call the same frontier models. What separates legal AI products now is the software around the model — the retrieval, the tools, the reasoning, the data.

    HAQQ runs on its own engine, Justinian. Think of it as Palantir for legal work: it connects to your existing stack, builds an ontology of your matters, and closes the loop between your documents and the answer.

    Under the hood, HAQQ orchestrates six models — two open-source and four proprietary — behind a router that reads each prompt and picks the right model for the task, balancing accuracy against cost, so no single model runs every job. Around the models sit HAQQ's own tools: the browser, the PDF extractor, the OCR. The models are one component. The extraction, the citation layer, and the routing are ours. That is why it holds up on bad PDFs and refuses to hallucinate where thinner wrappers don't.

    The part no one else is solving: emerging markets

    European and American law is relatively easy for AI. It is digitized, well-documented, and endlessly scraped. That is why every tool looks decent on a UK contract.

    Now point the same tools at the Middle East, or South America, or parts of Africa and Asia. The law is not cleanly digitized. The training data is thin. And the models hallucinate — the same failure you see on Middle Eastern prompts shows up everywhere the data is sparse.

    HAQQ was built for exactly this. We have inscribed and extracted the primary law for these jurisdictions ourselves, which is a moat the general-purpose legal tools do not have. Being strongest in the Middle East in an independent benchmark is not luck. It is the design.

    What legal AI benchmarks should measure next

    Running through this evaluation sharpened our view on what a legal AI benchmark should test. A few things we think the whole field should adopt:

    • Percentile, not pass or fail. "Above average" tells you almost nothing. The 50th and the 99th percentile are both above average, and they are not the same product. Show where a tool sits against the frontier.
    • Difficulty and defensibility. Generating a contract is easy to build. Reliable OCR across bad scans, or extraction that cites the exact clause, can take twelve to eighteen months. Weight how hard a capability is to build, because that is what separates a real lead from a demo.
    • Price and ROI. Quality is close to table stakes at the top now. The question buyers actually ask is: at what cost? Cost per contract, per research task, per page. Measure it, and reward the vendors who are transparent and affordable.
    • Intent inference. Real users don't write long, perfect prompts. They type "draft an NDA" or "review these documents." A good tool infers what they mean and gets to the answer without ten rounds of back-and-forth, the opposite of a chatbot tuned to keep you talking.
    • Citations and verified sources. The web is full of outdated law. What matters is whether a tool cites the specific, current provision. This is where the gap between a general LLM and real legal software shows up fastest.
    • Reasoning on review. For contract review, the value is not only the redline. It is the suggestions, the scenarios, and the visible thought process behind them. Score whether the tool shows its work.
    • Hallucination rate, front and center. For every reason above, this is the number lawyers care about most. It belongs in the headline, not a footnote.

    What's next for this benchmark

    This was the first round, and it covered contract drafting and data extraction. There is more coming.

    HAQQ is currently designated a Contender for Q3 2026 in all five certification categories: Best Legal AI Application (Overall), Contract Drafting, Information Extraction, The High Bar, and User Experience. To be exact about what that means, because the wording matters: Contender is the evaluator's in-quarter designation for a participating application, and certifications are only confirmed once the quarter closes. If no application clears the threshold in a category, no certification is awarded there at all. We will publish where we land, win or lose.

    Next month they expand into legal research and analysis, and after that into more regions, past the traditional US and UK focus, into emerging-market jurisdictions across South America, Africa, Asia, and Southeast Europe. Those are exactly the regions we focus on. They run on a quarterly cadence, because the models move fast enough that any benchmark is half-obsolete the moment a new one ships.

    The honest version

    We are proud of this result and we are not going to oversell it. A 27.0% overall pass rate is a real bar cleared on a deliberately brutal test — and it also means that on roughly three out of four hard tasks, there was still at least one lawyer-authored criterion we missed. Both of those sentences are true, and anyone quoting the first without the second is misreading the benchmark.

    We also can't tell you how far we sit from the single best tool in the field, because the evaluator never shows us competitors individually — only as an average. Being above that average is not the same as being first, and we won't imply otherwise until placements are published. What we do know: on the work that looks like real legal practice, on the hardest tasks in the set, and on the jurisdictions most tools ignore, HAQQ held up under an outside microscope. We'll keep showing up for every round and keep publishing what they find.

    • Justinian — the engine behind these results
    • How HAQQ compares to other legal AI
    • 300-task frontier-model legal benchmark
    • HAQQ-LAB civil-law / MENA benchmark

    HAQQ AI kostenlos testen

    Erleben Sie KI-gestützte juristische Recherche und Entwurf

    H

    HAQQ Research

    Benchmark Series

    Verwandte Ressourcen

    Justinian — the engine behind these resultsHow HAQQ compares to other legal AI300-task frontier-model legal benchmarkHAQQ-LAB civil-law / MENA benchmark

    Verwandte Artikel

    Warum ChatGPT Anwälte enttäuscht: Notizen von 3 US-Anwälten

    Warum ChatGPT Anwälte enttäuscht: Notizen von 3 US-Anwälten

    Spellbook-Alternativen: Die besten KI-Tools zur Vertragserstellung 2026

    Spellbook-Alternativen: Die besten KI-Tools zur Vertragserstellung 2026

    Rechts-KI-Statistiken 2026: Wie viele Anwälte KI wirklich nutzen

    Rechts-KI-Statistiken 2026: Wie viele Anwälte KI wirklich nutzen

    Häufig gestellte Fragen

    Who ran the HAQQ legal AI benchmark?

    An independent third-party evaluator with no commercial stake in HAQQ, running a published methodology applied identically to every application in the cohort. We do not name the evaluator here. Their reports are private by default, but a vendor may publish its own results, which is why HAQQ's full numbers appear in this article while competing products are shown only as an aggregate average.

    How did HAQQ perform in the benchmark?

    HAQQ passed 27.0% of tasks overall and scored 2.53 out of 3 on form, above both the legal-AI average and the industry average. On contract drafting it passed 24.2% against 20.6% for the legal-AI average and 23.9% for general-purpose AI, with a form score of 2.56 against 2.44 for both. On data extraction it passed 30.0% against 29.6% and 29.3%. On The High Bar, the benchmark's hardest subset, it passed 5.9% against 1.1% for the average legal AI.

    What is The High Bar in the legal AI benchmark?

    The High Bar is the subset of tasks that no more than 2 of the 10 evaluated applications could complete: 34 tasks scored across 68 attempts. It exists to separate exceptional performance from ordinary competence. HAQQ passed 5.9% of those attempts, compared with 1.1% for the average legal AI and 2.2% for general-purpose AI — more than five times the legal-AI field. Absolute numbers are low by design, because these are tasks the whole industry fails.

    How is the legal AI benchmark scored?

    61 real tasks contributed by practising lawyers, across 21 jurisdictions, presented as native Word files, PDFs, scans and documents with tracked changes rather than pre-converted text. Every task has binary pass/fail criteria written by lawyers and fixed before testing, and a task passes only when every applicable criterion is met. Substance and form are scored separately and never blended. Every task is run twice independently, two LLM judges grade each output, and any disagreement that could change a pass is escalated to a qualified lawyer who reviews it blind.

    Does HAQQ hallucinate?

    HAQQ scored 100% on the benchmark's missing-reference checkpoint: when key schedules or referenced sources were unavailable, it preserved the gap instead of inventing the contents. On the broader hallucination-resistance tasks it passed 28.0% and met 61.7% of criteria, above average but not perfect. No AI is universally hallucination-free, but inventing absent content is the failure mode that matters most in law, and it is the one HAQQ is built to resist.

    What is the Justinian engine?

    Justinian is HAQQ's own hybrid architecture. It connects to your existing stack, builds an ontology of your matters, and orchestrates six models (two open-source, four proprietary) behind a router, with HAQQ's own OCR, extraction and citation tools around them. It does not rely on a model's built-in OCR, which is why it holds up on scanned and low-resolution documents.

    Is HAQQ good for Middle Eastern and emerging-market law?

    Yes. On drafting governed by Omani law, HAQQ produced the strongest results in the independent benchmark, at 50.0%. HAQQ has inscribed and extracted primary law for jurisdictions that are poorly digitized, across the Middle East and other emerging markets, where general-purpose tools tend to hallucinate.

    Can I see HAQQ's benchmark score?

    Yes. HAQQ's own results are published in full: 27.0% overall pass rate, 2.53 out of 3 on form, 24.2% on contract drafting, 30.0% on data extraction and 5.9% on The High Bar. Competing legal AI products appear only as an aggregate average, because the evaluator does not disclose any application's individual result to another vendor.

    Has HAQQ been certified by the benchmark?

    Not yet. HAQQ is currently designated a Contender for Q3 2026 in all five certification categories: Best Legal AI Application (Overall), Contract Drafting, Information Extraction, The High Bar, and User Experience. Contender is the evaluator's in-quarter designation; certifications are confirmed only when the quarter closes, and if no application clears the threshold in a category, no certification is awarded there.

    Was kommt als Nächstes?

    HAQQ AI kostenlos testen

    Erleben Sie KI-gestützte juristische Recherche und Entwurf

    ROI berechnen

    Sehen Sie, wie viel Zeit und Geld HAQQ Ihrer Kanzlei spart

    380+ juristische Prompts durchsuchen

    Sofort einsatzbereite Prompts für jede juristische Aufgabe

    Zurück zum Blog

    Vorheriger Artikel

    Legal-AI-Preise 2026: jeder veröffentlichte Preis und jeder Anbieter, der keinen nennt

    Nächster Artikel

    Änderungen der KI-Verordnung 2026: Was die Verordnung (EU) 2026/1744 geändert hat

    Setzen Sie das ein

    Stellen Sie HAQQ die Frage, die dieser Artikel bei Ihnen aufgeworfen hat.

    HAQQ across all devices
    HAQQ Legal AI Platform Logo

    Ihr juristischer KI-Zwilling und Kanzleimanagement-System für Entwurf, Abrechnung und Gewinnen.

    Download on theApp StoreGet it onGoogle Play

    Produkt

    • HAQQ Legal AI Chat
    • HAQQ eFirm
    • Mobile App
    • HAQQ eBar
    • HAQQ eWallet
    • Justinian AI Engine
    • HAQQ für Unternehmen
    • Sicherheit
    • Preise

    Lösungen

    • Alle Lösungen
    • Nach Rolle
    • Für Sie
    • Nach Anwendungsfall
    • Nach Funktion
    • Nach Kanzleigröße
    • Nach Land
    • Nach Stadt
    • Spezialisiert
    • Vergleichen Sie uns
    • ROI-Rechner

    Ressourcen

    • Blog
    • HAQQ Academy
    • Prompt-Bibliothek
    • Klausel-Bibliothek
    • Dokumentenbibliothek
    • Rechts-KI Skills
    • Kostenlose Tools
    • Legal AI Index
    • Changelog
    • Status

    Unternehmen

    • Das Team
    • Karriere
    • Presse & Events
    • Partnerschaft
    • Studierende
    • Startup-Programm
    • Kontakt
    • Support
    • Sprachenar en fr es it de pt
    • Kontaktinfo@haqq.ai
    • Statusbetriebsbereit·fundiert
    • Nutzungsbedingungen
    • Datenschutzrichtlinie
    • Cookie-Richtlinie
    • Datenverarbeitung
    • humans.txtlawyers.txtsecurity.txt
    © 2026 HAQQ Inc. Alle Rechte vorbehalten.Produkt intern von HAQQ entwickelt. Website mit modernen Web-Tools gebaut.