Skip to content
    HAQQ
    • Preços
    Comece Grátis
    Comece GrátisAgendar uma Demo
    Entrar
    Central de recursos

    Explorar

    • Competências IA Jurídicas338
    • Blog168
      • Alterações ao Regulamento de IA da UE 2026: o que mudou com o Regulamento (UE) 2026/1744
      • Benchmark de IA jurídica 2026: o desempenho da HAQQ numa avaliação independente
      • IA Jurídica Corporativa: O Que as Grandes Organizações Verificam Antes de Confiar
      • A Melhor IA Jurídica para Advogados de Imigração
      • IA de Pesquisa Jurídica com Fontes Citadas: Leis e Jurisprudência entre Jurisdições
      • Escritório de Advocacia 3.0: Porque os Escritórios Precisam de um Sistema Operativo Jurídico
      • A IA Pode Substituir Advogados? Não - e Eis Porque Essa é a Pergunta Errada
      • O Aconselhamento Jurídico por IA É Seguro e Preciso? O Que Verificar Antes de Confiar Numa IA Jurídica
      • Alucinação de IA Jurídica: A Citação Falsa Que Passa em Todas as Verificações
      • IA para Advogados de Divórcio: O Manual em 8 Fases (2026)
      • Alternativas ao Spellbook: As Melhores Ferramentas de Redação de Contratos com IA em 2026
      • Startups de IA Jurídica em São Francisco: O Mapa de 2026 (Harvey, Eve, GC AI, Ivo)
      • HAQQ vs. Onit: IA Jurídica vs. Gestão Jurídica Empresarial (2026)
      • Análise do CoCounsel 2026: preços, benchmark e alternativas
      • Harvey vs Legora vs CoCounsel: uma rubrica de 50 pontos
      • Ver tudo62
    • Ferramentas Gratuitas20
    • Como usar o HAQQ27
    • Certificação Advogado IA29
    • Soluções174
    • Livros brancos2
    • Pesquisa29
    1. Início
    2. Blog
    3. Benchmark de IA jurídica 2026: o desempenho da HAQQ numa avaliação independente
    Voltar ao BlogIA & Tech Jurídica

    Benchmark de IA jurídica 2026: o desempenho da HAQQ numa avaliação independente

    Um avaliador independente testou a HAQQ em tarefas jurídicas difíceis. Superou as médias da IA jurídica e do setor, liderou no direito do Médio Oriente e não alucinou.

    July 23, 2026
    11 min de leitura
    |
    HAQQ Research
    Benchmark de IA jurídica 2026: o desempenho da HAQQ numa avaliação independente

    TL;DR — We handed HAQQ to an independent third party and asked them to test it the way a skeptical buyer would: 61 real legal tasks across 21 jurisdictions, every task run twice, no coaching, no special access.

    HAQQ passed 27.0% of tasks overall, above both the legal-AI average and the industry average. On The High Bar — the subset of tasks almost nothing in the field can finish — HAQQ passed 5.9% against 1.1% for the average legal AI: more than five times the field. It was the strongest tool tested on Omani-law drafting, scored 65.2% on scanned and low-quality documents, and hit 100% on the tasks engineered to bait it into inventing content that wasn't there.

    We're publishing our own numbers below. Competing products appear only as an aggregate average — that's the evaluator's disclosure rule, not our choice.

    Why an independent benchmark

    Most legal AI benchmarks are run by the companies selling the product. Ours included, and ours is biased by definition, because it's ours. So we did the opposite: we gave HAQQ to an independent evaluator with no stake in the result and asked them to grade it against the field.

    When a third party with nothing to gain runs the test, the number means something different. That is where the market is heading. Buyers, and increasingly the LLMs that recommend tools, look for trusted independent sources over vendor marketing. We would rather earn the score than claim it.

    How the benchmark works

    A score is only worth as much as the method behind it, so here is the method in full. None of this is our design — it is the evaluator's published methodology, applied identically to every application in the cohort.

    • 61 real tasks, contributed by practising lawyers, spanning contract drafting and information extraction across 21 jurisdictions including the United States, the United Kingdom, India, Singapore, the Netherlands and Oman.
    • Documents arrive as they do in practice — native Word files, PDFs, scans, photographs, and files with tracked changes. Nothing is pre-converted into clean text, so reading the original file is itself part of what's being measured.
    • Binary pass/fail criteria, written by lawyers, fixed before testing. A task passes only when every applicable criterion is met. Where several answers would be professionally defensible, the criteria allow for that — while still failing specific omissions and errors.
    • Two axes, never blended: substance and form. Substance asks whether the work is correct and complete; form asks whether a lawyer could use it as delivered. A polished answer that is legally wrong still fails.
    • Every task is run twice, independently, and the score is the average of the two runs — so a single lucky generation can't carry a result.
    • Two independent LLM judges grade every output, and any disagreement that could change whether a task passes is escalated to a qualified lawyer who reviews the output blind, without knowing which product produced it. The human verdict is final and is preserved for audit.
    • The High Bar is the subset that no more than 2 of the 10 evaluated applications could complete: 34 tasks, scored across 68 attempts. It exists specifically to separate exceptional performance from ordinary competence.

    Two consequences worth internalising before reading any number below. First, pass rates across this benchmark look brutally low compared to the marketing numbers you see elsewhere — that is the all-criteria-or-nothing rule doing its job, and it applies to everyone equally. Second, we never see our competitors' individual scores. The evaluator only ever shows us an aggregate legal-AI average, which is why this post compares HAQQ to averages rather than to named products.

    Key facts

    • Overall pass rate: 27.0%, above both the legal-AI average and the industry average. A task passes only when every lawyer-authored criterion is met, which is why absolute numbers are low across the entire field.
    • The High Bar: 5.9% vs 1.1%. On the hardest subset, HAQQ passed more than five times as often as the average legal AI, and nearly three times the general-purpose average.
    • Strongest tool tested on Omani-law drafting (50.0%) — not competitive, the strongest.
    • 100% on missing-reference handling. On tasks engineered to make a model invent absent schedules and references, HAQQ preserved the gap every single time.
    • 2.53/3 on form — clarity 2.73, structure 2.65 — above both averages. Correct and usable as delivered.
    • The task set was deliberately hard; the whole field scores low, because most tools still fail on genuinely difficult legal work.

    What the benchmark found

    Here are the head-to-head numbers. "Legal AI avg" is the aggregate of the other legal AI applications evaluated; "general-purpose avg" is the frontier chat assistants. Form is scored 1 to 3, where 3 is best.

    MeasureHAQQLegal AI avgGeneral-purpose avg
    Contract drafting — pass rate24.2%20.6%23.9%
    Contract drafting — form (1-3)2.562.442.44
    Data extraction — pass rate30.0%29.6%29.3%
    The High Bar — pass rate5.9%1.1%2.2%

    And the measures where the evaluator reported HAQQ on its own:

    MeasureHAQQ
    Overall pass rate27.0%
    Overall form (1-3)2.53
    Contract drafting — criteria met78.6%
    OCR and scan handling — criteria met65.2%
    Omani-law drafting50.0%
    Missing-reference handling100.0%
    Repeatable success across both runs21.3%

    Contract drafting

    The evaluator scored drafting on two axes: substance (is the answer accurate and complete?) and form (is the deliverable polished and ready to put in front of a client?).

    HAQQ beat both the general-purpose tools and the legal-AI average on each one: 24.2% pass rate against 23.9% and 20.6%, and 2.56/3 on form against 2.44 for both fields. Across individual criteria rather than whole tasks, HAQQ satisfied 78.6% of what the lawyers asked for. Not just technically correct. Correct and client-ready.

    The High Bar: the hardest tasks

    Inside the benchmark sat a smaller set of the hardest problems, formally called The High Bar: the 34 tasks that no more than 2 of the 10 evaluated applications could complete. Multi-step work, multiple reference documents, complex fact patterns. The kind of thing where most AI simply breaks.

    HAQQ passed 5.9% of those 68 attempts. The average legal AI passed 1.1%; the general-purpose average was 2.2%. That is more than five times the legal-AI field and nearly three times the frontier chat assistants — the widest margin HAQQ posted anywhere in the evaluation.

    Read those numbers honestly: 5.9% is a low number in absolute terms, and it should be. These are tasks specifically selected because the entire industry fails them. The gap is the signal here, not the level. This is the part we care about most, because it looks like real legal work rather than a demo.

    Three things that stood out

    1. Middle Eastern legal work

    On drafting governed by Omani law, HAQQ produced the strongest results in the benchmark, at 50.0%. Not competitive. The strongest. That is the thing we set out to build, so it was not a surprise to us, but it is good to see it confirmed by someone with no reason to flatter us.

    2. Scanned and low-resolution documents

    Legal work runs on bad PDFs: faxed contracts, scanned filings, photographed pages. On scanned, photographed, redacted and otherwise imperfect source material, HAQQ met 65.2% of criteria and passed 33.3% of tasks outright — above average on both. It more often extracted legible content while flagging what it couldn't read, because it does not lean on a model's built-in OCR. It runs its own.

    3. It didn't make things up

    The benchmark included tasks engineered to bait a model into inventing what isn't there: missing schedules, absent references. HAQQ scored 100% on the missing-reference checkpoint. When key schedules were missing, it consistently preserved the gap instead of inventing their contents.

    One honest caveat, stated precisely: that 100% is on one specific measure — recognising when a referenced source or requested fact is simply unavailable. On the benchmark's broader hallucination-resistance tasks, HAQQ scored 28.0% pass and 61.7% of criteria met: above average, not perfect. No AI is universally hallucination-free. Inventing absent content happens to be the failure mode that ends careers in law, which is why it is the one we obsess over.

    Why hallucination is the whole game in legal

    In vibe-coding, a hallucination is a bug you catch. In law, a hallucination is a fake case cited to a judge. You have seen the headlines. For a contract or a court filing, a single invented clause or a cross-reference to a law that doesn't exist isn't a rough edge. It is disqualifying.

    So "mostly right" is not a product. The real bar is whether a lawyer can trust the output enough to build on it. Independent testing on absent-content tasks is one of the few honest ways to measure that, and it is exactly where HAQQ was built to be strong.

    The architecture, not the model

    Here is the shift the whole industry is living through: it is no longer the model that wins. Everyone can call the same frontier models. What separates legal AI products now is the software around the model — the retrieval, the tools, the reasoning, the data.

    HAQQ runs on its own engine, Justinian. Think of it as Palantir for legal work: it connects to your existing stack, builds an ontology of your matters, and closes the loop between your documents and the answer.

    Under the hood, HAQQ orchestrates six models — two open-source and four proprietary — behind a router that reads each prompt and picks the right model for the task, balancing accuracy against cost, so no single model runs every job. Around the models sit HAQQ's own tools: the browser, the PDF extractor, the OCR. The models are one component. The extraction, the citation layer, and the routing are ours. That is why it holds up on bad PDFs and refuses to hallucinate where thinner wrappers don't.

    The part no one else is solving: emerging markets

    European and American law is relatively easy for AI. It is digitized, well-documented, and endlessly scraped. That is why every tool looks decent on a UK contract.

    Now point the same tools at the Middle East, or South America, or parts of Africa and Asia. The law is not cleanly digitized. The training data is thin. And the models hallucinate — the same failure you see on Middle Eastern prompts shows up everywhere the data is sparse.

    HAQQ was built for exactly this. We have inscribed and extracted the primary law for these jurisdictions ourselves, which is a moat the general-purpose legal tools do not have. Being strongest in the Middle East in an independent benchmark is not luck. It is the design.

    What legal AI benchmarks should measure next

    Running through this evaluation sharpened our view on what a legal AI benchmark should test. A few things we think the whole field should adopt:

    • Percentile, not pass or fail. "Above average" tells you almost nothing. The 50th and the 99th percentile are both above average, and they are not the same product. Show where a tool sits against the frontier.
    • Difficulty and defensibility. Generating a contract is easy to build. Reliable OCR across bad scans, or extraction that cites the exact clause, can take twelve to eighteen months. Weight how hard a capability is to build, because that is what separates a real lead from a demo.
    • Price and ROI. Quality is close to table stakes at the top now. The question buyers actually ask is: at what cost? Cost per contract, per research task, per page. Measure it, and reward the vendors who are transparent and affordable.
    • Intent inference. Real users don't write long, perfect prompts. They type "draft an NDA" or "review these documents." A good tool infers what they mean and gets to the answer without ten rounds of back-and-forth, the opposite of a chatbot tuned to keep you talking.
    • Citations and verified sources. The web is full of outdated law. What matters is whether a tool cites the specific, current provision. This is where the gap between a general LLM and real legal software shows up fastest.
    • Reasoning on review. For contract review, the value is not only the redline. It is the suggestions, the scenarios, and the visible thought process behind them. Score whether the tool shows its work.
    • Hallucination rate, front and center. For every reason above, this is the number lawyers care about most. It belongs in the headline, not a footnote.

    What's next for this benchmark

    This was the first round, and it covered contract drafting and data extraction. There is more coming.

    HAQQ is currently designated a Contender for Q3 2026 in all five certification categories: Best Legal AI Application (Overall), Contract Drafting, Information Extraction, The High Bar, and User Experience. To be exact about what that means, because the wording matters: Contender is the evaluator's in-quarter designation for a participating application, and certifications are only confirmed once the quarter closes. If no application clears the threshold in a category, no certification is awarded there at all. We will publish where we land, win or lose.

    Next month they expand into legal research and analysis, and after that into more regions, past the traditional US and UK focus, into emerging-market jurisdictions across South America, Africa, Asia, and Southeast Europe. Those are exactly the regions we focus on. They run on a quarterly cadence, because the models move fast enough that any benchmark is half-obsolete the moment a new one ships.

    The honest version

    We are proud of this result and we are not going to oversell it. A 27.0% overall pass rate is a real bar cleared on a deliberately brutal test — and it also means that on roughly three out of four hard tasks, there was still at least one lawyer-authored criterion we missed. Both of those sentences are true, and anyone quoting the first without the second is misreading the benchmark.

    We also can't tell you how far we sit from the single best tool in the field, because the evaluator never shows us competitors individually — only as an average. Being above that average is not the same as being first, and we won't imply otherwise until placements are published. What we do know: on the work that looks like real legal practice, on the hardest tasks in the set, and on the jurisdictions most tools ignore, HAQQ held up under an outside microscope. We'll keep showing up for every round and keep publishing what they find.

    • Justinian — the engine behind these results
    • How HAQQ compares to other legal AI
    • 300-task frontier-model legal benchmark
    • HAQQ-LAB civil-law / MENA benchmark

    Experimente HAQQ AI grátis

    Experimente a redação e pesquisa jurídica com IA

    H

    HAQQ Research

    Benchmark Series

    Recursos relacionados

    Justinian — the engine behind these resultsHow HAQQ compares to other legal AI300-task frontier-model legal benchmarkHAQQ-LAB civil-law / MENA benchmark

    Artigos relacionados

    Por que o ChatGPT falha com advogados: notas de 3 advogados dos EUA

    Por que o ChatGPT falha com advogados: notas de 3 advogados dos EUA

    Alternativas ao Spellbook: As Melhores Ferramentas de Redação de Contratos com IA em 2026

    Alternativas ao Spellbook: As Melhores Ferramentas de Redação de Contratos com IA em 2026

    Estatísticas de IA jurídica 2026: quantos advogados usam IA de fato

    Estatísticas de IA jurídica 2026: quantos advogados usam IA de fato

    Perguntas frequentes

    Who ran the HAQQ legal AI benchmark?

    An independent third-party evaluator with no commercial stake in HAQQ, running a published methodology applied identically to every application in the cohort. We do not name the evaluator here. Their reports are private by default, but a vendor may publish its own results, which is why HAQQ's full numbers appear in this article while competing products are shown only as an aggregate average.

    How did HAQQ perform in the benchmark?

    HAQQ passed 27.0% of tasks overall and scored 2.53 out of 3 on form, above both the legal-AI average and the industry average. On contract drafting it passed 24.2% against 20.6% for the legal-AI average and 23.9% for general-purpose AI, with a form score of 2.56 against 2.44 for both. On data extraction it passed 30.0% against 29.6% and 29.3%. On The High Bar, the benchmark's hardest subset, it passed 5.9% against 1.1% for the average legal AI.

    What is The High Bar in the legal AI benchmark?

    The High Bar is the subset of tasks that no more than 2 of the 10 evaluated applications could complete: 34 tasks scored across 68 attempts. It exists to separate exceptional performance from ordinary competence. HAQQ passed 5.9% of those attempts, compared with 1.1% for the average legal AI and 2.2% for general-purpose AI — more than five times the legal-AI field. Absolute numbers are low by design, because these are tasks the whole industry fails.

    How is the legal AI benchmark scored?

    61 real tasks contributed by practising lawyers, across 21 jurisdictions, presented as native Word files, PDFs, scans and documents with tracked changes rather than pre-converted text. Every task has binary pass/fail criteria written by lawyers and fixed before testing, and a task passes only when every applicable criterion is met. Substance and form are scored separately and never blended. Every task is run twice independently, two LLM judges grade each output, and any disagreement that could change a pass is escalated to a qualified lawyer who reviews it blind.

    Does HAQQ hallucinate?

    HAQQ scored 100% on the benchmark's missing-reference checkpoint: when key schedules or referenced sources were unavailable, it preserved the gap instead of inventing the contents. On the broader hallucination-resistance tasks it passed 28.0% and met 61.7% of criteria, above average but not perfect. No AI is universally hallucination-free, but inventing absent content is the failure mode that matters most in law, and it is the one HAQQ is built to resist.

    What is the Justinian engine?

    Justinian is HAQQ's own hybrid architecture. It connects to your existing stack, builds an ontology of your matters, and orchestrates six models (two open-source, four proprietary) behind a router, with HAQQ's own OCR, extraction and citation tools around them. It does not rely on a model's built-in OCR, which is why it holds up on scanned and low-resolution documents.

    Is HAQQ good for Middle Eastern and emerging-market law?

    Yes. On drafting governed by Omani law, HAQQ produced the strongest results in the independent benchmark, at 50.0%. HAQQ has inscribed and extracted primary law for jurisdictions that are poorly digitized, across the Middle East and other emerging markets, where general-purpose tools tend to hallucinate.

    Can I see HAQQ's benchmark score?

    Yes. HAQQ's own results are published in full: 27.0% overall pass rate, 2.53 out of 3 on form, 24.2% on contract drafting, 30.0% on data extraction and 5.9% on The High Bar. Competing legal AI products appear only as an aggregate average, because the evaluator does not disclose any application's individual result to another vendor.

    Has HAQQ been certified by the benchmark?

    Not yet. HAQQ is currently designated a Contender for Q3 2026 in all five certification categories: Best Legal AI Application (Overall), Contract Drafting, Information Extraction, The High Bar, and User Experience. Contender is the evaluator's in-quarter designation; certifications are confirmed only when the quarter closes, and if no application clears the threshold in a category, no certification is awarded there.

    O que vem a seguir?

    Experimente HAQQ AI grátis

    Experimente a redação e pesquisa jurídica com IA

    Calcule seu ROI

    Veja quanto tempo e dinheiro o HAQQ economiza para seu escritório

    Explore Prompts jurídicos

    Prompts prontos para cada tarefa jurídica

    Voltar ao Blog

    Artigo anterior

    Preços de IA jurídica em 2026: todos os preços publicados e todos os fornecedores que não publicam nenhum

    Próximo artigo

    Alterações ao Regulamento de IA da UE 2026: o que mudou com o Regulamento (UE) 2026/1744

    Ponha isto a trabalhar

    Faça ao HAQQ a pergunta que este artigo lhe levantou.

    HAQQ across all devices
    HAQQ Legal AI Platform Logo

    O seu Gémeo IA Jurídico & Sistema de Gestão de Prática para redigir, faturar e vencer.

    Download on theApp StoreGet it onGoogle Play

    Produto

    • Chat HAQQ IA Jurídica
    • HAQQ eFirm
    • App Móvel
    • HAQQ eBar
    • HAQQ eWallet
    • Motor de IA Justinian
    • HAQQ para Enterprise
    • Segurança
    • Preços

    Soluções

    • Todas as soluções
    • Por função
    • Para você
    • Por caso de uso
    • Por funcionalidade
    • Por tamanho do escritório
    • Por país
    • Por cidade
    • Especializadas
    • Compare-nos
    • Calculadora ROI

    Recursos

    • Blog
    • HAQQ Academy
    • Biblioteca de Prompts
    • Biblioteca de Cláusulas
    • Biblioteca de documentos
    • Competências IA Jurídicas
    • Ferramentas Gratuitas
    • Índice de IA Jurídica
    • Registo de alterações
    • Estado

    Empresa

    • Conheça a Equipa
    • Carreiras
    • Imprensa e Eventos
    • Parceria
    • Estudantes
    • Programa para Startups
    • Contacto
    • Suporte
    • Idiomasar en fr es it de pt
    • Contatoinfo@haqq.ai
    • Estadooperacional·fundamentado
    • Termos de Serviço
    • Política de Privacidade
    • Política de Cookies
    • Processamento de Dados
    • humans.txtlawyers.txtsecurity.txt
    © 2026 HAQQ Inc. Todos os direitos reservados.Produto desenvolvido internamente pela HAQQ. Site construído com ferramentas web modernas.