TL;DR — We handed HAQQ to an independent third party and asked them to test it the way a skeptical buyer would: 61 real legal tasks across 21 jurisdictions, every task run twice, no coaching, no special access.
HAQQ passed 27.0% of tasks overall, above both the legal-AI average and the industry average. On The High Bar — the subset of tasks almost nothing in the field can finish — HAQQ passed 5.9% against 1.1% for the average legal AI: more than five times the field. It was the strongest tool tested on Omani-law drafting, scored 65.2% on scanned and low-quality documents, and hit 100% on the tasks engineered to bait it into inventing content that wasn't there.
We're publishing our own numbers below. Competing products appear only as an aggregate average — that's the evaluator's disclosure rule, not our choice.
Why an independent benchmark
Most legal AI benchmarks are run by the companies selling the product. Ours included, and ours is biased by definition, because it's ours. So we did the opposite: we gave HAQQ to an independent evaluator with no stake in the result and asked them to grade it against the field.
When a third party with nothing to gain runs the test, the number means something different. That is where the market is heading. Buyers, and increasingly the LLMs that recommend tools, look for trusted independent sources over vendor marketing. We would rather earn the score than claim it.
How the benchmark works
A score is only worth as much as the method behind it, so here is the method in full. None of this is our design — it is the evaluator's published methodology, applied identically to every application in the cohort.
- 61 real tasks, contributed by practising lawyers, spanning contract drafting and information extraction across 21 jurisdictions including the United States, the United Kingdom, India, Singapore, the Netherlands and Oman.
- Documents arrive as they do in practice — native Word files, PDFs, scans, photographs, and files with tracked changes. Nothing is pre-converted into clean text, so reading the original file is itself part of what's being measured.
- Binary pass/fail criteria, written by lawyers, fixed before testing. A task passes only when every applicable criterion is met. Where several answers would be professionally defensible, the criteria allow for that — while still failing specific omissions and errors.
- Two axes, never blended: substance and form. Substance asks whether the work is correct and complete; form asks whether a lawyer could use it as delivered. A polished answer that is legally wrong still fails.
- Every task is run twice, independently, and the score is the average of the two runs — so a single lucky generation can't carry a result.
- Two independent LLM judges grade every output, and any disagreement that could change whether a task passes is escalated to a qualified lawyer who reviews the output blind, without knowing which product produced it. The human verdict is final and is preserved for audit.
- The High Bar is the subset that no more than 2 of the 10 evaluated applications could complete: 34 tasks, scored across 68 attempts. It exists specifically to separate exceptional performance from ordinary competence.
Two consequences worth internalising before reading any number below. First, pass rates across this benchmark look brutally low compared to the marketing numbers you see elsewhere — that is the all-criteria-or-nothing rule doing its job, and it applies to everyone equally. Second, we never see our competitors' individual scores. The evaluator only ever shows us an aggregate legal-AI average, which is why this post compares HAQQ to averages rather than to named products.
Key facts
- Overall pass rate: 27.0%, above both the legal-AI average and the industry average. A task passes only when every lawyer-authored criterion is met, which is why absolute numbers are low across the entire field.
- The High Bar: 5.9% vs 1.1%. On the hardest subset, HAQQ passed more than five times as often as the average legal AI, and nearly three times the general-purpose average.
- Strongest tool tested on Omani-law drafting (50.0%) — not competitive, the strongest.
- 100% on missing-reference handling. On tasks engineered to make a model invent absent schedules and references, HAQQ preserved the gap every single time.
- 2.53/3 on form — clarity 2.73, structure 2.65 — above both averages. Correct and usable as delivered.
- The task set was deliberately hard; the whole field scores low, because most tools still fail on genuinely difficult legal work.
What the benchmark found
Here are the head-to-head numbers. "Legal AI avg" is the aggregate of the other legal AI applications evaluated; "general-purpose avg" is the frontier chat assistants. Form is scored 1 to 3, where 3 is best.
| Measure | HAQQ | Legal AI avg | General-purpose avg |
|---|---|---|---|
| Contract drafting — pass rate | 24.2% | 20.6% | 23.9% |
| Contract drafting — form (1-3) | 2.56 | 2.44 | 2.44 |
| Data extraction — pass rate | 30.0% | 29.6% | 29.3% |
| The High Bar — pass rate | 5.9% | 1.1% | 2.2% |
And the measures where the evaluator reported HAQQ on its own:
| Measure | HAQQ |
|---|---|
| Overall pass rate | 27.0% |
| Overall form (1-3) | 2.53 |
| Contract drafting — criteria met | 78.6% |
| OCR and scan handling — criteria met | 65.2% |
| Omani-law drafting | 50.0% |
| Missing-reference handling | 100.0% |
| Repeatable success across both runs | 21.3% |
Contract drafting
The evaluator scored drafting on two axes: substance (is the answer accurate and complete?) and form (is the deliverable polished and ready to put in front of a client?).
HAQQ beat both the general-purpose tools and the legal-AI average on each one: 24.2% pass rate against 23.9% and 20.6%, and 2.56/3 on form against 2.44 for both fields. Across individual criteria rather than whole tasks, HAQQ satisfied 78.6% of what the lawyers asked for. Not just technically correct. Correct and client-ready.
The High Bar: the hardest tasks
Inside the benchmark sat a smaller set of the hardest problems, formally called The High Bar: the 34 tasks that no more than 2 of the 10 evaluated applications could complete. Multi-step work, multiple reference documents, complex fact patterns. The kind of thing where most AI simply breaks.
HAQQ passed 5.9% of those 68 attempts. The average legal AI passed 1.1%; the general-purpose average was 2.2%. That is more than five times the legal-AI field and nearly three times the frontier chat assistants — the widest margin HAQQ posted anywhere in the evaluation.
Read those numbers honestly: 5.9% is a low number in absolute terms, and it should be. These are tasks specifically selected because the entire industry fails them. The gap is the signal here, not the level. This is the part we care about most, because it looks like real legal work rather than a demo.
Three things that stood out
1. Middle Eastern legal work
On drafting governed by Omani law, HAQQ produced the strongest results in the benchmark, at 50.0%. Not competitive. The strongest. That is the thing we set out to build, so it was not a surprise to us, but it is good to see it confirmed by someone with no reason to flatter us.
2. Scanned and low-resolution documents
Legal work runs on bad PDFs: faxed contracts, scanned filings, photographed pages. On scanned, photographed, redacted and otherwise imperfect source material, HAQQ met 65.2% of criteria and passed 33.3% of tasks outright — above average on both. It more often extracted legible content while flagging what it couldn't read, because it does not lean on a model's built-in OCR. It runs its own.
3. It didn't make things up
The benchmark included tasks engineered to bait a model into inventing what isn't there: missing schedules, absent references. HAQQ scored 100% on the missing-reference checkpoint. When key schedules were missing, it consistently preserved the gap instead of inventing their contents.
One honest caveat, stated precisely: that 100% is on one specific measure — recognising when a referenced source or requested fact is simply unavailable. On the benchmark's broader hallucination-resistance tasks, HAQQ scored 28.0% pass and 61.7% of criteria met: above average, not perfect. No AI is universally hallucination-free. Inventing absent content happens to be the failure mode that ends careers in law, which is why it is the one we obsess over.
Why hallucination is the whole game in legal
In vibe-coding, a hallucination is a bug you catch. In law, a hallucination is a fake case cited to a judge. You have seen the headlines. For a contract or a court filing, a single invented clause or a cross-reference to a law that doesn't exist isn't a rough edge. It is disqualifying.
So "mostly right" is not a product. The real bar is whether a lawyer can trust the output enough to build on it. Independent testing on absent-content tasks is one of the few honest ways to measure that, and it is exactly where HAQQ was built to be strong.
The architecture, not the model
Here is the shift the whole industry is living through: it is no longer the model that wins. Everyone can call the same frontier models. What separates legal AI products now is the software around the model — the retrieval, the tools, the reasoning, the data.
HAQQ runs on its own engine, Justinian. Think of it as Palantir for legal work: it connects to your existing stack, builds an ontology of your matters, and closes the loop between your documents and the answer.
Under the hood, HAQQ orchestrates six models — two open-source and four proprietary — behind a router that reads each prompt and picks the right model for the task, balancing accuracy against cost, so no single model runs every job. Around the models sit HAQQ's own tools: the browser, the PDF extractor, the OCR. The models are one component. The extraction, the citation layer, and the routing are ours. That is why it holds up on bad PDFs and refuses to hallucinate where thinner wrappers don't.
The part no one else is solving: emerging markets
European and American law is relatively easy for AI. It is digitized, well-documented, and endlessly scraped. That is why every tool looks decent on a UK contract.
Now point the same tools at the Middle East, or South America, or parts of Africa and Asia. The law is not cleanly digitized. The training data is thin. And the models hallucinate — the same failure you see on Middle Eastern prompts shows up everywhere the data is sparse.
HAQQ was built for exactly this. We have inscribed and extracted the primary law for these jurisdictions ourselves, which is a moat the general-purpose legal tools do not have. Being strongest in the Middle East in an independent benchmark is not luck. It is the design.
What legal AI benchmarks should measure next
Running through this evaluation sharpened our view on what a legal AI benchmark should test. A few things we think the whole field should adopt:
- Percentile, not pass or fail. "Above average" tells you almost nothing. The 50th and the 99th percentile are both above average, and they are not the same product. Show where a tool sits against the frontier.
- Difficulty and defensibility. Generating a contract is easy to build. Reliable OCR across bad scans, or extraction that cites the exact clause, can take twelve to eighteen months. Weight how hard a capability is to build, because that is what separates a real lead from a demo.
- Price and ROI. Quality is close to table stakes at the top now. The question buyers actually ask is: at what cost? Cost per contract, per research task, per page. Measure it, and reward the vendors who are transparent and affordable.
- Intent inference. Real users don't write long, perfect prompts. They type "draft an NDA" or "review these documents." A good tool infers what they mean and gets to the answer without ten rounds of back-and-forth, the opposite of a chatbot tuned to keep you talking.
- Citations and verified sources. The web is full of outdated law. What matters is whether a tool cites the specific, current provision. This is where the gap between a general LLM and real legal software shows up fastest.
- Reasoning on review. For contract review, the value is not only the redline. It is the suggestions, the scenarios, and the visible thought process behind them. Score whether the tool shows its work.
- Hallucination rate, front and center. For every reason above, this is the number lawyers care about most. It belongs in the headline, not a footnote.
What's next for this benchmark
This was the first round, and it covered contract drafting and data extraction. There is more coming.
HAQQ is currently designated a Contender for Q3 2026 in all five certification categories: Best Legal AI Application (Overall), Contract Drafting, Information Extraction, The High Bar, and User Experience. To be exact about what that means, because the wording matters: Contender is the evaluator's in-quarter designation for a participating application, and certifications are only confirmed once the quarter closes. If no application clears the threshold in a category, no certification is awarded there at all. We will publish where we land, win or lose.
Next month they expand into legal research and analysis, and after that into more regions, past the traditional US and UK focus, into emerging-market jurisdictions across South America, Africa, Asia, and Southeast Europe. Those are exactly the regions we focus on. They run on a quarterly cadence, because the models move fast enough that any benchmark is half-obsolete the moment a new one ships.
The honest version
We are proud of this result and we are not going to oversell it. A 27.0% overall pass rate is a real bar cleared on a deliberately brutal test — and it also means that on roughly three out of four hard tasks, there was still at least one lawyer-authored criterion we missed. Both of those sentences are true, and anyone quoting the first without the second is misreading the benchmark.
We also can't tell you how far we sit from the single best tool in the field, because the evaluator never shows us competitors individually — only as an average. Being above that average is not the same as being first, and we won't imply otherwise until placements are published. What we do know: on the work that looks like real legal practice, on the hardest tasks in the set, and on the jurisdictions most tools ignore, HAQQ held up under an outside microscope. We'll keep showing up for every round and keep publishing what they find.
- Justinian — the engine behind these results
- How HAQQ compares to other legal AI
- 300-task frontier-model legal benchmark
- HAQQ-LAB civil-law / MENA benchmark
Essayer HAQQ AI gratuitement
Découvrez la rédaction et la recherche juridique par IA



