Legal AI Benchmark 2026: How HAQQ Performed in an Independent Evaluation
An independent evaluator benchmarked HAQQ on hard legal tasks. It beat the legal-AI and industry averages, led on Middle Eastern law, and didn't hallucinate.
Why an independent benchmark
Most legal AI benchmarks are run by the companies selling the product. Ours included, and ours is biased by definition, because it's ours. So we did the opposite: we gave HAQQ to an independent evaluator with no stake in the result and asked them to grade it against the field.
When a third party with nothing to gain runs the test, the number means something different. That is where the market is heading. Buyers, and increasingly the LLMs that recommend tools, look for trusted independent sources over vendor marketing. We would rather earn the score than claim it.
Key facts
- HAQQ finished above both the legal-AI average and the general-purpose (industry) average across the tested categories.
- Strongest tool in the benchmark on Middle Eastern legal reasoning — not competitive, the strongest.
- Zero hallucinations on the benchmark's absent-content tasks, which were engineered to make a model invent missing schedules and references.
- The task set was deliberately hard; the industry average came in low, because most tools still fail on genuinely difficult legal work.
What the benchmark found
The evaluator shared qualitative placements rather than raw scores (the full numbers sit in a private report). Here is where HAQQ landed:
| Area | Result |
|---|---|
| Overall (drafting + extraction) | Above legal-AI average and industry average |
| Contract drafting — substance | Above legal-AI and industry average |
| Contract drafting — form | Above legal-AI and industry average |
| Challenge set (hardest tasks) | Well above average vs both fields |
| Middle Eastern legal work | Strongest tool tested |
| Scanned / low-resolution PDFs | High-quality parsing via in-house OCR |
| Absent-content tasks | Zero hallucinations |
Contract drafting
The evaluator scored drafting on two axes: substance (is the answer accurate and complete?) and form (is the deliverable polished and ready to put in front of a client?).
HAQQ beat both the general-purpose tools and the legal-AI average on each one. Not just technically correct. Correct and client-ready.
The challenge set: the hardest tasks
Inside the benchmark sat a smaller set of the hardest problems: multi-step tasks, multiple reference documents, complex fact patterns. The kind of work where most AI simply breaks.
On that subset, HAQQ scored well above average against both general-purpose AI and other legal AI. This is the part we care about most, because it looks like real legal work rather than a demo.
Three things that stood out
1. Middle Eastern legal work
On tasks involving Middle Eastern legal reasoning and analysis, HAQQ was the strongest tool in the benchmark. Not competitive. The strongest. That is the thing we set out to build, so it was not a surprise to us, but it is good to see it confirmed by someone with no reason to flatter us.
2. Scanned and low-resolution documents
Legal work runs on bad PDFs: faxed contracts, scanned filings, photographed pages. HAQQ parsed low-resolution documents and returned high-quality results, because it does not lean on a model's built-in OCR. It runs its own.
3. It didn't make things up
The benchmark included tasks engineered to bait a model into inventing what isn't there: missing schedules, absent references. On those tasks, HAQQ produced zero hallucinations. When information was missing, it left the gap instead of filling it with fiction.
Why hallucination is the whole game in legal
In vibe-coding, a hallucination is a bug you catch. In law, a hallucination is a fake case cited to a judge. You have seen the headlines. For a contract or a court filing, a single invented clause or a cross-reference to a law that doesn't exist isn't a rough edge. It is disqualifying.
So "mostly right" is not a product. The real bar is whether a lawyer can trust the output enough to build on it. Independent testing on absent-content tasks is one of the few honest ways to measure that, and it is exactly where HAQQ was built to be strong.
The architecture, not the model
Here is the shift the whole industry is living through: it is no longer the model that wins. Everyone can call the same frontier models. What separates legal AI products now is the software around the model — the retrieval, the tools, the reasoning, the data.
HAQQ runs on its own engine, Justinian. Think of it as Palantir for legal work: it connects to your existing stack, builds an ontology of your matters, and closes the loop between your documents and the answer.
Under the hood, HAQQ orchestrates six models — two open-source and four proprietary — behind a router that reads each prompt and picks the right model for the task, balancing accuracy against cost, so no single model runs every job. Around the models sit HAQQ's own tools: the browser, the PDF extractor, the OCR. The models are one component. The extraction, the citation layer, and the routing are ours. That is why it holds up on bad PDFs and refuses to hallucinate where thinner wrappers don't.
The part no one else is solving: emerging markets
European and American law is relatively easy for AI. It is digitized, well-documented, and endlessly scraped. That is why every tool looks decent on a UK contract.
Now point the same tools at the Middle East, or South America, or parts of Africa and Asia. The law is not cleanly digitized. The training data is thin. And the models hallucinate — the same failure you see on Middle Eastern prompts shows up everywhere the data is sparse.
HAQQ was built for exactly this. We have inscribed and extracted the primary law for these jurisdictions ourselves, which is a moat the general-purpose legal tools do not have. Being strongest in the Middle East in an independent benchmark is not luck. It is the design.
What legal AI benchmarks should measure next
Running through this evaluation sharpened our view on what a legal AI benchmark should test. A few things we think the whole field should adopt:
- Percentile, not pass or fail. "Above average" tells you almost nothing. The 50th and the 99th percentile are both above average, and they are not the same product. Show where a tool sits against the frontier.
- Difficulty and defensibility. Generating a contract is easy to build. Reliable OCR across bad scans, or extraction that cites the exact clause, can take twelve to eighteen months. Weight how hard a capability is to build, because that is what separates a real lead from a demo.
- Price and ROI. Quality is close to table stakes at the top now. The question buyers actually ask is: at what cost? Cost per contract, per research task, per page. Measure it, and reward the vendors who are transparent and affordable.
- Intent inference. Real users don't write long, perfect prompts. They type "draft an NDA" or "review these documents." A good tool infers what they mean and gets to the answer without ten rounds of back-and-forth, the opposite of a chatbot tuned to keep you talking.
- Citations and verified sources. The web is full of outdated law. What matters is whether a tool cites the specific, current provision. This is where the gap between a general LLM and real legal software shows up fastest.
- Reasoning on review. For contract review, the value is not only the redline. It is the suggestions, the scenarios, and the visible thought process behind them. Score whether the tool shows its work.
- Hallucination rate, front and center. For every reason above, this is the number lawyers care about most. It belongs in the headline, not a footnote.
What's next for this benchmark
This was the first round, and it covered contract drafting and data extraction. There is more coming.
At the end of August, the evaluator will certify the top-ranked applications across six categories, among them contract drafting, data extraction, product experience, the challenge set of hardest tasks, and overall best legal AI application. We are in the running, and we will publish where we land, win or lose.
Next month they expand into legal research and analysis, and after that into more regions, past the traditional US and UK focus, into emerging-market jurisdictions across South America, Africa, Asia, and Southeast Europe. Those are exactly the regions we focus on. They run on a quarterly cadence, because the models move fast enough that any benchmark is half-obsolete the moment a new one ships.
The honest version
We are proud of this result and we are not going to oversell it. "Above average" is a real bar cleared on a hard test, not a coronation. We don't yet know how far we sit from the very top, and we'll say so when the evaluator publishes placements. What we do know: on the work that looks like real legal practice, and on the jurisdictions most tools ignore, HAQQ held up under an outside microscope. We'll keep showing up for every round and keep publishing what they find.
- Justinian — the engine behind these results
- How HAQQ compares to other legal AI
- 300-task frontier-model legal benchmark
- HAQQ-LAB civil-law / MENA benchmark
FAQ
Who ran the HAQQ legal AI benchmark?
An independent third-party evaluator with no commercial stake in HAQQ. Their detailed reports are private by default, so we share the findings we are permitted to share and do not name the evaluator here.
How did HAQQ perform in the benchmark?
Overall, HAQQ finished above both the legal-AI average and the general-purpose (industry) average. It beat both averages on contract drafting for substance and form, scored well above average on the hardest challenge-set tasks, and was the strongest tool tested on Middle Eastern legal work.
Does HAQQ hallucinate?
On the benchmark's absent-content tasks, designed to bait a model into inventing missing schedules and references, HAQQ produced zero hallucinations, preserving the gaps instead of filling them. No AI is universally hallucination-free, but inventing absent content is the failure mode that matters most in law, and it is the one HAQQ is built to resist.
What is the Justinian engine?
Justinian is HAQQ's own hybrid architecture. It connects to your existing stack, builds an ontology of your matters, and orchestrates six models (two open-source, four proprietary) behind a router, with HAQQ's own OCR, extraction and citation tools around them. It does not rely on a model's built-in OCR, which is why it holds up on scanned and low-resolution documents.
Is HAQQ good for Middle Eastern and emerging-market law?
Yes. HAQQ was the strongest tool in the independent benchmark on Middle Eastern legal reasoning. HAQQ has inscribed and extracted primary law for jurisdictions that are poorly digitized, across the Middle East and other emerging markets, where general-purpose tools tend to hallucinate.
Can I see HAQQ's benchmark score?
The evaluator's detailed scores are in a private report. Certifications for top-ranked applications across six categories are expected at the end of August; we will publish where HAQQ lands once results are final.