← Back to HAQQ Blog

Legal AI Benchmark 2026: How HAQQ Performed in an Independent Evaluation

By HAQQ Research · · 9 min read · Ai-legal-tech

An independent evaluator benchmarked HAQQ on hard legal tasks. It beat the legal-AI and industry averages, led on Middle Eastern law, and didn't hallucinate.

Why an independent benchmark

Most legal AI benchmarks are run by the companies selling the product. Ours included, and ours is biased by definition, because it's ours. So we did the opposite: we gave HAQQ to an independent evaluator with no stake in the result and asked them to grade it against the field.

When a third party with nothing to gain runs the test, the number means something different. That is where the market is heading. Buyers, and increasingly the LLMs that recommend tools, look for trusted independent sources over vendor marketing. We would rather earn the score than claim it.

Key facts

What the benchmark found

The evaluator shared qualitative placements rather than raw scores (the full numbers sit in a private report). Here is where HAQQ landed:

AreaResult
Overall (drafting + extraction)Above legal-AI average and industry average
Contract drafting — substanceAbove legal-AI and industry average
Contract drafting — formAbove legal-AI and industry average
Challenge set (hardest tasks)Well above average vs both fields
Middle Eastern legal workStrongest tool tested
Scanned / low-resolution PDFsHigh-quality parsing via in-house OCR
Absent-content tasksZero hallucinations

Contract drafting

The evaluator scored drafting on two axes: substance (is the answer accurate and complete?) and form (is the deliverable polished and ready to put in front of a client?).

HAQQ beat both the general-purpose tools and the legal-AI average on each one. Not just technically correct. Correct and client-ready.

The challenge set: the hardest tasks

Inside the benchmark sat a smaller set of the hardest problems: multi-step tasks, multiple reference documents, complex fact patterns. The kind of work where most AI simply breaks.

On that subset, HAQQ scored well above average against both general-purpose AI and other legal AI. This is the part we care about most, because it looks like real legal work rather than a demo.

Three things that stood out

On tasks involving Middle Eastern legal reasoning and analysis, HAQQ was the strongest tool in the benchmark. Not competitive. The strongest. That is the thing we set out to build, so it was not a surprise to us, but it is good to see it confirmed by someone with no reason to flatter us.

2. Scanned and low-resolution documents

Legal work runs on bad PDFs: faxed contracts, scanned filings, photographed pages. HAQQ parsed low-resolution documents and returned high-quality results, because it does not lean on a model's built-in OCR. It runs its own.

3. It didn't make things up

The benchmark included tasks engineered to bait a model into inventing what isn't there: missing schedules, absent references. On those tasks, HAQQ produced zero hallucinations. When information was missing, it left the gap instead of filling it with fiction.

In vibe-coding, a hallucination is a bug you catch. In law, a hallucination is a fake case cited to a judge. You have seen the headlines. For a contract or a court filing, a single invented clause or a cross-reference to a law that doesn't exist isn't a rough edge. It is disqualifying.

So "mostly right" is not a product. The real bar is whether a lawyer can trust the output enough to build on it. Independent testing on absent-content tasks is one of the few honest ways to measure that, and it is exactly where HAQQ was built to be strong.

The architecture, not the model

Here is the shift the whole industry is living through: it is no longer the model that wins. Everyone can call the same frontier models. What separates legal AI products now is the software around the model — the retrieval, the tools, the reasoning, the data.

HAQQ runs on its own engine, Justinian. Think of it as Palantir for legal work: it connects to your existing stack, builds an ontology of your matters, and closes the loop between your documents and the answer.

Under the hood, HAQQ orchestrates six models — two open-source and four proprietary — behind a router that reads each prompt and picks the right model for the task, balancing accuracy against cost, so no single model runs every job. Around the models sit HAQQ's own tools: the browser, the PDF extractor, the OCR. The models are one component. The extraction, the citation layer, and the routing are ours. That is why it holds up on bad PDFs and refuses to hallucinate where thinner wrappers don't.

The part no one else is solving: emerging markets

European and American law is relatively easy for AI. It is digitized, well-documented, and endlessly scraped. That is why every tool looks decent on a UK contract.

Now point the same tools at the Middle East, or South America, or parts of Africa and Asia. The law is not cleanly digitized. The training data is thin. And the models hallucinate — the same failure you see on Middle Eastern prompts shows up everywhere the data is sparse.

HAQQ was built for exactly this. We have inscribed and extracted the primary law for these jurisdictions ourselves, which is a moat the general-purpose legal tools do not have. Being strongest in the Middle East in an independent benchmark is not luck. It is the design.

Running through this evaluation sharpened our view on what a legal AI benchmark should test. A few things we think the whole field should adopt:

What's next for this benchmark

This was the first round, and it covered contract drafting and data extraction. There is more coming.

At the end of August, the evaluator will certify the top-ranked applications across six categories, among them contract drafting, data extraction, product experience, the challenge set of hardest tasks, and overall best legal AI application. We are in the running, and we will publish where we land, win or lose.

Next month they expand into legal research and analysis, and after that into more regions, past the traditional US and UK focus, into emerging-market jurisdictions across South America, Africa, Asia, and Southeast Europe. Those are exactly the regions we focus on. They run on a quarterly cadence, because the models move fast enough that any benchmark is half-obsolete the moment a new one ships.

The honest version

We are proud of this result and we are not going to oversell it. "Above average" is a real bar cleared on a hard test, not a coronation. We don't yet know how far we sit from the very top, and we'll say so when the evaluator publishes placements. What we do know: on the work that looks like real legal practice, and on the jurisdictions most tools ignore, HAQQ held up under an outside microscope. We'll keep showing up for every round and keep publishing what they find.

FAQ

Who ran the HAQQ legal AI benchmark?

An independent third-party evaluator with no commercial stake in HAQQ. Their detailed reports are private by default, so we share the findings we are permitted to share and do not name the evaluator here.

How did HAQQ perform in the benchmark?

Overall, HAQQ finished above both the legal-AI average and the general-purpose (industry) average. It beat both averages on contract drafting for substance and form, scored well above average on the hardest challenge-set tasks, and was the strongest tool tested on Middle Eastern legal work.

Does HAQQ hallucinate?

On the benchmark's absent-content tasks, designed to bait a model into inventing missing schedules and references, HAQQ produced zero hallucinations, preserving the gaps instead of filling them. No AI is universally hallucination-free, but inventing absent content is the failure mode that matters most in law, and it is the one HAQQ is built to resist.

What is the Justinian engine?

Justinian is HAQQ's own hybrid architecture. It connects to your existing stack, builds an ontology of your matters, and orchestrates six models (two open-source, four proprietary) behind a router, with HAQQ's own OCR, extraction and citation tools around them. It does not rely on a model's built-in OCR, which is why it holds up on scanned and low-resolution documents.

Is HAQQ good for Middle Eastern and emerging-market law?

Yes. HAQQ was the strongest tool in the independent benchmark on Middle Eastern legal reasoning. HAQQ has inscribed and extracted primary law for jurisdictions that are poorly digitized, across the Middle East and other emerging markets, where general-purpose tools tend to hallucinate.

Can I see HAQQ's benchmark score?

The evaluator's detailed scores are in a private report. Certifications for top-ranked applications across six categories are expected at the end of August; we will publish where HAQQ lands once results are final.