In short: most legal AI products are architecturally identical. Your question goes to a general-purpose model with a prompt template wrapped around it, and you cannot tell from the interface. The difference that matters happens before the model is called. Four questions at the end of this post will tell you which kind of system any vendor is selling you, in about ten minutes of a demo.
Every lawyer evaluating legal AI eventually asks the same question, usually about eight minutes into a demo: how is this different from just using ChatGPT?
It is the right question, and most vendors answer it badly. They talk about training data, or a proprietary corpus, or a fine-tune. Those answers are hard to verify and, more often than not, not where the difference actually lives.
The honest answer is architectural, and it comes down to one thing: where does the thinking happen?
The wrapper problem is real, and invisible from the outside
A large share of legal AI products work like this. You type a question. The product wraps it in a prompt template, some instructions about being a helpful legal assistant, maybe a few examples, and forwards the whole thing to a general-purpose model. The model answers. The product formats the answer nicely and shows it to you.
There is nothing dishonest about this. It can be genuinely useful. But it means the product's quality is almost entirely the model's quality, and the model is the same one your opposing counsel can access for twenty dollars a month.
The uncomfortable part for a buyer is that you cannot tell from the interface. A wrapper and a purpose-built engine look identical: a text box, a streaming response, some citations underneath. The difference is upstream, where nobody can see it, which is exactly why so few vendors get pressed on it.
Where the reasoning actually belongs
When a question arrives at HAQQ, it does not go to a general model first. It goes through a step we own and control, whose entire job is to work out what the request actually is and to configure everything that happens next.
That step resolves the things a competent associate would establish before starting work, and it resolves them as structured data rather than as prose:
- What is actually being asked. Not the words typed at 11pm between two calls, but the question underneath them, restated precisely and completely. This runs natively in Arabic and French as well as English, rather than as a translation bolted on afterwards.
- What kind of legal work this is. Drafting, research, review, comparison, a procedural question. Each wants a different shape of answer, and deciding that up front is very different from hoping a general model infers it.
- Which jurisdiction is in play, and separately, what language the governing law is written in. Those two are frequently not the same, and treating them as one is a quiet and common source of error.
- The facts, as objects. Parties, dates, obligations, amounts, governing law, document type, extracted out of prose and turned into structured data before anything expensive reads it.
- What the question needs, and what it does not. Which specialised capabilities to bring to bear, how much reasoning depth the question actually warrants, and whether it falls within scope at all.
Only then does the general-purpose model run. And it arrives already prepared: the relevant capabilities loaded, the facts extracted, the jurisdiction identified, the scope decided.
The model is the last step in the chain, not the first.
Why this produces a better answer, not just a cheaper one
Two mechanisms, and the second surprised us more than the first.
Attention is a budget
A model's context window is finite, and everything you put into it competes for the model's attention. Loading every capability a system has into every request consumes a serious share of that window before the lawyer's actual document gets a look in. Bringing only the relevant ones leaves that room for the thing that matters: the contract, the statute, the facts.
Parsing is not reasoning
This is the larger effect. When a model receives a paragraph of prose, a meaningful part of its work goes into figuring out what the paragraph contains. When it receives structured facts, these parties, this governing law, this obligation, this date, that work is already done, and the whole budget goes to the legal question.
The practical consequence is that the same general-purpose model produces measurably better legal work inside a system like this than it does inside a wrapper. The model did not change. What we handed it did.
The parts nobody puts in a demo
Two things that never make it into a sales deck, and both of which decide whether the product works on a real Tuesday.
Document intake. A large share of legal work does not arrive as clean text. It arrives as a photograph of a stamped document, a scan of a fax, a PDF someone printed and re-scanned crooked. We treat reading those as a first-class engineering problem rather than a preprocessing afterthought, because if a clause number is misread on the way in, no amount of downstream reasoning recovers it. The system will reason impeccably about the wrong number.
Memory as structure, not transcript. The engine holds the relationships between your matters over time and draws on them when a question calls for it. That is different from pasting your last twenty messages back into the prompt. It is what lets the system know that the counterparty in the NDA you are reviewing today is the same one from a dispute two months ago.
What this architecture does not fix
We would rather say this than have you discover it.
It does not eliminate error. Any system built on a language model can produce a wrong answer, and a vendor who tells you otherwise is either not being careful with words or is hoping you will not check. What good architecture does is make wrongness rarer and detectable: claims generated against retrieved sources, citations you can open, and explicit flags where the law is genuinely ambiguous rather than a confident resolution of something unresolved.
It does not remove the lawyer. Everything here is designed around the assumption that a professional reviews the output and signs it. That is not a limitation we are apologising for. It is the design constraint, and it is the reason the architecture looks the way it does.
It does not make coverage universal. Legal AI quality varies by jurisdiction, by how well digitised the source law is, and by how much of it is public at all. Anyone who quotes you a single global coverage number should be asked what it counts. We are asked this frequently, and we would rather answer it about your jurisdiction than wave a number at you.
How to test any vendor for this in a demo
You do not need to see anyone's architecture diagram. Four questions, about ten minutes, and they work on every vendor in the category, including us.
1. Ask it something outside law
A general model wrapped in a legal prompt will usually answer a medical or financial question, because underneath, it is a general model. A system that establishes intent before generating should decline, and say why.
2. Ask the same question in two languages
Not a translated version. The same legal question, once in English and once in Arabic or French. If the substance of the answer changes, language is being handled after the reasoning rather than before it.
3. Give it a bad scan
Photograph a document at an angle, in poor light, and upload it. This is the single most predictive test in the whole demo, because it is the most common real input and the least demoed one.
4. Ask what it does not know
Push toward an ambiguous, unsettled, or thinly sourced area of law. A system optimised to always produce an answer will produce one. A system built for professional work will tell you the ground is uncertain, and show you why.
If a vendor is comfortable with all four, the architecture is probably real. If the demo steers around them, that is information too.
Key takeaways
- The meaningful difference between legal AI products is architectural and invisible from the interface. Ask where the reasoning happens.
- Doing structured work before the model call improves output quality, not just cost. A model handed extracted facts spends its budget on law instead of on parsing.
- The unglamorous components decide real-world performance. Document intake quality is more predictive of whether a tool works on a Tuesday than any benchmark score.
- Four demo questions, off-topic, two languages, a bad scan, and an unsettled area of law, will tell you what kind of system you are being sold.
- The Justinian engine, and how it reasons
- 45 red flags when evaluating a legal AI vendor
- Human in the loop legal AI
- Context engineering for legal AI
Try HAQQ AI Free
Experience AI-powered legal drafting and research



