The Brakes Are Untested
Pump your brakes! The AI goldrush has come to the practice of law. In the last two years the number of products marketed as "Legal AI" has gone from a handful to hundreds. It seems like every conference has a booth — every bar journal has a webinar. Wall Street and private venture capital are spending billions in this sector, and firms of all sizes are being pitched legal AI software by swarms of salespeople promising the moon. Of course, vendors are chasing first-mover advantage, because the money is real — for them.
And the demand is rational. Lawyers are being asked to do more with less: leaner staffing, thinner margins, faster turnaround, larger dockets. A tool that promises to do the intake, draft the memo, pull the cases, cite the opinion, and do it in a fraction of the time is not just a temptation — it seems heaven-sent. Of course, firms are buying and adopting it at a faster clip than any software acquisition in history. But whatever happened to caveat emptor?
Here is what most of those products are, under the hood. Modern general-purpose AI models (Claude, ChatGPT, Gemini) are built by three companies: Anthropic, OpenAI, and Google. Those companies rent access to their models through an API — a sort of key to access their large language model, or LLM. Indeed, a large fraction of what is marketed as legal AI is what the industry calls a Claude wrapper or a GPT wrapper: the vendor takes a general-purpose model built by someone else (the AI), writes a system prompt telling it to behave like a legal professional, sometimes adds a retrieval layer that feeds in legal documents, and puts a slicker interface around it. Nothing wrong with that as a construction as far as it goes, and to be fair, much of modern software is built that way.
The problem is what it looks like from the buyer's side. From the buyer's side it looks proprietary. It's marketed with the vocabulary of platforms and engines and reasoning frameworks. But what the law firm is often actually paying for is a subscription markup to use the publicly available AI, wrapped in someone else's user interface and prompt on top. That is not a scandal. But it has an implication most buyers miss: the AI your firm is using is, in many cases, the same underlying model your opposing counsel is using, through a different-colored product wrapped by a different startup. The model is not what differentiates. The prompt, the system of retrieval, and the guardrails built in — which, in the legal profession, are the rules of procedure, statutes, and case law of the jurisdiction — are what differentiate. But these guardrails are often left out because, as one programmer said to me, "Claude just knows." In any case, whatever logic these AI wrappers use is very rarely disclosed, hidden under the banner of intellectual property. None of them are measured independently. This is a recipe for disaster, because AI never says it doesn't know — it just makes stuff up, or hallucinates.
And we already know what happens when nobody is measuring. Judges have addressed AI errors in more than fourteen hundred filings. A Mississippi order disqualified all four lawyers on a case and cancelled the trial after both sides' filings were found to contain AI-fabricated citations. A New York appeals court sanctioned an attorney and his firm $10,500 for a brief citing nonexistent cases. One of the most prestigious firms in the country apologized to a judge for more than forty AI-fabricated errors in a single filing. Every one of those sanctions landed on a signature, not on software. The pattern isn't a scandal. It's a symptom. When the buyer can't tell grounded from fluent, and the seller isn't graded by anyone independent, the errors are structurally certain.
The lawyers being sanctioned right now for AI-fabricated citations are not fools. They are professionals who did what the market told them to do, under time and cost pressure, with a tool nobody handed them a manual for. The mechanism by which a professional convinces themselves they are being careful, while the record shows they were not, does not change because the case-checker is a language model instead of a stack of memos on a Sunday night. It just scales.
I started out trying to build a legal-AI tool. I wanted to see what was on the other side of the marketing — what the actual output looked like on real Minnesota cases, tested at the pincite instead of the headnote. What I found first wasn't a product. It was a gap. There is no independent grading of any of these tools. Every vendor cites its own accuracy number. No two vendors calculate that number the same way. None of the numbers are anchored to what a court actually held. The buyers most exposed to the downside — the firms, and the carriers who underwrite them — are choosing based on demos.
That is the tool I ended up building. Not another legal AI. An independent auditor of the ones already on the market. Ratia — a name from ratio decidendi, the reasoning that decides the case — measures a tool's output against the court-adjudicated record: does the case exist, was the holding read correctly, is it still good law? Reported honestly, by court level, without a flattering blend, in a format an insurance carrier or an ethics panel can read.
The goldrush is not going to stop. Neither are the sanctions. The gap between the two is where the discipline cases will keep coming from. Somebody has to pump the brakes.
If you are evaluating a legal-AI tool before you buy, or you already run one and would like to know what it does on your jurisdiction's law before a court finds out for you, early access to Ratia is opening to a small group of design partners. You can request access at ratialegal.com.