Vetting AI for Attorneys
The Landscape Has Changed—You Should too
When ChatGPT launched in late 2022, generative AI was a curiosity. Three-plus years later, it is a baseline expectation. Clients ask whether you use it (to save them money), judges increasingly ask you to disclose it, malpractice carriers ask how you control it, and bar regulators have published opinions telling you what they expect when you do. The right question is no longer “should I try AI?” It is “which AI, for which task, under which controls, and how do I know it is working?”
This paper is a vetting framework, not a buyer’s guide. The tools change too fast for any fixed list to stay current. Casetext, which a few years ago was the consumer-facing name for legal AI, was acquired by Thomson Reuters in 2023 and shut down as a standalone product on April 1, 2025, with its capabilities folded into CoCounsel inside the Westlaw ecosystem. New entrants (Harvey, Spellbook, Paxton, Vincent AI, Alexi, Briefpoint, Clio Work, MyCase IQ, EvenUp, Robin AI, and dozens of others) are released, repriced, or rebuilt faster than we’d ever imagine software changing. What stays constant is what you should examine when deciding whether a tool belongs in your firm: the use case, the ethics floor, the vendor’s data handling, the tool’s accuracy on your work, and the operational fit. Walk through those, in that order, and you will pick well.
Start With the Job, Not the Tool
Almost every disappointing AI rollout starts the same way: a partner reads about a tool, the firm buys seats, and then someone asks what the seats are for. Reverse the order. Define the job first.
The categories below cover where lawyers are getting real value today. Within each category, the strongest tools as of mid-2026 fall into three buckets: legal-specific platforms (built on top of general-purpose models with legal-grade retrieval, citation, and security), general-purpose AI assistants (ChatGPT, Claude, Gemini, Microsoft Copilot—typically used inside an enterprise tenant), and embedded AI inside the software you already own (Microsoft 365 Copilot, Clio Work, MyCase IQ, NetDocuments ndMAX, iManage Insight+). Each bucket has a different security posture and price tag.
Legal Research and Memo Drafting
Cite-aware legal-research platforms are among the best-developed corner of the market. Westlaw AI-Assisted Research and CoCounsel (Thomson Reuters), Lexis+ AI and Protégé (LexisNexis), Vincent AI (vLex), Paxton, and Alexi all combine a foundation model with retrieval-augmented generation against a curated legal corpus. They answer questions, draft memos, and produce citations you can click to verify. Their value depends entirely on whether the corpus covers your jurisdiction and practice area.
Contract Drafting, Review, and Analysis
Spellbook, Robin AI, Diligen, LinkSquares, Ironclad’s AI features, and CoCounsel handle redlining, clause comparison, playbook enforcement, and risk flagging in transactional work. Harvey is the dominant choice in BigLaw for matter-level workflows that span drafting, due diligence, and litigation support. For solos and small firms, Spellbook and Briefpoint hit a sweet spot of price and integration with Microsoft Word.
Discovery and Document Review
eDiscovery vendors integrated generative AI quickly. Relativity aiR, Everlaw AI Assistant, DISCO Cecilia, Reveal, and Logikcull use AI for first-pass review, privilege screening, deposition summaries, and narrative generation. These products are mature, sitting on top of years of TAR (technology-assisted review) work, and they are the AI deployments most likely to deliver cost savings on litigation matters.
Drafting, Editing, and Office Productivity
For everyday firm office work, general-purpose models inside an enterprise tenant do the job at a fraction of the cost of legal-specific tools. Microsoft 365 Copilot, ChatGPT Enterprise/Team, Claude Team/Enterprise, and Google Workspace’s Gemini features all offer no-training contractual terms and SOC 2-level controls. Pick one tenant-level deployment, give your staff training, and you will recover the cost on time saved rapidly.
Client Intake, Communication, and Translation
AI-driven intake (Lawmatics AI, Clio Grow, Smith.ai’s chat) and translation use cases are low-risk and high-value. They sit at the edge of the firm and rarely touch protected work product. They are also where AI is most likely to interact with the public, so they need clear disclosure that the user is talking to a bot.
Predictive Analytics
Settlement valuation, motion-outcome prediction, judge analytics, and damages modeling have moved from research projects to commercial products (EvenUp for personal injury demand letters, Lex Machina and Trellis for analytics, Solomon for valuation). These tools are useful but only as good as their training data, so ask what data they were trained on and whether your jurisdiction is represented.
Ethical Floor: ABA Formal Opinion 512 and the State Bars
In July 2024, the ABA Standing Committee on Ethics and Professional Responsibility issued Formal Opinion 512, its first formal guidance on generative AI (“GAI”). The opinion does not invent new rules; it explains how existing Model Rules apply to GAI use. Six areas:
Competence (Rule 1.1)
Lawyers must have a reasonable understanding of the capabilities and limitations of the AI tools they use. You do not need to become an AI expert, but you do need to know enough to supervise the output and recognize when the tool is wrong.
Communication (Rule 1.4)
In some matters, you must tell the client you are using AI—particularly when the use is material to the engagement, when the client’s confidential information is being processed by the tool, or when the client’s reasonable expectations require disclosure.
Fees (Rule 1.5)
If AI cut the work in half, the bill should reflect that. Charging hourly for work the AI did in seconds is a fee problem. Many firms are rebuilding flat-fee structures around AI-assisted work; others are charging a separate “AI license” or technology fee, which the ABA suggests is permissible if disclosed and reasonable.
Confidentiality (Rule 1.6)
Before inputting client confidences into any GAI tool, evaluate the tool’s data handling. The ABA expressly rejected boilerplate consent in engagement letters as a substitute for tool-by-tool informed consent when client confidences are involved. Self-learning models that train on user inputs require particular care.
Candor Toward the Tribunal (Rules 3.1, 3.3)
Every citation, every quotation, every authority your AI tool produces must be verified before it appears in a filing. There is no “the AI made me do it” defense to a Rule 11 motion.
Supervisory Responsibilities (Rules 5.1, 5.3)
Partners and managers must train their teams on appropriate AI use and put policies in place. The duty of supervision extends to non-lawyer staff and to the AI tool itself—treated, for these purposes, like a non-lawyer assistant.
Vendor Due Diligence Checklist
Vetting an AI vendor is a security review with extra steps. Six things to ask, every time:
Data Handling
Where does my input go? In which country, on whose infrastructure, and for how long? Confirm data residency if you have clients (especially European, healthcare, or government) who require it.
Is my input used to train models? The default consumer answer is yes; the default enterprise answer should be no. Get the no-training commitment in writing in the master agreement, not just in marketing copy.
How long is my data retained? Most enterprise AI vendors offer zero-retention or short-retention options. If your firm handles regulated data, demand them.
Who else can see it? List of subprocessors, support-staff access policies, and audit logs you can pull.
Encryption. In transit (TLS 1.2+) and at rest, with customer-managed keys available for sensitive deployments.
Security and AI Governance Certifications
Independent attestations are not a guarantee, but their absence is a red flag. The shortlist:
SOC 2 Type II—the table-stakes security audit. Ask for the current report, not just the badge.
ISO/IEC 27001—information security management. Common in mature vendors.
ISO/IEC 42001—the AI management system standard, published in late 2023 and now the emerging differentiator. Vendors who hold it have built governance, risk, and lifecycle processes specific to AI, not just generic infosec.
NIST AI Risk Management Framework alignment—a U.S. voluntary framework that most credible enterprise vendors map to.
HIPAA, CJIS, FedRAMP—if your practice touches healthcare, criminal-justice data, or government clients.
Output Grounding and Indemnification
Retrieval-augmented generation (RAG). For research and document-Q&A tools, does the tool retrieve from a defined corpus and cite its sources, or does it answer from the model’s training data alone? Cite-aware tools are not hallucination-free (see below), but they are the only ones you can spot-check efficiently.
Output ownership. You should own the outputs, with no claim back to the vendor.
Indemnification. Microsoft, Google, OpenAI, and Adobe all offer some form of copyright indemnity for outputs of their commercial models, conditioned on compliance with their use policies. Legal-tech vendors generally offer narrower indemnities. Read what is and is not covered, especially around hallucinated citations.
Contract Terms Worth Negotiating
No-training clause (specific to your data, not just “we will not train without consent”).
Data deletion on termination, with timeline and audit right.
Subprocessor change notice, so you find out when a new model provider enters the chain.
Security incident notification, with a hard deadline (24–72 hours is now common).
Service levels and uptime credits for tools that touch active matters.
Auditable model and prompt logs—the ability to look back at what the firm asked and what the tool answered.
Operational Fit
A tool that nails every box above but cannot integrate with NetDocuments, iManage, Clio, MyCase, or your core software of choice will sit likely unused. Before you commit, confirm the integration story, the per-seat versus per-use pricing model, the training and onboarding support, the vendor’s product roadmap, and their financial stability and ownership.
Jurisdictional and Cross-Border Considerations
If you have clients in the EU, the EU AI Act applies to your tools and your processes whether your firm is in the EU or not. The Act took effect in August 2024 with phased compliance dates; the main high-risk obligations land on August 2, 2026, with some extensions agreed in May 2026 (a 16-month postponement for new or substantially modified Annex III high-risk systems, and a 12-month postponement for AI that is a safety component of products governed by EU product-safety rules). “Administration of justice” is one of the Act’s high-risk categories.
Accuracy and Hallucinations: Still a Problem, Even in Legal-Specific Tools
The single most-cited Stanford RegLab paper on legal AI (Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, May 2024, peer-reviewed in the Journal of Empirical Legal Studies in 2025) found that the two leading cite-aware tools—Lexis+ AI and Westlaw AI-Assisted Research—hallucinated on 17% to 33% of 202 carefully constructed test queries. Both vendors had marketed those products with claims that effectively promised hallucination-free output. Both vendors have improved their products since the study, and the methodology generated meaningful debate, but the headline finding stands: cite-aware tools still hallucinate, and the verification burden never leaves the lawyer.
Dozens of federal and state judges have adopted standing orders requiring disclosure of AI use, certification of verification, or both. The orders are inconsistent across jurisdictions; check your judge’s standing order before every filing.
The operational lesson is simple: AI output is a draft, not an authority. Verify every citation against the actual reporter; verify every quotation against the source; verify every statute and rule against the current code. Make verification someone’s named responsibility, not a vague “the team will check.”
Run a Real Pilot Before You Sign
Vendors tell you their product is the right one. Your job is to test that claim against your work, with measurable success criteria, before the master agreement is signed. A workable pilot has five steps:
Define success first. What does “good enough” look like? Time saved, error rate, partner sign-off rate, billable-hour reduction, client satisfaction? Pick metrics and write them down before you start.
Choose representative tasks. Real matters, real documents, real research questions. Not the polished demo Q&A the vendor used in the sales meeting.
Pilot with a small group. Test the tool with two to four lawyers and one paralegal for 30 to 60 days, using structured weekly check-ins. Larger pilots tend to dilute feedback.
Red-team for failure modes. Try to break the tool the way an adversary would. Feed it jurisdictionally adjacent questions (state law it does not cover, prior versions of statutes, repealed authorities). Ask it for cases on a fact pattern with no good answer and see if it invents one. Have it summarize a document it has not been given. Have it cite “the leading case on X” in a niche subspecialty. Document every failure mode and how it presents.
Decide and document. Adopt, reject, or re-pilot with a different scope, as necessary. Whatever you decide, write down what you tested, what you found, and why you chose what you chose. That memo is your evidence of competence-level diligence if a malpractice carrier or bar regulator ever asks.
What Hasn’t Changed
Three-plus years of breakneck change have not altered the underlying point: AI is a powerful, error-prone assistant that does its best work when a competent lawyer supervises it. The tools that exist today are dramatically better than what existed when ChatGPT was a holiday-season novelty, and the tools that will exist in the future will be better still. None of that relieves you of the duty to verify, to maintain client confidentiality, to communicate honestly with the court, and to charge a reasonable fee for the work you actually did.
Pick tools that respect those duties. Vet them the way you would vet co-counsel. And do not outsource judgment.
Updated June 2026

