A law firm can make useful progress on AI within a quarter without promising a production deployment. It can select a bounded use, establish professional and technical controls, test performance on representative material and make an evidenced decision to proceed, revise or stop.

That definition of progress matters. A deadline should not encourage the firm to expose client information or rely on unverified output. Nor should regulatory uncertainty become a reason to leave unapproved use invisible.

The current SRA warning notice on misuse of AI, published in August 2026, says firms and solicitors remain responsible for meeting professional standards when using AI. It highlights inaccurate information and client confidentiality, including the safeguards needed before client information enters an AI system. Read the full current notice and obtain legal, compliance, privacy, security and professional-indemnity advice suited to the proposed use.

The three options below are candidate investigations, not universal recommendations for every 300-person firm. Size does not determine readiness. Data, work type, client terms, technology, capability and risk do.

We have companion versions for mid-market financial services firms and management consultancies.

1. Evaluate AI-assisted document review on a controlled benchmark

Document review is a broad category. Tools may classify documents, extract clauses, compare terms, summarise material or identify items for human attention. The firm should select one task whose correct result can be defined and checked.

For example, a corporate team might evaluate extraction of a defined set of clauses from safely prepared historical agreements. The goal is not to let the system decide legal significance. It is to establish whether the tool can support a bounded first pass without reducing the quality and control of the review.

Prepare the benchmark

Create a representative set of documents whose use is lawful and permitted, with client confidentiality, privilege, retention and supplier access addressed. Synthetic or appropriately prepared material may reduce exposure, though specialists should confirm whether it is suitable.

Have qualified lawyers establish the reference answers and note legitimate ambiguity. Include difficult examples, poor scans, missing clauses, unusual drafting and documents outside scope. A curated set of easy cases will overstate production performance.

Define the human role

State who reviews each output, what source evidence they see, how disagreement is recorded and whether any result can enter client work. Human review is meaningful only when the reviewer has competence, time and authority.

Protect reviewer capacity. One of the strongest source observations concerned a technically capable supervising associate repeatedly pulled back into live matters. The pilot drifted and produced inconclusive data. AI evaluation is delivery work, not an extracurricular task.

Measure quality and work

Use task-appropriate measures: missed relevant items, false flags, extraction accuracy, review time, correction effort and performance by document type. Examine whether reviewers become over-reliant or spend extra time verifying low-confidence output.

Any time released should be described carefully. A faster first pass does not automatically become billable capacity or a lower client fee. Decide how the operating and pricing model should respond.

The quarter should end with a decision record: suitable for a larger controlled test, suitable only for specified documents, currently below threshold, or inappropriate for the work.

2. Test permissions-aware search across internal knowledge

Law firms hold valuable precedents, guidance and work product that can be difficult to find. Natural-language search and retrieval can help, provided the system respects matter permissions, confidentiality, information barriers, retention and the status of the material.

The source article captured the human workaround well: people ask the colleague who has been at the firm for years and remembers where everything is. That colleague often adds context a search result lacks. A new tool should preserve or replace that context deliberately.

Choose a contained collection

Start with one practice area and an approved set of internal guidance or precedents. Avoid indexing the entire document estate before permissions, provenance and quality have been tested.

Review:

  • whether the material should be in the collection;
  • who can retrieve it;
  • document version and approved status;
  • matter and client restrictions;
  • how withdrawn or outdated material is removed;
  • what the supplier stores or uses; and
  • whether answers cite the underlying source.

Do not assume an existing platform integration preserves every permission correctly. Test with users who should and should not see specific material.

Evaluate retrieval and answer behaviour

Build realistic questions from fee earners. Measure whether the right sources appear, whether important sources are missed and whether generated summaries stay faithful to them. Include queries with no supported answer and require the system to respond safely.

Compare the tool with current search and human help. The aim may be faster discovery, wider reuse or reduced dependence on one person. Each requires different evidence.

Design adoption around real work

Training should use actual tasks and demonstrate limitations. Visible partner participation can help, though popularity must not substitute for quality. Provide a route to report a weak result and make the response visible.

Adoption alone is not success. A widely used search tool that retrieves out-of-date or unauthorised material creates greater risk at scale. Review use, answer quality, permission incidents and whether people still need to recreate work.

3. Prototype a supervised response service for narrow client questions

This is the highest-risk option of the three. It may be unsuitable for many firms or matters. The safe quarterly outcome may be a service design and offline evaluation rather than any client-facing use.

Choose a narrow category such as status information drawn from an approved system, access guidance or explanation of an administrative process. Avoid legal advice, interpretation or free-form answers based on uncontrolled material.

Start with the service boundary

Write down what the system may answer, the authoritative source, what it must never answer and how it escalates. Define the client groups, channels, authentication and accessibility needs.

The source proposed reviewing every response initially. That is a useful starting control and no guarantee of safety. Reviewers need source evidence and a clear accountability route. The team should assess whether drafting with AI saves any time once review, correction and record-keeping are included.

Build the controls before the interface

Address confidentiality, privilege, personal data, accuracy, supervision, client communication, records, complaints, supplier terms, security and incident response. Involve the firm’s relevant compliance roles and qualified advisers from the beginning.

Set an approved response library or retrieval source. Log inputs, sources, generated drafts, human changes and final communications where lawful and appropriate. Prevent unsupported output from being sent automatically.

Test offline and include refusal

Use representative questions, adversarial wording, ambiguous matters and attempts to obtain information outside permission. Measure supported accuracy, inappropriate answer rate, escalation, reviewer effort and accessibility.

An effective system should refuse or escalate many questions. A high escalation rate does not automatically mean failure if the boundary is intentionally narrow. The relevant question is whether the service safely resolves the approved tasks and improves the client journey.

Only consider a live test after the firm’s authorised decision-makers accept the evidence and controls. Tell clients what they need to know about the service and preserve access to a person.

Resource the work as a matter

The hardest constraint is often protected expert time. Each evaluation needs:

  • a business and professional owner;
  • legal, compliance, privacy and security input;
  • representative users;
  • technical and supplier support;
  • an evaluation lead independent enough to report failure;
  • approved data and test material; and
  • a dated decision forum.

Do not run all three because the title lists three. Rank them by material need, evidence availability, possible harm and existing control. One well-run evaluation is more valuable than three demonstrations.

Define success before testing and include a stop threshold. Report positive, negative and inconclusive findings. “We do not know yet” can lead to a better benchmark; it should not become indefinite pilot status.

How to build confidence in AI without overpromising covers the communication and trust problem. What using AI in your business actually means provides broader context.

Make the quarter end in a governance decision

By the final review, leadership should be able to see:

  • the use and people affected;
  • obligations and controls;
  • benchmark and limitations;
  • performance and harm evidence;
  • operating cost and required capability;
  • adoption or workflow implications; and
  • the recommended next state.

Our AI readiness scorecard for legal services examines data quality, technical environment, fee-earner readiness and compliance posture. Check that the tool and destination remain current before relying on them. Our AI governance framework provides a related structure.

Distinction offers an AI opportunity assessment and has a commercial interest in subsequent work. The assessment should be capable of recommending a different use case, prerequisite work or no deployment.

This quarter’s useful result is not being seen to move faster than another firm. It is one professional, evidence-based decision about where AI helps, where it fails and what the firm is prepared to operate responsibly.