A 200-person consulting firm invested in document summarisation and rolled it out to its consulting teams. Within six weeks, use had collapsed.

The tool functioned. The document collection did not. Current and outdated versions sat together in a SharePoint folder that had not been governed for years. Consultants tried the service, received unreliable summaries and went back to manual work.

The firm treated this as a technology failure. It was a data and knowledge-governance failure exposed by technology. The wasted budget mattered. The deeper cost was lost internal confidence: the next AI proposal arrived in a room that had learned to be sceptical.

Readiness work is meant to prevent that category error.

Readiness is a profile, not a pass mark

"We will know when we try" sounds practical. It is a little like saying you will discover whether a building is structurally sound during the renovation. Sometimes you will. By then, the money is committed and the foundations dictate the available choices.

AI readiness is not binary. It is also not a single maturity number.

Different workflows make different demands. Retrieving approved knowledge depends heavily on source quality, permission and content ownership. Assisting a professional decision adds evidence, oversight and accountability. A client-facing assistant adds service, security, accessibility and communication risks. Every use still needs governance and human capability; the emphasis changes.

Assess five dimensions independently:

  1. data and knowledge;
  2. systems and security;
  3. people and change;
  4. governance and accountability;
  5. skills and operation.

Then compare the profile with one defined use case. The purpose is to decide what the firm can test safely, which gap blocks it and whether AI is the right intervention at all.

How to score without fooling yourself

Use a 1-to-5 scale as an evidence prompt, rather than a scientific benchmark.

  • 1: Ad hoc. The capability is largely absent, uncontrolled or dependent on individuals.
  • 2: Recognised. The need is understood and some local practices exist.
  • 3: Defined. Ownership and a repeatable approach exist for part of the firm or use case.
  • 4: Operating. The approach is used consistently and produces reviewable evidence.
  • 5: Assured and improving. Outcomes, failures and change are monitored, with independent challenge where the risk warrants it.

Do not average the five dimensions. A 4 in data and a 2 in governance is different from two 3s. The gap may be a veto for one use case and manageable for another.

Score current evidence. A policy still in draft, a data-cleaning project not started or a training budget awaiting approval belongs in the plan, rather than today's score.

Dimension 1: data and knowledge

AI systems do not make a source authoritative by reading it. They can make weak information easier to retrieve and repeat.

Assess the specific sources a use case would rely on.

Quality and status

Are records accurate enough for the decision? Can the system distinguish current, draft, superseded and duplicate material? Who decides which precedent or policy is approved?

A high score needs evidence from sampling, validation and correction, rather than confidence from the system owner.

Ownership and lineage

Can somebody explain where an important field or document came from, who may change it and which downstream uses depend on it? Is there a route to correct an error at source?

Lineage matters when an output needs to cite or defend its evidence. A generated answer without traceable sources may be unsuitable even when it sounds right.

Access and permission

Can the application obtain only the information an authorised user may access? Existing repositories often contain permissions designed for human browsing, inherited access and informal sharing. Retrieval can expose those weaknesses at scale.

Test representative roles and edge cases. "The documents are in Microsoft 365" does not answer whether the intended application can use them lawfully and securely.

Coverage and representativeness

Does the source include the range of cases the tool will encounter, including exceptions and older clients? Missing data can produce systematically weaker service for some groups.

The readiness question is narrower than "is our data clean?" Ask whether this source is fit for this task and how the team will know when it is not.

Dimension 2: systems and security

Modern platforms may provide APIs and embedded AI features. That does not make every integration easy or every feature approved.

Assess the complete service, including identity, data flow and operation.

Integration capability

Can the necessary systems exchange data through supported interfaces? Are rate limits, failure, version changes and recovery understood? Who owns each connection?

Manual export can be acceptable for a short, low-risk evaluation. It is a fragile production design if nobody controls version, access or deletion.

Identity and access

Will the AI service respect user and matter-level permission? Can access be removed promptly? Are privileged actions separated and logged?

Test what a user must not see, as well as the expected result.

Supplier and model service

Which supplier processes prompts, outputs and uploaded information? Where can data be stored or reviewed? Is it used for training? Which settings, deployment types and contract terms change the answer?

Provider terms and features change. Keep an owner and review date rather than relying on a procurement statement made at pilot launch.

Reliability and observability

Can the team see errors, latency, source failures, cost and unusual use? Is there a safe fallback when the model or integration is unavailable? A demo can succeed without any of these production capabilities.

Existing tools may offer a sensible low-friction starting point. Evaluate them against the use case and risk instead of assuming a bundled feature is automatically cheaper or safer.

Dimension 3: people and change

"Culture" can become a vague explanation for any failed rollout. Assess observable conditions instead.

User-side value

Do the people expected to change understand how the new workflow helps them perform a real task? Has the team observed the current work and its exceptions?

If AI creates more checking, entry or uncertainty than the existing route, reluctance may be rational.

Leadership behaviour

Is a senior owner present for trade-offs, risk and resources? Do leaders follow the same use rules they expect from colleagues? Can the sponsor protect a bounded pilot from a stream of additions?

Sponsorship is behaviour, rather than a name on a slide.

Participation and challenge

Can users, subject experts, operations and risk colleagues change the design while decisions remain open? Is it safe to report a wrong output or say the use case lacks value?

Enthusiasm is not readiness. A team needs sceptics who can improve the test and users who represent ordinary working conditions.

Change capability

What happened in recent CRM, portal or workflow changes? Which adoption practices worked? Where did local incentives or management behaviour preserve the old route?

Past results are evidence, not destiny. Use them to design ownership, rollout, support and feedback.

Partnership firms may need broader influence and consensus than hierarchical organisations. That is a governance constraint to design for, rather than proof that partners resist change.

Dimension 4: governance and accountability

Governance should make responsible experimentation easier by clarifying boundaries and decisions.

Permitted use

Does practical guidance tell colleagues which tools and information are approved, what human review is required and where to take an uncertain case? Can the firm update and communicate the guidance as services change?

An unread policy is weak evidence. Test staff understanding and observe use.

Named accountability

Who owns the use case outcome, technical service, data, professional quality and incident response? A committee can challenge and coordinate. Individuals still need authority to act.

Risk classification and approval

Does the firm distinguish a low-consequence drafting aid from a client communication or decision affecting rights and money? Higher-impact uses require stronger evidence, oversight and authority.

The classification should examine the actual workflow and affected people. Calling something "internal" does not make it harmless when it shapes advice later.

Monitoring and response

Which failures, complaints, exceptions and outcome differences will be reviewed? Can the tool be restricted or stopped? Who tells clients, regulators or insurers where required?

Governance that ends at launch is approval, rather than control.

Regulatory and professional obligations

Legal, financial and other regulated services must map relevant duties to the use case. The companion article on AI governance covers this in more depth. If the firm cannot explain confidentiality, competence, fairness, record, oversight and supplier responsibilities, keep client-facing use outside scope until it can.

Dimension 5: skills and operation

The firm does not necessarily need machine-learning engineers. It needs enough capability to commission, evaluate, use and operate the chosen service.

Domain-led evaluation

Can subject experts create representative tests, recognise subtle errors and define acceptable performance? Generic AI benchmarks do not establish professional fitness.

Make evaluation part of the work, with time and ownership, rather than asking busy experts to glance at a few outputs.

Technical ownership

Can internal people or a durable partner configure integrations, manage releases, investigate failures and control cost? Does the firm retain enough knowledge to challenge and replace the partner?

AI literacy

Do users understand what the tool can and cannot do, which sources it uses, how to review output and when to refuse it? Prompt tips alone are insufficient.

Training should fit each role and use case. Someone approving client advice needs different knowledge from an administrator using a bounded extraction tool.

Service support

Who handles access, user questions, wrong outputs, source corrections and incidents? What happens outside the pilot team's availability? How are new colleagues enabled and departing colleagues removed?

Learning and improvement

Can the team use observed errors and user feedback to change the source, workflow, training or model? Is there a release and re-evaluation process?

A tool is not ready for operation because its first version works. The firm needs a way to keep it fit as sources, suppliers and work change.

Add a use-case gate

A readiness profile without a problem can turn into a generic improvement programme. Put one use case through a gate.

Value

Which person, task and outcome improve? What baseline exists? Could process simplification or conventional automation solve it more reliably?

Feasibility

Which sources, systems and skills are required? What is the smallest evaluation that could expose the biggest uncertainty?

Risk

Who could be harmed by error, omission, delay, bias or data exposure? Which professional or regulatory duties apply? Can a human realistically detect the important failures?

Operability

Who will own, support, monitor and fund the service after the pilot? Which supplier changes or usage costs could alter the case?

A firm can score strongly overall and still reject a poor use case. Another can have low general maturity and run a narrow, low-risk learning exercise with synthetic or non-client data. Readiness is contextual.

Interpret the pattern carefully

Avoid automatic recipes such as "strong data means automate the back office" or "weak data means start with meeting summaries". Back-office work can carry high consequence, and meeting tools can process confidential information.

Instead, identify the lowest dimension that materially constrains the chosen use case.

If source status is weak, curate a bounded set and assign ownership before retrieval. If governance is weak, define approved tools, data boundaries and accountability before client or confidential use. If cultural evidence is weak, research the workflow and test with a representative group. If technical operation is weak, choose a managed service or limit the exercise to evaluation until ownership exists.

Everything at 2 or 3 is common. Choose the gap that unlocks one valuable, proportionate test. Do not launch five workstreams to raise a maturity diagram.

Scores of 4 or 5 across the board still need evidence. A firm with strong process can overestimate model quality or choose a use case with no user value.

Counter self-assessment bias

The CTO may think data is cleaner than it is. The managing partner may overestimate adoption. Compliance may assess the policy rather than observed practice. None requires bad faith; people see the organisation from different positions.

Ask several roles to score independently and cite evidence. Compare gaps in the room. Sample data, observe a task, inspect a supplier setting and test a user journey. Record uncertain calls instead of forcing consensus.

The AI Readiness Assessment uses these five dimensions to produce a profile. Treat the result as a structured conversation, rather than an external benchmark or assurance opinion.

A facilitated session can help when leadership, technology, operations and risk need to reconcile their evidence. The value is often in the disagreement: it exposes an assumption that would otherwise arrive during delivery.

The useful question is not "are we ready for AI?" It is "what are we ready to test, which evidence supports that view and what must be true before we go further?"

That question produces a map instead of a badge.