AI creates an awkward buying problem for service firms. The demonstrations are compelling, the range of possible uses is vast and the consequences of a poor implementation may surface only after money and trust have been spent.
Our response is a deliberately constrained approach. We begin with a commercial or operational problem, test one bounded use case and put governance into the design. We also expect to say “not yet” when data, ownership or assurance is too weak for a responsible pilot.
This article sets out the commitments behind that approach. It is intended to give a prospective client something concrete to challenge us on, rather than another set of general claims about responsible AI.
Six principles that govern the work
1. Start with the problem
“How should we use AI?” is too broad to guide an investment. A better starting point is a workflow in which delay, error, repetition or constrained capacity is causing a measurable problem.
We map that workflow before recommending a model or platform. The apparent AI opportunity may sit beside the real bottleneck. Faster document analysis achieves little if approval still takes two weeks. Automated drafting may accelerate the first version while leaving review, permissions and source quality unresolved.
A useful problem statement names the affected people, present performance, desired change and constraints. It can be tested without assuming AI is the answer.
2. State limitations in operational terms
AI limitations become meaningful when connected to consequences. A fabricated citation in legal research, an unsupported statement in a proposal and an incorrect client classification carry different risks. Each needs a proportionate control.
We define foreseeable failure modes, who reviews the output, which sources can be used, what is logged and what happens when confidence is low. Human oversight needs a real task, enough information and authority to intervene. Adding “subject to human review” to a specification is insufficient.
3. Prove value before scaling
A firm-wide implementation plan built before a use case has met real data, users and controls is largely hypothetical. We prefer a focused pilot with a decision at the end.
The pilot should be large enough to expose integration, adoption and quality problems, yet bounded enough to stop safely. Success criteria are agreed before building. The scaling decision then rests on observed business performance, assurance evidence and total operating effort.
A stopped pilot can be a good outcome. It may show that the problem is too small, source data is unreliable or a simpler workflow change produces most of the value.
4. Build client capability
The people who will own an AI-enabled service need to understand it well enough to operate, challenge and change it. That includes the model's role, source data, evaluation method, review path, monitoring and incident response.
We therefore expect client product owners, subject specialists, technology colleagues and risk owners to participate during the work. Documentation and knowledge transfer happen throughout the engagement. Dependence on one consultancy, vendor or individual is itself an operational risk.
5. Design governance with the service
Governance affects architecture and workflow. Permissions determine what information a system may retrieve. Accountability affects where approval sits. Record-keeping shapes the technical design. These decisions cannot sensibly wait until go-live.
Before implementation, we seek clear answers to questions such as:
- Who owns the use case and its outcomes?
- Which data and sources are permitted?
- Where is review mandatory, and what must the reviewer see?
- How are outputs, decisions and changes recorded?
- How can a user contest or escalate a result?
- What triggers a pause, rollback or retirement?
- Which legal, regulatory, contractual and professional duties apply?
The answers should match the use case. A low-risk internal summarisation aid and a client-facing recommendation engine do not need identical controls.
6. Measure the business outcome
Technical measures matter for diagnosis and assurance. They are not the investment case.
We connect them to the outcome the firm is trying to improve: less rework, shorter cycle time, better retrieval, increased capacity or more consistent handling. We also measure displaced effort. A tool that saves an author ten minutes and adds twenty minutes of expert review has moved the cost rather than reduced it.
Quality needs a defined benchmark. “The output looks good” cannot support a scaling decision. A representative evaluation set, agreed acceptance criteria and documented exceptions provide a stronger basis.
What the process looks like
The exact shape depends on the risk and complexity of the use case. Most engagements move through five decisions.
Readiness and problem definition
We examine the workflow, data, systems, governance, skills and incentives relevant to the proposed use. This is a specific readiness profile, rather than an average maturity score for the whole organisation.
The output identifies prerequisites, evidence gaps and constraints. Where a foundation is missing, the recommendation may be to fix it first. Poorly structured source content, unclear data rights or absent ownership will not be cured by adding a model.
Use-case choice
Candidate uses are assessed for commercial relevance, feasibility, risk and ability to learn. The best first use case is rarely the grandest one. It has an identifiable owner, suitable data, frequent enough demand and consequences that can be controlled.
This stage ends with a decision to pilot, prepare the foundations, use a simpler intervention or stop.
Pilot and assurance design
Before building, we agree the baseline, intended users, in-scope data, success thresholds, evaluation approach, human controls and stop conditions. We also specify how the pilot will coexist with the current service and how affected people will be supported.
Security, privacy, accessibility and regulatory review are part of the work where relevant. They are design inputs, rather than paperwork added after the technical team has finished.
Implementation and adoption
Implementation combines configuration or engineering with workflow change. Real users test representative tasks and edge cases. Subject specialists help evaluate quality. Owners learn how to inspect failures and change controlled elements of the service.
Adoption is treated as evidence. If capable users repeatedly avoid the tool, the response is to understand why. The cause may be poor quality, extra effort, weak incentives, inadequate training or a failure to fit the work.
Scale, change or stop
At the end of the pilot, results are reviewed against the agreed criteria. The choices are explicit: scale the use, change and retest it, retain it within a narrow scope, or stop.
Scaling brings new questions. More users, data and decisions can change the risk profile. Monitoring, support capacity, costs and vendor dependencies must therefore be reassessed rather than extrapolated from a small trial.
What we need from a client
Responsible delivery depends on participation from the firm. A senior sponsor must be able to make decisions and resolve access problems. A product or process owner needs to own the outcome. Subject specialists must be available to define quality and test results. Technology, information security, data protection, legal or compliance colleagues should join according to the use case.
We also need access to representative data and the current workflow, including its workarounds. A polished process map that omits how work is actually completed gives a pilot a false foundation.
Time commitments should be agreed before starting. An AI project run entirely at the edge of everyone's day will struggle to make sound decisions or transfer capability.
What we will decline
The principles become useful when they constrain behaviour. We will recommend a pause or a different intervention when a use case lacks a meaningful owner, lawful and suitable data, an adequate review path or a credible business outcome.
We will not treat a vendor demonstration as proof in a client's environment. We will not present generated output as reliable merely because it is fluent. We will not remove human accountability through vague language about automation. We will not conceal ongoing model, infrastructure, evaluation and support costs behind a pilot price.
These are practical boundaries, open to scrutiny. A prospective client should ask us to show where they appear in the scope, architecture, acceptance criteria and scaling decision.
The question now is whether the approach I've described here is the right fit for your firm's specific situation. If you want to find out, book an AI discovery conversation. It takes 45 minutes and ends with a clear view of whether and how we can help. No pitch. No pressure. Just an honest assessment of where you are and what makes sense next.



