Getting Claude to produce a good answer takes an afternoon. Getting Claude into production, acting on live systems, under audit, at the scale an enterprise actually runs at, is a different discipline entirely. Most organisations discover this the hard way: an impressive pilot, a stalled rollout, and a quiet return to the way things were before.
This is the method that gets past that point. Not a framework borrowed from generic software delivery with "AI" added to the slide, but the specific discipline of taking a reasoning system into a regulated, governed, European enterprise environment.
Why most Claude projects stall
The pattern repeats with enough consistency that it is worth naming precisely, because each failure mode has a distinct cause and a distinct fix.
The wrong use case gets chosen first. Pilots often start from what looks impressive rather than what is measurable: a general-purpose assistant, a broad "chat with our documents" tool, an ambition rather than a process. Nobody can tell whether it worked, because nobody defined what working meant. The fix is choosing use cases the way this playbook describes below: specific, owned, and measured in a number the business already tracks.
Governance arrives as an afterthought. Security and legal get looped in once the pilot already works, at which point every architectural decision that would have made their review straightforward has already been made the other way. The review becomes a negotiation instead of a walkthrough, and in regulated sectors, a negotiation an AI pilot usually loses. Governance designed from the first session costs less than governance retrofitted at the end, in every case we have seen.
Change management is underestimated. A working system meets a workforce that was never brought into its design, and the system gets routed around. Adoption is not a training session bolted onto go-live. It has to be designed alongside the architecture, with the people who will use the system daily involved before launch, not informed after it.
None of these is a failure of the technology. Claude's reasoning is not the constraint in the vast majority of stalled deployments. The constraint is that the surrounding discipline, use case selection, governance design, adoption planning, was treated as optional.
The four-stage method
Every implementation we run follows the same four stages, in the same order, because each stage exists to remove a specific reason the ones after it fail.
Identify. Find the use case where reasoning over your record changes a number the business already tracks. Not the use case that would make the best demo: the decision your organisation makes hundreds or thousands of times a week, with a clear owner, a measurable baseline, and errors that are recoverable rather than catastrophic. This stage typically takes days, not weeks, and it is the stage most commonly skipped or rushed. It should not be.
Prove. A guided pilot on your real data, not an export or a synthetic dataset, with success defined in business terms before a single configuration decision gets made. This is where the honest answer to "does this actually work for us" gets found, on your systems, under your constraints, in weeks rather than quarters.
Productionise. The pilot becomes an operation: permissions mapped to your existing access model, audit trails that satisfy your compliance function without a special request, escalation paths for the cases the system should not resolve alone, and the governance posture your regulatory environment requires, built in rather than bolted on. This is the stage where most implementations that reach this far still fail, usually because the first two stages were rushed and this one inherits the debt.
Scale. From one proven use case to a governed portfolio: performance monitoring, continuous optimisation, and a pipeline of the next candidates identified using the same discipline as the first. Value compounds here, or it does not happen at all, because a single successful pilot that never scales is not a transformation, it is an anecdote.
Governance is not the last chapter
For any organisation operating in the EU, governance is not a compliance appendix added at the end of the method above. It shapes the architecture from the Identify stage onward.
Two things matter here, and they are not the same thing. The GDPR has applied for years and does not pause for a new technology: where prompts, context and outputs live, what leaves the EU and under what safeguard, and whether your data protection function can actually answer those questions in writing. The EU AI Act is newer and more specific to AI systems, with obligations that scale to the risk category a given use case falls into. If your use case touches employment decisions, candidate assessment, credit, or several other categories the Act enumerates, the obligations are substantial, and the timeline for them has recently shifted, in ways worth reading carefully rather than assuming you already know. We cover exactly what changed, and what did not, in a dedicated piece on the recent deadline revision.
The practical implication for this playbook: classify your use case against these categories during the Identify stage, not after Productionise is already underway. It is a conversation that takes an afternoon at this point in the process. Left until later, it can take the project apart.
Choosing the right first use case
A short, honest checklist, because getting this one decision right changes everything downstream:
- High volume. A decision made rarely is hard to measure and slow to prove. A decision made constantly gives you a real baseline within days.
- Measurable outcome. You should be able to name the number this use case moves before you configure anything. If you cannot, the use case is not ready yet.
- Recoverable errors. The first production use case should be one where a mistake is inconvenient, not catastrophic. Save the harder, higher-stakes use cases for stage two of Scale, once the foundation and the trust are both established.
- A named owner. Someone in the business, not just in IT, whose job gets measurably easier or whose numbers measurably improve. Without this person, adoption has no advocate once the project team moves on.
Where this fits with Salesforce
If your record of truth is Salesforce, the Identify and Productionise stages above connect to a specific, well-trodden architecture: reasoning that acts on your platform data under your existing permission model, rather than on exports and copy-paste. We cover that connection in detail in our Claude and Salesforce integration piece, and it is where we most often see the fastest path from Identify to a genuinely provable pilot.
Start where you actually are
Bring a stalled pilot or a blank page, either is a reasonable place to begin. A senior consultant will give you an honest read on where your organisation sits against this method, and what the next stage would actually look like, within one working day.