The most expensive AI deployment I've ever watched at a customer of ours was not a deployment that failed.
It was a deployment that succeeded fast.
A vendor, not us, shipped a generative AI assistant into a hospital network in 21 days. The demo was perfect. The procurement team was thrilled by the speed. The press release wrote itself.
Six months later, that customer had a compliance incident, three rounds of architecture rework, a re-papered contract with their model provider, and a junior CIO who quietly took a job at another health system.
The vendor wasn't bad. Their tech wasn't bad. They had just compressed a 90-day deployment into 21.
After 18 months of shipping AI into hospitals, law firms, and insurance carriers, I've learned, the slow way, that the first 90 days of a regulated-industry AI deployment determine the next three years.
Here's the playbook we now run on every one.
The decisions that lock in a three-year deployment are mostly made before the first agent reaches production. Race through the first 90 days and you spend the next 18 months retrofitting under deal pressure.
Phase 1 - Days 1–14: Eval co-authoring
What happens? The customer's compliance, operations, and most-experienced clinical, legal, or underwriting staff sit down with our engineering team and write the eval suite together. Not after the pilot. Not after the first model is deployed. Together, in week one.
The output. A signed, versioned eval suite that defines what the AI will and will not do, what it will refuse, what it will escalate, and what the customer considers an acceptable answer in each major workflow.
Why this goes first. Because if you skip the eval co-authoring, every decision you make in the next 11 weeks is based on what you assume the customer wants, not what they will sign off on. We have watched competing vendors spend month 3 rebuilding everything they built in months 1 and 2, because they got the eval requirements at the wrong time.
The failure mode if you skip it. You ship something the engineering team approved but the customer's compliance team doesn't trust. The deployment "works." It just doesn't get past the first audit conversation. Or the renewal conversation.
Phase 2 - Days 15–30: Policy layer review
What happens? We sit down with the customer's General Counsel, Chief Information Security Officer, and any external compliance counsel. We walk them through every part of the policy layer, identity, access controls, classification labels, redaction rules, retention windows, data residency, audit logging, before any agent runs against real data.
The output. A signed policy specification, written in plain English that the GC can read and a regulator can audit. This document becomes the appendix the contract references.
Why this goes second. Because the policy and the eval together define the entire surface area of risk. If you ship an agent before the policy is signed, you have a default-on stack of unreviewed access patterns. Most vendors don't realize this until a regulator asks the customer to produce documentation.
The failure mode if you skip it. Six months from now, the customer's compliance team will find an access pattern they never approved, and the relationship will get rewritten in a meeting nobody on your team is invited to.
Phase 3 - Days 31–60: Limited rollout with human review on every decision
What happens? The agent goes live in production, but every single decision it makes routes through a named human reviewer on the customer's side. Not a sample. Not just the high-stakes ones. Every. Single. Decision.
The output. Two artifacts. First, a stack of reviewed decisions that becomes the training data for the customer's eventual confidence thresholds. Second, a calibrated escalation interface that's been stress-tested against real workflows, not synthetic ones.
Why this matter? The escalation interface I wrote about a few issues ago doesn't get built in a quiet engineering room. It gets built by watching what real reviewers actually do with real borderline cases. By day 60, we know which decision types the agent can handle confidently and which always need a second pair of eyes, and the customer has watched us learn it. That shared learning is worth more than every benchmark on the public scorecards combined.
The failure mode if you skip it. Your "autonomous" deployment turns out to be wrong on the exact edge cases that matter most. The customer finds out three weeks after they stopped reviewing, usually from one of their own customers.
The first 60 days aren't a rollout. They're a contract. Every decision the agent makes, every escalation it logs, every refusal it produces - those become the artifacts the customer signs on day 90.
Phase 4 - Days 61–90: Scale-up and sign-off
What happens? The customer reduces human review from every decision to a sampled audit. The eval suite gets one final review and is signed by the customer's CRO or compliance lead. The audit-replay flow is run end-to-end against a real historical decision, in front of the customer's auditor, with the regulator's expected schema as the output.
The output. Four artifacts, four signatures: a signed eval suite, a signed policy specification, a tested escalation interface, and a working audit-replay packet. That is the deployment.
Why this is the moment. By day 90, the customer's contract is referencing artifacts, not features. The model can change. The vendor can change engineering teams. The agent can be upgraded. The artifacts are durable. That is what a three-year enterprise AI deployment actually looks like.
The failure mode if you skip it. You are a 12-month customer pretending to be a 3-year customer. The renewal conversation at month 12 is going to surprise you in ways you cannot un-surprise.
The Builder's Takeaways
1. Write the eval suite with the customer in week one, not week six. The eval is the contract. If you don't co-author it early, you'll co-author it under deal pressure later, which costs five times more and trains the customer to distrust your process. Make the first kickoff meeting an eval workshop, not a demo.
2. Sign the policy before the agent runs against real data. Identity, redaction, retention, residency, audit logging, all of it. Get the General Counsel's signature on a policy document in plain English. The policy doc becomes the contract appendix. The customer who signs the policy is the customer who renews.
3. Treat days 31–60 as a stress test of your escalation interface, not a rollout. Every decision the agent makes routes through a human reviewer. What you learn in that phase is what makes the deployment defensible for the next three years. Don't reduce human review until the customer reduces it. Their reduction is the signal that the architecture works.
Final statement
There is a version of enterprise AI deployment that takes 21 days and ends in a compliance incident.
And there is a version that takes 90 days and ends in a three-year contract.
The teams that ship the first version look fast for one quarter and slow for the rest of the deployment. The teams that ship the second version look slow for one quarter and compound for years.
Most of the decisions that matter in a three-year deployment are made before day 60.
The first 60 days aren't a rollout. They're a contract.
