If you ran the Three-Signals Test against your back office, you have a list. If you read the Q4 shortlist, you may have already picked a candidate. This piece covers the part nobody writes about: the 90 days between picking a workflow and knowing, with a number, whether the automation deserves to live.

Most credit union AI efforts fail in the middle, where the pilot has no kill date, no baseline, and no owner, and it drifts until the budget cycle quietly ends it. The fix is a plan with dates and exit criteria written down before anyone signs a vendor agreement. Here is the one we use.

Days 0 to 15: Scope one workflow and write the kill criteria

Pick exactly one candidate. Not two, not a platform decision, one workflow. Which one matters less than you would think; any candidate that passes the three signals will teach you what you need to learn.

This phase produces the artifacts below, and none of them require a vendor:

A baseline measurement. Count what the workflow costs today: items per week, minutes per item, error and rework rate, and current cycle time from arrival to resolution. Pull two or three months of history if your systems have it. If you cannot measure the workflow today, stop, because you will not be able to prove the pilot worked either.

Kill criteria, in writing. Decide now what result ends the pilot. Examples: accuracy below 95 percent on classification after tuning, cycle time improvement under 20 percent, staff overriding more than a third of outputs in week four of the parallel run. If you wait until the pilot is running to write these, they will get negotiated downward to match whatever the system is already doing.

A named owner. One person with the authority to stop the pilot and the obligation to report the numbers. A committee cannot do this job; someone’s name has to be on it.

Days 15 to 30: Compliance prework, before the contract

Skipping this phase is how pilots become exam findings. The good news: for a back-office workflow with human review, the requirements are manageable. We covered the regulator’s posture in detail in what NCUA expects before you deploy AI on member data; the pilot-scale version is short.

Third-party risk review. Run the vendor through your existing due-diligence process, the same one everything else goes through. Financial condition, SOC 2 or equivalent, breach history, subcontractor and model-provider disclosure.

A data map. One page: what member data leaves your environment, where it goes, whether the vendor trains models on it, and how deletion works at termination. Insist on the training answer in writing before you sign anything. Vendors who dodge it usually have a reason.

Model risk documentation, sized to the risk. For a sub-five-minute back-office task with human review, a short memo covering what the model does, its known failure modes, and the review control is proportionate. Do not import a bank holding company’s model risk framework for a document-routing pilot.

The examiner file. Open a folder on day 15 and put everything above in it. When the question comes on your next exam, and it will, you hand the examiner the folder.

Days 30 to 60: Build the pilot with review built in

Now the vendor work starts. Whatever the tooling category, some design rules hold:

Human review on every output, at first. Start with a person confirming every classification, every generated document, every routing decision. You will loosen this deliberately, with data, in the parallel run.

Log everything. Every input, every model output, every human override, with timestamps. The override log tells you where the model is weak, and it is the evidence base for the day-90 decision.

No new interfaces for members. First pilots stay internal. The moment a pilot touches member-facing channels, your compliance surface triples and your timeline goes with it. That is a second-project problem.

Expect integration against your core or imaging system to consume most of this phase. That is normal. If a vendor promises live-in-a-week against your core, ask which credit union on your core they did it for, then call that credit union.

Days 60 to 80: The parallel run

Run the automation alongside the existing process for three weeks minimum. Staff keep doing the work the old way; the system does it in parallel; the owner compares outputs weekly against the baseline from phase one.

Watch the agreement rate between the system and staff, the override rate where staff rejected the system’s output, and cycle time on the items where the system led. Publish the numbers internally each week. Most pilots that fail do so quietly, through neglect, and a Friday email with the week’s agreement and override rates makes that a lot harder.

The parallel run is also where you loosen review deliberately: if agreement holds above your threshold for two consecutive weeks on a category, move that category to spot-check review and note the date in the examiner file.

Days 80 to 90: Scale, fix, or kill

Hold the decision meeting on the calendar date you set on day zero, not when things feel ready. The pilot ends in one of three ways: you scale it, you fix what broke and rerun it, or you kill it.

Scale. The numbers cleared the criteria. Approve production, keep the logging, set a quarterly review, and pick pilot number two from your shortlist.

Fix. The numbers missed narrowly and the override log shows a specific, addressable weakness. One fix cycle is legitimate: 30 days, same criteria, same meeting. A second fix cycle usually means the pilot already failed and nobody wants to say so.

Kill. The numbers missed the criteria. Shut it down, write the one-page postmortem, and keep the artifacts; the baseline and the compliance file transfer directly to the next candidate. Killing in 90 days costs less than many institutions spend just deciding whether to start.

The outcome this plan exists to prevent is month seven of a pilot nobody can name a number for.

The pattern under the plan

The 90 days also leaves you with reusable machinery: a baseline method, a proportionate compliance file, a review-and-loosen pattern, and a decision meeting people take seriously. The second pilot runs faster because of it. The institutions in our case study review that made automation stick all had some version of this machinery; none of them started with a platform purchase. One honest limit: this plan assumes a back-office workflow with human review. A member-facing deployment or a credit decision tool needs a heavier compliance track than the day 15 to 30 phase described here.

More screening and scoping material lives in the back-office automation pillar.

If you want help pressure-testing your pilot plan, or you want a second set of eyes on the kill criteria before you commit, Advisor Labs runs a 45-minute back-office AI audit: book a conversation.

Get one substantive analysis like this every Tuesday: subscribe to the AiForCU newsletter.