Pilot
Designed so that it can tell you to stop.
A pilot that cannot produce a negative result is a purchase with extra steps. This one has a readiness gate that can end it in week one, six stopping rules, and a primary measure chosen before any data is collected.
Phases
Three phases, each with an exit gate
- 00
Phase Zero, readiness
Before anything runs. Fixed fee starting at $7,500.
Retrain and validate on your own history. A model trained on this repository's synthetic generator says nothing about your clinic until it has been retrained on your data. It runs on de-identified operational data: per-session counts by appointment type, complexity flags and realized technician minutes. No patient identifiers leave your clinic.
Exit gate
The model must beat both the fixed-ratio rule and the historical mean on error, with interval coverage near 80 percent. If it does not, the pilot stops. That the fixed ratio is adequate for you is a legitimate finding.
- 01
Phase One, shadow mode
4 weeks
Schedulers see forecasts and recommended plans. The fixed-ratio process stays the plan of record, so behaviour does not change and no patient is exposed.
Exit gate
Forecast accuracy measured prospectively, and scheduler trust. Not outcomes, because nothing has changed yet.
- 02
Phase Two, site-paired comparison
8 to 12 weeks
Pair sites on volume, subspecialty mix and staffing profile. One site in each pair staffs from the forecast with scheduler override, the other keeps the fixed ratio.
Exit gate
Pairing is at site level because technicians are shared within a site-day, so finer randomisation contaminates the control.
Source docs/pilot-evaluation-design.md, Design
The gate that most vendors do not offer
Phase Zero can end the engagement. If a model retrained on your own history does not beat both your fixed ratio and a simple historical mean, there is no case for continuing, and finding that out costs you a few weeks rather than a year. Your fixed ratio being adequate is a real possible outcome. The finding is yours to keep either way.
Phase Zero
What Phase Zero costs and what it touches
A fixed fee starting at $7,500
Phase Zero is a fixed fee starting at $7,500. The range above the floor depends on how many sites and how much history you bring. If the forecast does not beat the ratio you use today, the engagement stops there and the finding is yours to keep.
De-identified operational data only
Phase Zero runs on de-identified operational data: per-session counts by appointment type, complexity flags and realized technician minutes. No patient identifiers leave your clinic. Production use with protected health information is a later roadmap item, under a business associate agreement, a security review and the compliance work that goes with it.
Measures
Defined before they are collected
| Measure | Operational definition | Direction |
|---|---|---|
| Provider idle time | Provider minutes in session with no patient ready. | Lower |
| Patient cycle time | Mean arrival-to-departure minutes per session. | Lower |
| Overtime hours | Paid technician minutes beyond scheduled shift end. | Lower |
| Technician workload balance | Standard deviation of assigned minutes across technicians per week. | Lower |
| Reassignment frequency | Technician moves after publication on the service day. | Lower |
| Delayed testing | Ordered tests not completed because no qualified technician was free. | Lower |
| Throughput | Patients completed per session and site-day. | Higher |
| Competency rule compliance | Share of assignments covering every required competency. | Floor of 100 percent, not an improvement measure |
Source docs/pilot-evaluation-design.md, Success measures
The recommended primary measure
Provider idle time per session: minutes in session with no patient ready. It is the measure most directly caused by staffing alignment, it is cheaper to capture than cycle time, and it is less confounded by front-desk variation. You designate it, or another, before anything is collected.
Two measures that will be under-powered
Delayed testing and reassignment frequency are low-count events. They are reported descriptively with intervals rather than significance tested, because a test on a handful of events invites a conclusion the data cannot carry. Patient and staff satisfaction need an instrument you already own, and the cadence is yours to set.
Stopping rules
Any one of these stops the pilot and reverts to the fixed ratio
- 01
Any competency, availability, site eligibility, weekly hours or supervision violation reaching the floor.
- 02
Overtime at an intervention site materially above its paired control for two consecutive weeks.
- 03
A patient safety or care-delay concern raised by clinical staff and attributed to staffing, acted on before any statistical confirmation.
- 04
Forecast error falling below the fixed-ratio baseline over a rolling two-week window.
- 05
A fairness flag confirmed on investigation as a real disparity not explained by the competency rules.
- 06
A scheduler override rate so high that the recommendation is not in use, against a threshold you set before starting.
Source docs/pilot-evaluation-design.md, Stopping rules
Expansion
Agreed before the pilot begins, not argued afterwards
- 01
Zero hard-constraint violations across the pilot.
- 02
The primary measure improved against paired controls by at least the minimum effect set beforehand.
- 03
No secondary measure moved materially the wrong way. A cycle-time gain bought with overtime or worse workload balance is not success.
- 04
Staff satisfaction did not decline, and schedulers would keep using it.
- 05
The fairness report shows no confirmed disparity introduced or worsened.
- 06
Ongoing cost, priced from pilot experience rather than forecast, is justified by the improvement.
What happens when the result is mixed
If the first rule fails, the pilot stops permanently and the constraint layer is fixed before anything else is discussed. If the primary measure fails while the rest hold, your fixed ratio is close enough at those sites, and the sensible move is higher variance sites rather than expansion. If the primary measure holds but a secondary one, staff satisfaction, or fairness fails, the improvement cost more than the pilot was built to accept. Where results are positive but inside the noise, the answer is a time-boxed continuation at the same sites, not a decision.
Source docs/pilot-evaluation-design.md, Decision rule for expansion
Sizing
A planning estimate, not a calculation
A site-day is the unit of analysis. A pilot site running roughly ten provider sessions a day across four paired sites gives on the order of 400 site-days over ten weeks, which is enough for a moderate effect on a continuous session-level measure and not enough for rare events.
The required duration is computed in Phase Zero from the variance actually observed in your data. Until then, eight to twelve weeks is a planning estimate and is labelled as one. The minimum meaningful effect is set by your leadership beforehand, not read off the results: a smaller reduction in idle time is not a success even if it is statistically significant.
Source docs/pilot-evaluation-design.md, Sample size reasoning
