Quality assurance, compliance, and drift
Quality assurance has run on production volume for two consecutive quarters with documented actions.
Programs do not fail at launch. They decay around month eighteen.
The champion changes roles. The sampling schedule slips a quarter and then two. Documentation compliance falls quietly among the providers who were never enthusiastic. Nothing breaks visibly, and eighteen months later the program is a collection of devices with a policy nobody enforces.
Quality assurance exists to serve program safety, feedback, education, and assessment of practice patterns. It is not a compliance function and treating it as one guarantees that clinicians experience it as surveillance.
Sampling design
The commonly cited standard is review of roughly three percent of studies with a minimum of five studies per provider per year. That standard was designed for programs performing hundreds of studies annually and it breaks in both directions at enterprise scale.
At high volume, three percent becomes a reviewer burden nobody budgeted. At URMC's nursing bladder scanning volume alone, three percent of 70,000 studies over six months is 2,100 reviews, roughly seventeen per working day, for an application where the diagnostic stakes are modest and the marginal yield of review is low. At low volume, five studies per provider per year is too thin a sample to detect a competency problem before it produces harm.
Design sampling on risk rather than on a flat percentage. Weight by application consequence, by operator experience, by whether the finding drove a management decision, and by new-operator status. New operators get dense early sampling that thins as competency is demonstrated. High-consequence applications get sustained sampling regardless of operator tenure. Low-consequence high-volume applications get statistical sampling sufficient to detect drift rather than per-study review.
State the design, the rationale, and the review capacity it requires, and fund it. An unfunded quality plan is a quality plan that stops.
What gets reviewed
Separate the two questions, because they have different failure modes and different remedies.
Image quality and protocol adherence. Were the required views obtained, were they technically adequate, was the acquisition protocol followed. This is a training and technique question and the remedy is education.
Interpretation accuracy. Did the interpretation match the images, was the clinical conclusion supported. This is a judgment question and the remedy is different.
Add a third layer for programs at scale. Randomized over-read of a defined share of studies, blind re-read to measure inter-reader variability, and correlation against definitive imaging or clinical outcome for high-consequence findings. Correlation is the only method that detects the failure mode nobody else catches, which is a program that is internally consistent and consistently wrong.
Two AI-specific checks
Two failure modes are specific to AI-assisted programs and belong in the sampling design from the start.
For AI-guided acquisition by expanded operators, sample early and densely, measure the share of studies a credentialed reviewer deems adequate, and treat a rising inadequate rate as a retraining trigger. The tool raises a novice's floor. It does not hold it there without feedback.
For AI post-processing and automated quantification, track how often the interpreting clinician overrides the algorithm's measurement or inference. A high override rate is not a failure of the clinician. It is a signal about the tool in your population and your applications, and it belongs in both the vendor conversation and the decision to keep, retune, or retire the capability.
Feedback that changes behavior
Review that does not produce a documented action is data collection. Define what happens at each finding level, from individual feedback through targeted re-education to privilege modification, and document that the action occurred. The escalation path has to exist before it is needed, and it has to be approved by medical staff services rather than invented by the quality committee under pressure.
Compliance monitoring
Track storage compliance, meaning studies performed against studies archived, by department and by operator. Track documentation completeness against the elements required for a defensible claim. Track privileging currency. Report all three to governance on a fixed cadence and to department leadership by name.
The established provider problem deserves direct treatment. Clinicians who adopted POCUS before the program existed have workflows that predate every requirement, and they are frequently the most skilled and most respected users in the building. URMC names this explicitly as an ongoing challenge three years into deployment. Their finding is that physician champions and embedded fellowship-trained POCUS physicians moved compliance where policy alone did not. Peer influence works here. Compliance memos do not.
Drift
Build a named annual review of the program against its operating question, with a standing agenda covering volume trends, capture rate, denial patterns, quality findings, privileging currency, device inventory and age, and departmental participation. Assign the review to the governance body rather than to the director, so that the person accountable for the program is not the only person looking at whether it is still working.
Succession planning for the director and for department champions belongs in this review. The most predictable cause of program decay is a role change that nobody planned for.
In the practice setting
A practice cannot staff internal over-read and should not pretend otherwise. The workable model is a defined external review arrangement, whether through the vendor, a contracted reviewer, or a specialist relationship the practice already has, with a stated sample, a stated turnaround, and documented findings and actions.
Two things make this fail. Review that exists in a contract and never happens, and review that happens with no record of what was found or what changed as a result. Both are discovered at the same moment, which is when somebody asks.
Tie the sample to the pathway map. The applications where a missed finding changes management are the applications that get reviewed. Applications supporting a routine task with an available confirmatory path can be sampled thinly and honestly, provided the reasoning is written down.
What to require from your vendor
Quality analytics you own, exportable, and usable without the vendor's continued participation. Benchmark data from their installed base, with a clear statement of what the comparison population is.
No dependency on the vendor to operate your own quality program. If review workflow lives entirely inside a vendor platform, your quality function is a subscription. Ask what happens to historical quality data at contract end and get the answer in writing.
Failure mode
Quality assurance designed for a pilot is applied to an enterprise. Reviewer capacity is exceeded within two quarters, sampling silently stops, and the program discovers during an audit that it has no documented quality process covering the period in question.
Gate criteria
This workstream reaches steady state when quality assurance has run on production volume for two consecutive quarters with documented findings and documented actions, storage and documentation compliance are reported by department, and the review capacity is funded rather than absorbed. Quality is continuous, so this gate marks steady state rather than a stop.