Manjot Pal is the founder and CEO of Resonate AI. He holds multiple patents in ML & AI.
Healthcare leaders don’t need another impressive AI demo. They need a better way to determine whether a tool can survive real operations.
A 2025 report from MIT’s NANDA initiative found that only about 5% of task-specific enterprise GenAI tools in its sample reached successful implementation with sustained productivity or measurable financial impact. The number was directional, not a universal failure rate. Still, the underlying lesson matters: Most initiatives stalled because of brittle workflows, poor contextual learning and misalignment with day-to-day operations—not simply because the AI model wasn’t capable.
It’s tempting to conclude that AI isn’t ready. I think that’s the wrong lesson. Many pilots fail because the experiment doesn’t resemble the business it’s supposed to support.
Why A Smaller Pilot Is Not Always Safer
The pressure is highest in healthcare. The American Dental Association’s Health Policy Institute reported in 2026 that only 60% of surveyed dentists had adequate hygienist staffing. Among dentists recruiting hygienists, 91% said hiring was very or extremely challenging. Groups can’t simply hire their way out of every capacity problem, especially while expenses rise faster than reimbursement.
Multi-practice groups often choose their strongest location: an experienced manager, clean data, engaged employees and uncomplicated workflows. This reduces the chance that the pilot will fail. It also reduces the chance that the organization will learn what could break during rollout.
The U.S. Air Force learned a similar lesson decades ago. In a 1952 study, researcher Gilbert Daniels examined measurements from more than 4,000 pilots. When pilots were compared across 10 important physical dimensions, nobody fit the average across all 10. The answer wasn’t to calculate a better average but to make cockpits adjustable.
The average dental practice is just as imaginary. Practices inside one group can have different practice management systems, specialties, scheduling rules, insurance participation, call volumes, provider preferences and emergency protocols.
During a recent voice receptionist deployment across numerous dental practices in the same group, each location required different call routing, ring times, after-hours escalation and training. Neither practice was operating incorrectly. They were simply different. A pilot that treated one as representative of the other would’ve hidden the work required to scale.
Four Ways To De-Risk The Pilot
Start With The Decision, Not The Technology
Define what the organization needs to learn before selecting a location or vendor. For a voice receptionist, the objective might be to recover appointment demand from missed calls without increasing staff workload or compromising emergency escalation. Establish a baseline for missed calls, booking conversion, response time, staff touches and resulting production. Then define what would justify stopping, correcting, extending or expanding the pilot.
Pilot The Differences
Select two to four locations that expose meaningful operational variation: high and low volume, general and specialty care, centralized and local scheduling, different systems or different after-hours demand. Your best-run practice may be the safest place to test the product but the worst place to test whether it can scale.
Make Failure Visible Early
A pilot should intentionally test emergencies, cancellations, incomplete patient information, holiday hours, insurance questions, human handoffs and unsupported requests. Safety and compliance should be gates, not weighted scores. Strong ROI can’t offset mishandled patient information or an unsafe escalation. The pilot should also have a simple rollback path if those gates aren’t met.
Measure Value And Rollout Readiness Separately
Business measures may include incremental appointments, after-hours capture, request-resolution time and staff capacity returned. Rollout measures should include time to launch, configuration effort, training burden, manual exceptions, integration depth and support tickets per location. A pilot can show positive ROI while also proving that the vendor is too difficult to deploy across 100 practices.
This is also where vendor evaluation becomes practical. Product demos increasingly look similar. What separates vendors is what happens after the agreement is signed. Can the system adapt to each practice without becoming a custom software project? How quickly does the vendor respond when a workflow differs from the demo? Can results be reported by location rather than hidden inside an enterprise average?
The Goal Is Better Evidence
De-risking doesn’t mean making a pilot too small or controlled to fail. De-risking means making the pilot realistic enough to produce evidence that leaders can trust.
Sometimes, the right outcome is expansion. Sometimes, it’s a correction, another test or a decision to stop. Finding a material weakness before a 100-location rollout is exactly what the pilot was supposed to do.
The MIT figure shouldn’t stop healthcare organizations from experimenting with AI. It should change how those experiments are designed. Start narrow, but don’t make the environment artificially simple. The goal is to know whether AI can improve patient access and operations across every practice patients depend on.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?








