Surgical AI has moved fast. In just a few years, tools that once lived in research papers are now sitting inside operating room software, imaging platforms, and pre-op planning systems. The pitch is simple: better predictions, fewer complications, faster decisions. But the reality inside the OR is more complicated than the marketing suggests, and that gap is exactly where clinical judgment still earns its keep.
Andrew Ting MD has spent years watching this gap play out from both sides. His current work involves helping AI companies and startups train and refine models for the medical space, which means he sees firsthand where these tools shine and where they quietly fall apart. That dual vantage point, part clinician and part AI evaluator, is what makes his perspective on this topic worth paying attention to.

The Problem With “Validated”
Ask almost any AI vendor if their tool is validated, and the answer will be yes. But validated against what? A model trained and tested on a few thousand cases from three academic medical centers may perform beautifully in that narrow world and struggle the moment it meets a patient population it has never seen. Body types, comorbidities, prior surgeries, regional differences in disease presentation, all of these can shift outcomes in ways a model’s training data never captured.
Andrew Ting has pointed out that “validated” is often a marketing word before it is a statistical one. A tool can post strong accuracy numbers on a retrospective dataset and still misfire on the kind of messy, atypical case that shows up in a real operating room on a Tuesday afternoon. That is not a reason to dismiss the technology. It is a reason to ask harder questions about what the validation actually covered.
Where Models Tend to Fail
AI tools tend to struggle most when they encounter cases that differ from the data they were trained on. Unusual anatomical features, patients with several overlapping conditions, or scans that are incomplete or unclear can all create problems. In these situations, a model may suggest a surgical approach that appears reasonable based on the available data while missing something an experienced clinician would notice, such as scar tissue from an earlier procedure or an anatomical variation that is not fully visible on the imaging.
This is one area where Dr Andrew Ting’s work with AI developers can be especially valuable. He helps teams distinguish between a model’s statistical confidence and the level of confidence a clinician would place in a recommendation. Even when a system reports a 95 percent confidence score, its conclusion can still be wrong in a clinically significant way. In surgery, that distinction matters because an incorrect recommendation can have direct consequences for the patient.
A Practical Checklist Before Trusting a Recommendation
Clinicians evaluating a new AI tool, or deciding whether to lean on one already integrated into their workflow, should be asking a few consistent questions:
- What population was this trained on? If the demographics, case mix, or geographic scope don’t resemble your patient base, treat the output with caution.
- What does the model do with uncertainty? Some tools flag low-confidence predictions clearly. Others bury that signal, which can create false reassurance.
- How was this tested outside its original dataset? External validation across different hospitals or regions is a much stronger signal than internal testing alone.
- What happens when the model is wrong? Understanding failure modes matters more than understanding success rates.
- Does the interface make it easy to override the recommendation? A tool that buries the option to disagree with it quietly nudges toward over-reliance.
None of these questions require a data science background. They require the same skepticism a good clinician already applies to a new drug, a new device, or a new surgical technique before adopting it into practice.
The Real Value of AI in Surgical Planning
This does not mean AI has no place in surgery. When used appropriately, it can help identify patterns that might be overlooked during a long shift, surface risk factors hidden across years of medical records, and make planning more efficient in straightforward cases. The concern is not the technology itself, but how much authority clinicians give its recommendations. A model’s output should support a decision, not replace the judgment required to make one.
Andrew Ting MD has described this as part of a broader change in clinical judgment. Doctors are no longer only interpreting scans and considering a patient’s history. They also have to understand how AI tools arrive at their recommendations, recognize when those recommendations deserve closer scrutiny, and know when relying on the technology would be inappropriate. Those decisions depend on experience and an understanding of both medicine and the limits of the systems being used. That intersection has become an important part of Andrew Ting’s work.
The future of AI in surgery is unlikely to depend on which system can advertise the highest accuracy rate. What matters more is how carefully those systems are used in practice. Clinicians still need to question recommendations, consider what the model may have missed, and make the final decision with the patient’s individual circumstances in mind. In surgery, those decisions can have consequences that last far beyond the operating room.