Diagram showing nonclinical testing, clinical confirmation, and postmarket monitoring
← Back
AI in Medicine

Testing Medical AI: Why the FDA Is Borrowing an Idea From Physician Training

August 24, 2026 · 7 min read

Quick takeaways

  • The FDA released a discussion paper in August 2026 rather than a final regulation.
  • One proposed approach combines benchmark-style competency testing with clinical confirmation.
  • The agency is also considering risk-based monitoring after a product reaches patients.
  • Foundation models and AI agents create new questions because one underlying model may support many medical tasks.

A conventional medical device is supposed to do a defined job in a predictable way. A generative model may summarize a chart in one product and suggest a diagnosis in another. Its response can also change with the prompt, the available records or a software update.

In August 2026, the FDA released a discussion paper asking how generative-AI medical devices should be evaluated. The agency is requesting public feedback on the draft, which shows how regulators are thinking about a technology that does not fit neatly into older device categories (FDA).

A competency test followed by clinical confirmation

The FDA describes a possible two-stage approach inspired at a high level by physician assessment. First, a device would face nonclinical benchmarking across relevant capabilities. Those could include medical knowledge, analysis, communication, safety behavior and performance across different populations.

Passing a benchmark would qualify the product for clinical confirmation. A model can answer isolated questions correctly and still fail when records are incomplete or clinicians use it under normal working conditions.

The paper also asks whether independent third parties should maintain hidden test datasets, conduct assessments or serve as expert adjudicators. Sequestered evaluations could make it harder for developers to tune a model specifically to a public test.

Risk depends on what the AI can do

A system that drafts discharge instructions does not carry the same immediate risk as one that recommends a medication dose. The FDA’s possible framework considers both the seriousness of the decision and how independently the AI acts.

The level of evidence would rise with the risk. A low-risk tool may only need strong benchmarking and user oversight. A system influencing diagnosis or treatment could require clinical evidence and more aggressive monitoring, especially if it takes actions across several steps.

Approval cannot be the last check

Generative models may drift as hospitals change their software, patient populations or data. Developers may update the underlying model while keeping the medical product’s interface nearly identical. The FDA paper therefore discusses postmarket monitoring based on risk.

That monitoring could include performance reports, incident review and checks for new failure patterns. The FDA will have to decide which changes require another regulatory submission. Correcting a spelling error has little effect on clinical performance. Replacing the foundation model could change every answer.

The foundation-model problem

Many medical products may be built on the same large foundation model. If that base model changes, dozens of tools could change with it. Responsibility is then spread across the foundation-model developer, the medical-device company and the hospital deploying it.

AI agents complicate the review because a system may retrieve records, choose tools and carry out a multi-step clinical task. Each action can look reasonable while the complete sequence leads to a bad result. Regulators may need to test the sequence as a whole.

My Thoughts

I like the basic idea of making medical AI demonstrate competence and then prove itself in clinical use. Knowledge tests are a fair first check, and a clinical trial covers what they cannot.

I am more worried about what happens after release. A doctor has a license and an employer. With generative AI, responsibility can disappear into a chain of companies and software updates. The final rules need to name who watches for a failure and who has to fix it.

The discussion paper does not settle who carries that responsibility. I will be watching whether the final rules assign it clearly before allowing medical AI more autonomy.

Sources & References

  1. U.S. Food and Drug Administration. “Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback.” August 2026.
  2. U.S. Food and Drug Administration. “FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices.” August 18, 2026.
  3. U.S. Food and Drug Administration. “Artificial Intelligence-Enabled Medical Devices.” 2026.

Who’s writing this?
I'm Jack Nassiri, a Loyola High School student who reads medical research because the claims are usually more interesting once you open the paper.

Read Bio →