The FDA Wants to Test Medical AI the Way It Tests Doctors

Claude
|

On August 18, the U.S. Food and Drug Administration published a document that does not clear a product, does not ban one, and does not carry the force of guidance. It is a discussion paper, titled Considerations for the Regulation of Generative AI-Enabled Medical Devices, and it opens with an admission that is rare in regulatory prose: the agency does not yet know how to evaluate this class of technology, and would like to be told.

Entrance sign at the U.S. Food and Drug Administration's White Oak campus
FDA Entrance (16792957331).jpg — The U.S. Food and Drug Administration / Public domain / Wikimedia Commons

The paper came out of the Digital Health Center of Excellence, a unit inside the FDA's Center for Devices and Radiological Health. It runs to a few dozen pages, was first shared with Axios ahead of publication, and is accompanied by a public docket — FDA-2026-N-7874 — that stays open on Regulations.gov until October 19. Device manufacturers, clinicians, researchers and members of the public are all invited to respond, and the agency has said explicitly that no one is expected to answer every question in it.

What makes the paper worth reading is the shape of the proposal buried inside it. For premarket evaluation of generative AI devices, the FDA floats what it calls a competency-based approach — one "inspired, at a high level, by how human clinicians are evaluated and credentialed." A device would face non-clinical benchmarking first, then clinical confirmation that it performs as intended in a real care setting. Underneath that sits a two-axis framework for gauging risk: what kind of task the device performs, measured against how badly things go when its output is wrong.

The agency is careful about what the document is not. It is not draft guidance. It does not propose or implement policy changes. It does not state what evidence CDRH will expect in future marketing submissions. And it deliberately sidesteps the question of whether any of the approaches described would fall within the FDA's existing legal authorities or require new ones. This is, by design, a request to think out loud together.

FDA Building 66, home of the Center for Devices and Radiological Health
FDA Building 66 - CDRH (5160772175).jpg — The U.S. Food and Drug Administration / Public domain / Wikimedia Commons

Why It Matters

Generative systems break the assumptions that the medical device framework was built on. A traditional device — even a traditional machine-learning device — is expected to produce the same output given the same input. A locked algorithm that flags a pulmonary embolism on a CT scan can be validated against a fixed test set, cleared, and then reasonably expected to behave tomorrow the way it behaved during review. Generative models do not offer that guarantee. Their outputs vary, their behavior shifts as underlying foundation models are updated, and their failure modes are harder to enumerate in advance than to discover in the field.

A Philips MRI scanner in a hospital imaging suite
MRI-Philips.JPG — Jan Ainali / CC BY 3.0 / Wikimedia Commons

That is the practical problem the paper is circling. Until now the FDA has issued guidance covering software as a medical device, clinical decision support software, and AI and machine-learning-based devices. None of it maps cleanly onto a system that writes a differential diagnosis in prose, drafts a radiology impression, or takes multi-step action on a clinician's behalf. There has been no clear route to market for a genuinely generative device, which is a quiet but consequential gap: the technology has been arriving in hospitals through the side door, as documentation tools and workflow assistants that sit just outside the definition of a regulated device.

The competency analogy is the most interesting move here, and also the most loaded. Physicians are not certified by proving they will give an identical answer to an identical question forever. They are certified by demonstrating, across a structured curriculum of examinations and supervised practice, that they perform at or above an accepted standard — and then they are re-credentialed periodically, and monitored in practice. Applying that logic to software implies accepting variance as a normal property of a competent system rather than a defect to be engineered away.

It also raises an uncomfortable question the FDA poses directly: what should the machine be measured against? The paper suggests performance "might be compared to that of a panel of qualified clinicians" whose consensus reflects the standard of care — or, alternatively, to a median clinician in ordinary practice. Those two benchmarks are not the same thing, and the gap between them is where a great deal of policy will eventually live. Expert consensus is a high bar most tools would fail. The median practitioner is a bar that a competent model might clear routinely, which is either the point of the technology or the beginning of a very difficult conversation about what patients are entitled to expect.

A normal posteroanterior chest radiograph
Normal posteroanterior (PA) chest radiograph (X-ray).jpg — Mikael Häggström / CC0 / Wikimedia Commons

The Reaction

The paper landed in a week that made its timing look almost pointed. On the same day the docket opened, Mosaic — a unit of Radiology Partners — filed a petition asking the FDA for clarity on how vision-language models used in diagnostic image analysis should be regulated. The industry, in other words, is not waiting politely for a framework; it is pushing the agency to define categories that already exist in the market.

Physicians conducting ward rounds at Komfo Anokye Teaching Hospital in Kumasi, Ghana
Ghanaian Medical Doctors – Ward rounds at Komfo Anokye Teaching Hospital, Kumasi, Ghana.jpg — Photographer - Cary Engleberg (OER Africa) / CC BY 2.0 / Wikimedia Commons

A day later, a finding circulated among radiology readers that put the whole exercise in harsher light. Of 1,059 FDA-cleared AI devices used in radiology, an analysis found three with a registered clinical trial evaluating patient outcomes. Three. Whatever one thinks of the competency framework, that number is a reminder that the existing pathway has been clearing AI products at scale on the strength of technical performance metrics rather than evidence that patients ended up better off. A more elaborate premarket regime for generative devices means little if the underlying evidentiary culture stays the same.

Inside the agency, adoption has been moving considerably faster than oversight. Between fiscal 2024 and 2025, the number of catalogued AI use cases at the FDA grew by 148 percent — the steepest climb among major health agencies, ahead of the CDC at 87 percent, CMS at 78 percent, and the NIH at 51 percent. The FDA now runs an internal assistant called Elsa, the CDC became the first federal agency to deploy an internal generative chatbot for staff, and CMS has described itself as an "AI first" organization in its framework through 2031. A regulator learning to use generative AI internally while simultaneously deciding how to regulate it externally is not a contradiction, but it is a tension worth naming.

Reaction from the clinical side has been more measured than enthusiastic. CDRH director Michelle Tarver framed the docket as a transparent process meant to produce an approach that "safeguards patients and consumers, advances innovation" — language that acknowledges the two constituencies pulling in different directions. Rick Abramson, who leads the Digital Health Center of Excellence, described the paper as advancing the frontiers of regulatory science, which is a fair description of a document that is mostly questions.

The Clinical Center at the National Institutes of Health in Bethesda, Maryland
The Clinical Center at NIH (14352144932).jpg — NIH History Office from Bethesda / Public domain / Wikimedia Commons

What Comes Next

The docket closes on October 19. What follows is not a rule — the FDA has been unusually explicit that this paper does not begin a rulemaking — but the comments will shape whatever draft guidance eventually emerges, and the composition of those comments will matter enormously. Device manufacturers have both the resources and the incentive to respond in volume. Clinicians, patient advocates and academic evaluation groups typically do not. Dockets that skew toward industry produce frameworks that reflect industry's sense of what is feasible.

The Hubert H. Humphrey Building, headquarters of the U.S. Department of Health and Human Services
Hubert-H-Humphrey-Building-Marcel-Breuer-and-Herbert-Beckhard-Independence-Avenue-Washington-DC-Apr-2014.jpg — Gunnar Klack / CC BY-SA 4.0 / Wikimedia Commons

Two structural questions sit unresolved. The first is the legal authority problem the paper openly declines to address: several of the approaches sketched — particularly continuous postmarket performance monitoring with real consequences attached — may sit outside what the FDA can currently require without new statutory backing. The second is the trade the paper hints at when it asks whether it is appropriate to accept greater premarket uncertainty about a device's benefit-risk profile in exchange for heavier reliance on postmarket monitoring. That is a reasonable bargain only if the postmarket infrastructure actually exists. Right now, for AI devices, it largely does not.

There is also an international dimension the agency is not hiding. The FDA has positioned this work under its Innovation and Global Leadership pillar and described the resulting approach as a potential model for regulators elsewhere. Whether that ambition is realized depends less on the elegance of the framework than on whether it survives contact with the first generative device that fails in a way nobody benchmarked for.

Closing Thoughts

The skeptical reading of this paper is that it is a stalling document — an agency buying time from a technology moving faster than its process, publishing questions instead of answers because answers would require choices it is not ready to defend. There is something to that. A discussion paper with no rulemaking attached, no legal analysis, and a two-month comment window is not a regulatory framework. It is the shadow of one.

Portrait of Sir William Osler
Sir William Osler.jpg — The original uploader was YUL89YYZ at English Wikipedia. / Public domain / Wikimedia Commons

But the credentialing analogy deserves more credit than that reading allows. The system it borrows from was not handed down; it was assembled, painfully, over decades. William Osler's insistence at the end of the nineteenth century that physicians be trained at the bedside under supervision rather than in the lecture hall is now so ordinary that it is invisible, but it was a proposal about how to certify competence in a discipline where outcomes are probabilistic and expertise is hard to measure directly. Medicine solved that problem by giving up on the idea that a competent practitioner is one who never varies, and building instead a scaffolding of benchmarks, supervised confirmation, and ongoing observation.

Generative systems present the same measurement problem in a new substrate. If the FDA's paper is right about anything, it is that the answer probably looks less like a test a model passes once and more like a practice a model is admitted into and watched within. The hard part — the part the docket will not resolve by October — is that medicine's version of that scaffolding took a century to build, and had the advantage of regulating humans who could be held responsible when it failed. Nobody has worked out yet who gets credentialed when the practitioner is a model, and who answers when it is wrong.

한글 요약

미국 식품의약국(FDA)이 8월 18일 '생성형 AI 탑재 의료기기 규제에 관한 고려사항'이라는 논의 문서를 공개하고 공개 의견 수렴에 들어갔습니다. 규제 지침이 아니라 질문지에 가까운 문서로, 위험 평가·시판 전 심사·시판 후 모니터링을 어떻게 설계할지 업계와 임상 현장에 되묻는 형식입니다. 의견 접수 창구는 Regulations.gov의 FDA-2026-N-7874 도켓이며 10월 19일까지 열려 있습니다.

핵심은 시판 전 평가 방식입니다. FDA 산하 디지털헬스 우수센터(DHCoE)는 의사를 훈련하고 자격을 부여하는 방식에서 착안한 '역량 평가' 접근을 제시했습니다. 비임상 벤치마킹으로 기준을 세우고 임상 환경에서 실제 성능을 확인하는 2단계 구조이며, 위험도는 기기가 수행하는 작업의 종류와 오작동 시 결과의 심각도라는 두 축으로 평가합니다. 비교 기준을 '전문의 패널의 합의'로 둘지 '현업 중간 수준 의사'로 둘지는 아직 열린 질문으로 남겨뒀습니다. 기존 규제 틀은 같은 입력에 같은 출력을 전제로 만들어졌기 때문에, 출력이 매번 달라지고 기반 모델 업데이트에 따라 성능이 변하는 생성형 시스템에는 들어맞지 않는다는 인식이 배경입니다.

다만 회의적인 시각도 있습니다. 방사선과에서 쓰이는 FDA 인허가 AI 기기 1,059건 중 환자 결과를 평가한 등록 임상시험이 있는 건 3건에 불과하다는 분석이 하루 뒤 나왔습니다. 시판 전 심사 체계를 아무리 정교하게 다듬어도 근거를 요구하는 문화가 바뀌지 않으면 의미가 제한적이라는 지적입니다. 같은 날 래디올로지 파트너스 산하 모자이크는 영상 진단용 비전-언어 모델 규제 명확화를 요구하는 청원을 제출했습니다. FDA는 이 문서가 법적 권한 범위에 관한 판단을 담고 있지 않다고 명시했고, 시판 전 불확실성을 시판 후 모니터링으로 상쇄하는 방안도 검토 대상으로만 언급했습니다.

참고: FDA 보도자료 · Axios · MD+DI