Summary: A multicenter international study shows that neither experienced radiologists nor advanced multimodal large language models (LLMs) can consistently tell AI-generated “deepfake” X-rays from real images. Seventeen radiologists across six countries were tested using synthetic radiographs produced by ChatGPT (GPT-4o) and RoentGen.
Even when readers were alerted that synthetic images were included, average detection accuracy was only 75%. The results highlight a significant vulnerability in clinical imaging: fabricated X-rays could be used for fraud, forgeries in legal claims, or malicious manipulation of electronic medical records to disrupt care or create diagnostic errors.
Key Facts
- Deception rate: When not told to look for fakes, only 41% of radiologists spontaneously noticed anything suspicious in the AI-generated images.
- AI versus AI: The model used to create the deepfakes, GPT-4o, detected more of its own fabrications than other LLMs (GPT-5, Gemini 2.5 Pro, Llama 4 Maverick), but no model detected all synthetic images.
- Visual clues: Deepfake X-rays often appear unnaturally symmetric: overly smooth bone surfaces, abnormally straight spines, excessively clean-looking fractures, and uniform vessel or lung patterns.
- Experience not protective: Years of clinical experience (0–40 years) did not reliably predict better detection, although musculoskeletal radiologists achieved slightly higher accuracy than other subspecialists.
Source: RSNA
Overview: A study published in Radiology, the journal of the Radiological Society of North America (RSNA), evaluated how well radiologists and multimodal LLMs can distinguish real radiographs from AI-generated images. The study underscores the pressing need for safeguards, training, and technical solutions to preserve the integrity of medical imaging.

“Deepfake” describes images, videos, or audio that look authentic but were generated or altered using artificial intelligence. According to lead author Mickael Tordjman, M.D., postdoctoral fellow at Icahn School of Medicine at Mount Sinai, New York, the study shows that synthetic radiographs can convincingly mimic authentic clinical images and deceive even highly trained specialists.
Tordjman warned of real-world consequences: fabricated fractures could be presented in litigation, and an attacker who can insert synthetic images into hospital systems might change diagnoses or trigger unnecessary interventions. Such manipulations could undermine trust in digital medical records and endanger patient care.
Study details: Seventeen radiologists from 12 centers across six countries (United States, France, Germany, Turkey, United Kingdom, and United Arab Emirates) participated. Their experience ranged from novice to 40 years in practice. The dataset comprised 264 X-ray images: half authentic and half AI-generated. Readers evaluated two independent image sets with no overlap.
Phase 1 mixed authentic images with 77 ChatGPT-generated radiographs across multiple anatomical regions. In a second phase, after being told synthetic images were included, participants were asked to classify images as real or AI-generated. A third phase focused on chest X-rays: 55 authentic images and 55 generated by RoentGen, an open-source diffusion model developed by researchers at Stanford Medicine.
When not informed about synthetic content, only 41% of radiologists spontaneously flagged AI-generated images as suspicious. After disclosure, overall mean accuracy rose to 75% for distinguishing GPT-4o-generated images from authentic ones. Individual radiologist performance ranged from 58% to 92% on the ChatGPT-generated set.
Four multimodal LLMs were also tested (GPT-4o, GPT-5, Gemini 2.5 Pro, Llama 4 Maverick). Their accuracy ranged from 57% to 85% on the GPT-4o dataset and 52% to 89% on the RoentGen chest set. While GPT-4o and GPT-5 performed better than other models, none achieved perfect detection.
Common synthetic image features included excessive bilateral symmetry, uniform noise or grain, subtly odd soft-tissue textures, and overly smooth bone contours. Musculoskeletal specialists detected fakes more reliably, likely because of their focused experience with bone anatomy and fracture patterns.
Recommendations: To protect imaging integrity, the authors propose technical and procedural safeguards. Suggested measures include embedding invisible digital watermarks or cryptographic signatures into images at capture, and linking images to technologist credentials so any later tampering becomes detectable. Developing curated educational datasets and detection tools is also essential as AI advances toward generating synthetic 3D studies such as CT and MRI.
Key Questions Answered:
A: Motivations include financial fraud (e.g., falsified injuries for insurance claims) and malicious acts by attackers seeking to alter diagnoses or provoke unnecessary treatment. The clinical and legal consequences can be severe.
A: Yes. A synthetic image does not reflect the patient’s real condition. If a fake image masks disease or invents a false abnormality, the patient may miss necessary care or receive inappropriate interventions.
A: Recommended defenses include implementing cryptographic signatures and invisible watermarks at image acquisition, rigorous audit trails tying images to acquisition devices and technologists, and clinician training supported by curated detection datasets.
Editorial Notes:
- This article was edited by a Neuroscience News editor.
- The full journal paper was reviewed for this summary.
- Additional contextual information was provided by staff editors.
About this neurology and aging research news
Author: Linda Brooks
Source: RSNA
Contact: Linda Brooks – RSNA
Image: Image credited to Neuroscience News
Original Research: Open access. “The Rise of Deepfake Medical Imaging: Radiologists’ Diagnostic Accuracy in Detecting ChatGPT-generated Radiographs” by Mickael Tordjman et al., Radiology. DOI:10.1148/radiol.25209
Abstract
The Rise of Deepfake Medical Imaging: Radiologists’ Diagnostic Accuracy in Detecting ChatGPT-generated Radiographs
Background
Large language models and related generative AI tools can produce convincing synthetic medical images. These deepfakes raise concerns about misuse, diagnostic error, and data integrity in healthcare systems.
Purpose
To evaluate how well practicing radiologists and multimodal LLMs can distinguish synthetic radiographs produced by ChatGPT (GPT-4o) and RoentGen from authentic clinical images.
Materials and Methods
This retrospective diagnostic accuracy study (April–August 2025) included 17 radiologists from six countries. Phase 1 assessed 154 radiographs (77 ChatGPT-generated, 77 authentic) across multiple anatomic regions while readers were blinded to the study aim. Phase 2 asked readers—after disclosure—to classify images as synthetic or authentic. Four multimodal LLMs performed the same tasks. Phase 3 evaluated 110 chest radiographs (55 RoentGen-generated, 55 authentic). Statistical comparisons used McNemar and t tests.
Results
Only 41% of readers initially noticed synthetic images without being informed. After disclosure, mean accuracy was 75% for the GPT-4o dataset and 70% for the RoentGen chest dataset, with no statistically significant difference between them. No LLM detected all synthetic images; GPT-4o and GPT-5 outperformed Llama 4 Maverick and Gemini 2.5 Pro. Typical synthetic features included bilateral symmetry, uniform noise, unnatural soft-tissue texture, and overly smooth bones.
Conclusion
Deepfake radiographs produced by current LLM-based tools are difficult to distinguish from authentic images for both human experts and AI models. Training, detection tools, and secure imaging workflows—such as embedded cryptographic signatures and watermarking—are urgently needed. A curated deepfake dataset for education and tool development is available from the study authors.