Table of Contents
- Key Points
- Why This Research Matters
- Study Design: A Head-to-Head Comparison
- Who Took Part
- The 20 Clinical Cases
- What Doctors and the AI Were Asked to Do
- How GPT-4 Received Its Instructions
- The Primary Scoring Tool: Revised-IDEA (R-IDEA)
- Checking That the Scoring Tool Was Fair and Consistent
- Secondary Measures: Reasoning Quality, Diagnostic Accuracy, and Cannot-Miss Diagnoses
- Keeping the Scoring Objective and Unbiased
- How the Data Were Analyzed
- Study Limitations
- What This Means for Patients
- Frequently Asked Questions
- Source Information
Key Points
- The study compared GPT-4's written clinical reasoning with that of resident and attending physicians using the same 20 four-stage patient cases.
- Trained evaluators scored responses with the 10-point Revised-IDEA tool, blinded to whether GPT-4 or a human wrote each answer.
- Physicians came from two academic medical centers in Boston, so results may not represent community doctors or other regions.
- Cases came from a single educational platform using virtual patients, not real patient encounters.
- GPT-4 was tested on one fixed date with a single prompt design; different prompts or newer models could produce different results.
Why This Research Matters
Clinical reasoning is the thought process doctors use to figure out what is wrong with a patient. It involves weighing symptoms, risk factors, test results, and time course to reach the most likely diagnosis while keeping dangerous alternatives in mind.
Generative AI models like GPT-4 (the artificial intelligence system behind ChatGPT, created by OpenAI) are already being tested in health care settings. Patients and doctors alike need to know whether these tools can reason about medical cases the way skilled clinicians do.
This document is the online supplement to a research letter published in JAMA Internal Medicine. It provides the detailed methods behind the comparison, including survey instructions, the exact AI prompt used, and the full scoring rubric.
Understanding these methods matters for patients. It shows how rigorously AI must be tested before we can trust it with medical information — and what standards were applied when putting GPT-4 side by side with experienced physicians.
Study Design: A Head-to-Head Comparison
The study set up a direct comparison between human physicians and GPT-4. Both groups received identical clinical case material, one section at a time.
Each case unfolded in four stages. At every stage, respondents had to write a one-sentence summary of the case, called a problem representation, and a prioritized differential diagnosis (an ordered list of possible conditions) with justification. The AI was given exactly the same instructions as the doctors.
Trained clinician evaluators then scored the reasoning in each written response. The evaluators did not know whether a response came from GPT-4, an attending physician, or a resident physician in training.
The study followed the STROBE reporting guideline, a widely used checklist that ensures observational research is reported transparently and completely.
Who Took Part
Participants came from two academic medical centers in Boston, Massachusetts: Beth Israel Deaconess Medical Center and Massachusetts General Hospital.
The researchers recruited two groups of clinicians from the internal medicine departments at both hospitals:
- Residents — physicians in training after completing their first year of residency
- Attending physicians — fully trained faculty members practicing hospital medicine, primary care, or palliative care
Any physician whose training level was beyond the first year of residency was eligible to participate. No demographic information was collected beyond each resident's post-graduate year or, for attending physicians, the number of years since residency graduation.
The study was determined to be IRB exempt at both institutions, meaning it did not require full institutional review board approval because it involved no patients and posed minimal risk.
The 20 Clinical Cases
The researchers selected 20 clinical cases from NEJM Healer, an educational platform that uses realistic virtual patients to assess clinical reasoning.
NEJM Healer tests several reasoning competencies, including:
- Generating a problem representation
- Building a differential diagnosis
- Illness script instantiation (matching the patient's present findings to established patterns of disease)
- Management reasoning (deciding what to do next)
Each case was written by a practicing clinician in the relevant field and then edited by at least five additional physicians. All cases contained confirmed final diagnoses. Importantly, prior research has shown that performance on these cases correlates with clinical experience — meaning more experienced doctors tend to score higher.
The cases covered common problems seen in outpatient and acute care settings: pharyngitis (sore throat), headache, abdominal pain, cough, dyspnea (shortness of breath), chest pain, and arthralgia (joint pain).
Each case had four sections, with new information revealed at each stage:
- Triage Presentation — the initial reason the patient sought care
- Review of Systems — a full inventory of symptoms
- Physical Examination — the findings from a physical exam
- Diagnostic Testing — laboratory and imaging results
What Doctors and the AI Were Asked to Do
The survey was built in Qualtrics, a widely used online survey platform. Physician participants were told they were internal medicine clinicians expert at clinical reasoning, caring for the patient in the case at hand.
For each of the four case sections, doctors were instructed to provide two written items:
- A problem representation — one sentence summarizing the most important elements of the case so far
- A prioritized differential diagnosis with justification — a ranked list of possible diagnoses and the reasoning behind each
The instructions explicitly asked physicians to document their thinking just as they would in a real health care setting. That allowed the study team to evaluate genuine clinical reasoning documentation rather than artificially polished answers.
Survey drafts were refined through cognitive interviewing and pilot testing before the study launched. Each physician participant received the survey by email, along with one randomly selected clinical case.
How GPT-4 Received Its Instructions
Prompt engineering is the practice of designing instructions that get the best performance out of an AI model. Two of the study authors developed the GPT-4 prompt following OpenAI's official best-practice principles.
Critically, the prompt contained the identical instructions given to the human physicians. The only difference was technical formatting: the prompt told GPT-4 to press Enter after each section and used delimiters to separate the case text.
To avoid giving the AI an unfair advantage, the researchers used a zero-shot approach. That means GPT-4 received no worked examples of a good problem representation or differential diagnosis before starting. A "few-shot" approach, which provides sample answers for reference, would have favored the AI over humans.
Data collection took place on August 17 and 18, 2023. The prompt and Section 1 of each case were entered into a fresh chatbot session with GPT-4. The remaining case sections were then entered one at a time into the same session, and all responses were saved for scoring.
The Primary Scoring Tool: Revised-IDEA (R-IDEA)
The study's primary outcome was the Revised-IDEA score, known as R-IDEA. This tool was developed by Schaye and colleagues and evaluates clinical reasoning documentation in admission notes — the written records doctors produce when a patient is admitted to the hospital.
R-IDEA is a 10-point scale that assesses four core domains. The four domains and their point values are:
- Interpretive Summary (I) — 0 to 4 points
- Differential Diagnosis (D) — 0 to 2 points
- Explanation of the Lead Diagnosis (E) — 0 to 2 points
- Alternative Diagnosis Explained (A) — 0 to 2 points
For the interpretive summary, scorers looked for four specific features:
- Key risk factors from the history
- The chief complaint (the main reason for seeking care)
- The illness time course (how symptoms developed over time)
- Use of semantic qualifiers or unified medical concepts — for example, describing pain as "monoarticular" (affecting one joint) versus "polyarticular" (affecting many joints), or identifying "volume overload" or "cardiovascular risk factors" as unifying ideas
For the differential diagnosis domain, the scoring was:
- 0 points — no differential offered at all
- 1 point — differential is only implied, or is stated but only implicitly prioritized
- 2 points — differential is explicitly stated and explicitly prioritized
Scorers also judged how well respondents explained their lead diagnosis and their alternative diagnoses. A respondent earned points by linking specific objective data points from the case to the reasoning — not by vague statements. If a data point was not clearly tied to a diagnosis, it did not count.
For the explanation domains, scores depended on how many objective data points were used:
- 0 points — no explanation or no data points
- 1 point — one objective data point used
- 2 points — more than two objective data points used
The total R-IDEA score is the sum of all four domains, ranging from 0 to 10. The tool has been shown to evaluate reasoning across many different patient presentations, to score consistently with good interrater reliability (agreement between different raters), to predict real-world educational outcomes, and to reflect the natural progression from novice to expert.
Checking That the Scoring Tool Was Fair and Consistent
Before using R-IDEA in the main study, the researchers checked that it would work reliably in their own group of respondents.
Three authors experienced in evaluating clinical reasoning independently scored 29 case-section responses collected from 8 physicians who were not part of the main study. Each response was a written problem representation and differential diagnosis for one section of a case.
The three raters showed substantial agreement. The average Cohen's weighted kappa was 0.61 across the three paired combinations of scorers.
Cohen's weighted kappa is a statistical measure of agreement in which 0 means no better than chance and 1 means perfect agreement. A value of 0.61 is generally interpreted as substantial agreement — strong enough to trust that the scoring was consistent.
Secondary Measures: Reasoning Quality, Diagnostic Accuracy, and Cannot-Miss Diagnoses
The study evaluated three additional outcomes beyond the R-IDEA score.
Evidence of correct or incorrect clinical reasoning. The three clinician evaluators read each response and determined whether it contained one or more examples of correct or incorrect reasoning. This technique has been used previously in studies evaluating large language models in medicine.
Diagnostic accuracy. Accuracy was scored using a method developed by Chatterjee and colleagues. It is based on where the correct diagnosis appeared in the respondent's differential diagnosis list. The formula is:
Accuracy = 1 − (position of correct diagnosis − 1) ÷ (total number of diagnoses listed)
In plain terms:
- If the correct diagnosis was listed first, accuracy was 1, or 100%
- If the correct diagnosis was listed third out of 10 total diagnoses, accuracy was 0.8, or 80%
Cannot-miss diagnoses. Three physicians independently identified the "cannot-miss" diagnoses for the first section of each case. A cannot-miss diagnosis is a condition that, given the patient's presenting symptoms, poses an imminent threat to life or limb and absolutely must be considered in the differential.
If a diagnosis appeared on at least two of the three physicians' lists, it was included in the final list of cannot-miss diagnoses. The researchers then measured what proportion of these cannot-miss diagnoses each respondent included in their differential for the first case section. Two case sections had no identified cannot-miss diagnoses, so they were excluded from this analysis.
Keeping the Scoring Objective and Unbiased
Bias protection was a central feature of the study design.
For every case-section response, the scorers were blinded to whether the respondent was GPT-4, an attending physician, or a resident. The order in which responses were presented to scorers was also randomized, so no pattern could influence judgment.
The three evaluators independently scored two-thirds of the cases. In the event of disagreement on any response, the two scorers involved discussed the response and came to an agreement.
To make blinding work, a research team member who was not part of the scoring team edited all text outputs — both physician and GPT-4 — for spelling and grammar. This removed any telltale stylistic clues about whether a human or an AI had written the response.
How the Data Were Analyzed
The statistical plan was designed to compare GPT-4, residents, and attending physicians fairly while accounting for the fact that each respondent answered multiple case sections.
Descriptive statistics were calculated as median (interquartile range) for continuous outcomes and as frequency count (percentage) for binary outcomes. In the single instance where two attending physicians provided responses for the same case, all outcome scores for each case-section were averaged before analysis.
For the primary analysis, R-IDEA scores were divided into two categories: "low" (scores 0 to 7) and "high" (scores 8 to 10). Prior published literature defined scores below 5 as indicating low-quality clinical reasoning documentation. However, the minimum GPT-4 score in this study's dataset was 7. To stay aligned with the established literature while accommodating the actual data, the researchers adopted the nearest threshold of below 7 to define low scores. That choice allowed all three respondent groups to be represented in both the low and high categories.
The association between respondent type and R-IDEA score category was evaluated using a univariable logistic regression model with a random effect to account for the fact that the same participant answered multiple sections. The predicted probability of achieving a high score was calculated for each group, and standard errors on those probabilities were approximated using the Delta method, a standard statistical technique.
Several additional analyses were performed:
- Raw R-IDEA scores were compared among groups using Wilcoxon signed-rank tests, which allow pairwise score comparisons by case. The four case-section scores for each participant were averaged before analysis.
- Diagnostic accuracy was treated as a binary variable: low (below 75%) versus high (75 to 100%), analyzed with a univariable logistic regression model with random effects.
- Correct and incorrect clinical reasoning were each analyzed with univariable logistic regression models with random effects. A variance of zero in one respondent group prevented model convergence for the correct-reasoning evaluation, so a linear mixed model with random effect by case was substituted in that instance.
- Differences between respondent groups on cannot-miss diagnoses were evaluated using paired t-tests for each pair of respondents.
All reported P values were two-tailed, and P values of .05 or less were considered statistically significant. Analyses were performed using R version 4.0.3 (R Foundation for Statistical Computing).
Study Limitations
The supplement itself does not report the main study's results, so it cannot address all limitations of the findings. However, the methods reveal several important boundaries worth noting.
All physicians came from two academic medical centers in Boston, so their performance may not represent community physicians or those in other regions. Similarly, all cases came from a single educational platform, NEJM Healer, which uses virtual patients rather than real patient encounters.
The AI was tested under tightly controlled conditions on a fixed date with a single prompt design. Different prompts, different AI models, or newer versions of GPT-4 could produce different results.
The zero-shot approach was deliberately chosen to avoid favoring the AI. But real-world use of AI by doctors often involves more interactive prompting, so the study may not capture how GPT-4 would perform in actual practice.
Scoring relied on written documentation of reasoning. Doctors who communicate reasoning less clearly in writing, even when their thinking is sound, could receive lower scores.
What This Means for Patients
Patients are increasingly encountering AI in health care, whether through symptom checkers, patient portals, or news headlines. This study shows that AI systems can now be held to the same testing standards as human clinicians.
The R-IDEA tool that ranks a one-sentence problem summary, the prioritized differential list, and the reasoning behind each diagnosis reflects what good doctors actually do. When patients read a well-written note, they are seeing the product of these skills.
The study's careful blinding — so evaluators did not know whether they were reading a human or an AI response — matters for patients too. It means the comparison was designed to be fair, not designed to make either side look better.
If GPT-4's reasoning documentation proves comparable to or better than physicians' in the full study results, that does not mean AI will replace doctors. It suggests AI could serve as a decision support tool, helping clinicians catch diagnoses they might otherwise miss.
Patients should understand that studies like this are early steps. The real question for patient care is not whether AI can write a good differential on a training case, but whether it improves real outcomes — fewer missed diagnoses, fewer delays, and better communication at the bedside.
Frequently Asked Questions
What was this research actually testing?
It compared the written clinical reasoning of GPT-4, the AI behind ChatGPT, with that of human physicians. Both received the same 20 patient cases in four stages and wrote a one-sentence problem summary plus a prioritized list of possible diagnoses with reasons. Trained evaluators scored the reasoning without knowing who wrote each response.
Who took part in the comparison?
Physicians were recruited from internal medicine departments at two academic medical centers in Boston. They included residents, who are doctors in training after their first year, and attending physicians, who are fully trained faculty. Any physician beyond the first year of residency could join. No other personal details were collected.
What were the 20 clinical cases about?
The cases came from an educational platform using realistic virtual patients. They covered common outpatient and acute problems: sore throat, headache, abdominal pain, cough, shortness of breath, chest pain, and joint pain. Each case had a confirmed final diagnosis and unfolded in four stages, with new information revealed at each stage.
How was the reasoning scored?
Scorers used a 10-point tool called Revised-IDEA. It awards up to 4 points for the interpretive summary, 2 for the differential diagnosis, 2 for explaining the lead diagnosis, and 2 for explaining alternative diagnoses. Points for explanations required linking specific objective data from the case, not vague statements.
How did you make sure the scoring was fair?
Evaluators did not know whether a response came from GPT-4, an attending physician, or a resident. The order of responses was randomized. All text was edited for spelling and grammar to remove style clues. Three evaluators independently scored two-thirds of cases, and disagreements were resolved by discussion.
What does a diagnostic accuracy score of 80% mean?
Accuracy was based on where the correct diagnosis appeared in the ranked list. If it was listed first, accuracy was 100%. If it was third out of ten diagnoses, accuracy was 80%. The formula subtracts the diagnosis position from one and divides by the total number of diagnoses listed.
What is a cannot-miss diagnosis?
A cannot-miss diagnosis is a condition that, given the patient's symptoms, poses an imminent threat to life or limb and absolutely must be considered. Three physicians independently identified these for the first section of each case. A diagnosis was included if at least two of the three physicians listed it.
If I want a second opinion on how my doctor reasoned through my diagnosis, when should I ask for one?
Clinical reasoning is the thought process doctors use to weigh symptoms, risk factors, test results, and time course to reach the most likely diagnosis while keeping dangerous alternatives in mind. A second opinion is worth considering when you want the problem representation, the prioritized differential diagnosis, and the justification behind each possibility reviewed independently. The R-IDEA framework scores exactly these elements: an interpretive summary, an explicitly prioritized differential, and explanations linking objective data points to the lead and alternative diagnoses. A reviewer can check whether cannot-miss diagnoses were considered. Diagnostic Detectives Network provides independent expert second opinions.
Source Information
This patient-friendly article is based on peer-reviewed research published as supplementary online content accompanying the following original article:
Original title: "Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians" — Supplement
Authors: Cabral S, Restrepo D, Kanjee Z, et al.
Journal: JAMA Internal Medicine. Published online April 1, 2024.
DOI: 10.1001/jamainternmed.2024.0295
The full set of supplementary materials includes the survey instructions given to physicians (eTable 1), the GPT-4 prompt (eTable 2), the Revised-IDEA assessment tool by Schaye and colleagues (eTable 3), the complete methods (eMethods), and the reference list (eReferences), including prior work by Schaye et al., Abdulnour et al., Singhal et al., and Chatterjee et al.
© 2024 American Medical Association. All rights reserved. This patient summary was created to help non-specialist readers understand the research methods. The original peer-reviewed article remains the authoritative source.