Beyond the Breaking News

Advancing conversational diagnostic AI with multimodal reasoning

United States News News

Advancing conversational diagnostic AI with multimodal reasoning
United States Latest News,United States Headlines

Improvements in the Articulate Medical Intelligence Explorer, a large language model designed for diagnostic dialogue, enable the model to request, interpret and reason about multimodal medical data.

Real-world clinical practice is inherently multimodal, relying on the synthesis of patient history with visual information such as medical imagery and clinical documents. Although large language models have shown promise in diagnostic dialogue, their evaluation has been largely restricted to text-only interactions, failing to capture the complexity of modern remote care delivery.

Here we introduce a multimodal extension of the Articulate Medical Intelligence Explorer , capable of gathering, interpreting and reasoning about multimodal data within a diagnostic conversation. To achieve this, we developed a state-aware dialogue framework that dynamically guides history-taking based on diagnostic uncertainty and evolving patient states, emulating the structured reasoning of experienced clinicians.

We evaluated this updated, state-aware version of multimodal AMIE against primary care physicians in a randomized, blinded exploratory study comprising 105 simulated telehealth consultations, which included dermatology photographs, electrocardiograms and clinical documents. As assessed by 18 specialist physicians, multimodal AMIE outperformed PCPs not only in diagnostic accuracy but also in conversation quality, including history-taking and empathy. Specifically, multimodal AMIE demonstrated superior performance on 29 of 32 evaluation axes, including seven of nine metrics that assess multimodal reasoning.

These results validate the efficacy of state-aware reasoning in bridging the gap between text and visual information and demonstrate the potential for artificial intelligence systems to augment clinicians in complex, multimodal diagnostic settings. Healthcare delivery faces considerable challenges globally from aging populations and increasing care fragmentation through to clinician burnout and misalignment between clinical and financial incentivesRecently, we have witnessed the promise of generative AI systems in healthcare, especially those based on LLMs, as potential tools to augment healthcare delivery and address some of these challenges.

Research systems such as the AMIE, an LLM-based conversational diagnostic AI system, have demonstrated physician-like capabilities to steer text-based clinical conversations, attaining superior or similar performance compared to PCPs across a broad array of measurements for consultation quality in studies involving patient-actors. Although the evidence base for such autonomous diagnostic systems in real-world clinical settings remains limited, some conversational AI systems for triage and care navigation have shown promising results in specific applications.

For instance, Zeltzer et al.conducted an evaluation of an AI system deployed in a real-world virtual urgent care setting for common symptoms and showed that the initial AI recommendations were generally concordant or rated by expert adjudicators as better than the diagnosis and management recommendations from the treating physicians. Furthermore, studies on other conversational systems have also shown high patient acceptability, a crucial factor for successful adoption Despite the early promise of this technology, LLM-based medical AI systems have predominantly been studied and implemented as text-only chatbots.

This represents a substantial deviation from how multimodal information is routinely used in care. As demonstrated during the COVID-19 pandemic, it is integral for remote care participants to be able to converse about and interpret multimodal medical information. This can take the form of images such as patient-provided skin photographs, electrocardiogram tracings or laboratory reports, which can be shared via ubiquitous messaging platformsthat combine text and multimodal chat or even more commonly can occur via real-time video consultations.

We also note that these data provided by remote care participants are distinct from data obtained directly in the clinical environment, such as radiology imaging. Multimodal medical tests are essential in effective care and can meaningfully guide the course of a consultation.

However, evidence validating the use of LLMs for diagnostic conversations involving such multimodal data is scarce, revealing an important discrepancy between clinical needs and current technology. Patients may struggle to precisely articulate the results of an investigation using text alone, potentially omitting crucial details such as precise laboratory test values, and essential clinical information commonly resides only in non-textual formats.

For example, visual data such as photographs are paramount for remote assessment of dermatological conditions, and clinical documents, such as blood reports or specialist letters, contain objective findings that can be otherwise impractical to relay accurately via transcription to text chat. A text-only approach prevents AI from leveraging these rich, supplementary objective sources of information, hindering its ability to form a complete clinical picture.

An overreliance on text-only input not only increases the risk of diagnostic errors but also risks exacerbating disparities in access to telehealth, presenting barriers for individuals with lower digital literacy and language proficiency. By contrast, instant messaging apps that allow users to send text, voice and video messages are commonplace with billions of users.

Although real-time video call is more common in telemedicine, such multimedia instant messaging platforms enable considerably richer communication than text-only conversation and have seen reported use as tools for remote clinical consultations in which doctors and patients can exchange multimodal medical data The ability of LLMs to request, interpret and reason about multimodal medical data during clinical conversations has not been investigated or evaluated. To address this, we introduce multimodal AMIE—advancing the conversational diagnostic capabilities of the original systemby integrating multimodal medical perception.

We introduce a state-aware reasoning framework designed to orchestrate complex clinical conversation flows, which dynamically adapts responses based on intermediate model outputs, reflecting the evolving patient state and diagnostic uncertainty. This allows multimodal AMIE to strategically request, interpret and integrate multimodal artifacts, including skin images, ECGs and clinical documents, into the dialogue to refine diagnoses in a manner that emulates the diagnostic reasoning process of experienced clinicians in a telehealth setting.

To enable rapid iteration and validation of design choices, we developed a simulation environment for generating realistic patient scenarios and conducting automated turn-by-turn dialogue assessments. Finally, to evaluate the performance of our finalized system, we conducted a randomized, blinded human evaluation that emulates an objective structured clinical examination , a standardized assessment method used in both medical schools and physician licensing examinations worldwide.

We compared multimodal AMIE to PCPs in multimodal text-based chats and evaluated their behaviors on 105 carefully designed multimodal clinical scenarios with validated patient-actors. We note, however, that our study is not a randomized clinical trial with prespecified endpoints and preregistered statistical analysis. Rather, it is an exploratory study investigating the properties of multimodal diagnostic dialogue.

Furthermore, we designed a dedicated multimodal understanding and handling rubric to rigorously compare AI and human performance in artifact interpretation over numerous dimensions. FigureThis figure provides a schematic overview of the key components enabling and evaluating multimodal diagnostic conversations within the AMIE system, facilitated through a multimodal chat interface. , Multimodal state-aware reasoning. Multimodal AMIE employs a novel state-aware dialogue phase transition framework built on the publicly available Gemini 2.0 Flash model.

This dynamically controls the conversation flow through history-taking, diagnosis and management phases, guided by intermediate outputs reflecting patient state and diagnostic uncertainty, allowing strategic requests and interpretation of multimodal artifacts. , Simulation environment. A comprehensive simulation framework enables rapid development and automated evaluation. This framework involves generating realistic patient scenarios grounded in real images and metadata, using Gemini 2.0 Flash, simulating turn-by-turn multimodal dialogues between multimodal AMIE and patient agents, and using an auto-rater agent for assessment against clinical criteria.

, Randomized comparative study. The performance of multimodal AMIE was evaluated against PCPs in a randomized, double-blind, OSCE-style study. Trained patient-actors conducted synchronous chat consultations based on 105 diverse multimodal scenarios involving artifacts such as skin photographs, ECGs and clinical documents. , Evaluation results.

Specialist evaluations demonstrated that multimodal AMIE attains similar or superior performance to PCPs in handling and reasoning about multimodal data over multiple evaluation axes alongside strong performance in diagnostic accuracy and overall consultation quality. UI, user interface. To evaluate the clinical utility of multimodal AMIE, we conducted a randomized, blinded OSCE-style study comparing the system against board-certified PCPs in multimodal text chat consultations.

These findings indicate that, within the context of this randomized, blinded exploratory study, multimodal AMIE performance was similar to or higher than PCP performance across the assessed axes, as rated by specialists and patient-actors, including history-taking, diagnostic accuracy, management reasoning, communication skills, empathy and the understanding and handling of multimodal data. Subsequently, we report findings from our automated evaluation framework, using controlled ablation studies.

of multimodal AMIE was both more accurate and more comprehensive. DDx accuracy.

Both multimodal AMIE and PCPs have the opportunity to submit a DDx list for each dialogue. The top-accuracy was computed using an auto-rater that compares each of the 10 diagnoses with the ground truth. Multimodal AMIE outperformed PCPs in diagnostic accuracy on average image quality , non-multimodal specialist metrics and non-multimodal patient-centric metrics .

Each specialist and patient-actor separately rated two dialogues for the same scenario, and their relative ranking was derived from their scores. For the specialist metrics, each scenario was evaluated by three different expert physicians, yielding three pairwise rankings that were aggregated by taking the majority vote.

The ‘Tie’ label was assigned to cases where two or more specialists assigned the same ratings to both PCP and multimodal AMIE consultations or the cases for which the opinions on the relative performance were equally divided. Significant differences between multimodal AMIE and PCP segments, calculated using a one-sidedshow the underlying distributions of ratings for the same set of metrics.

For all panels, the error bars center on the mean and represent 95% confidence intervals estimated based on 10,000 bootstrap samples. . Multimodal AMIE’s leading diagnosis and broader DDx list more frequently contained the ground truth condition compared to the lists generated by PCPs. This overall difference in diagnostic accuracy was statistically significant .

We note that although we analyze accuracy up to top-10, neither multimodal AMIE nor PCPs always report 10 differential diagnoses. This work advances from prior comparison of multimodal AMIE in simulated consultations by incorporating the ability for patients to upload multimodal artifacts into text chat.

To understand the extent to which our results specifically reflected this new capability, and to explore the extent to which the comprehension and utilization of multimodal data contributed to the superior overall diagnostic accuracy of multimodal AMIE, we performed specific subgroup analyses using specialist physicians’ blinded consultation ratings , when independent specialists rated the provided artifact quality as low, the top-3 DDx accuracy decreased for both multimodal AMIE and PCPs as anticipated. However, multimodal AMIE demonstrated a significantly smaller performance drop than that of PCPs.

Second, regarding the impact of artifact use in reasoning , when specialists judged that the consulting agent appropriately used the visual artifact, diagnostic accuracy improved similarly for both. Finally, we examined consultations where specialists identified instances of the clinician or multimodal AMIE reporting findings not actually present in the artifact, which we refer to here as ‘hallucination’ ), ECG tracings and clinical documents , with additional effects estimated for experimental arm and modalities.

Both multimodal AMIE and PCPs were more accurate for scenarios based on clinical documents . Multimodal AMIE retained statistically significantly higher accuracy across all modalities. For the full statistical analysis details, see Supplementary Section, The panel shows the distributions of specialist physicians’ ratings and patient-actors’ ratings ; this analysis attempts to capture the overall user satisfaction of the consultation.

The distributions of ratings are displayed separately for the sessions with PCPs and multimodal AMIE and for modality types. , The distributions of scores for three MUH metrics are displayed. For illustration purposes, all responses from five-point rating scales were mapped to a generic five-point scale ranging from ‘Very favorable’ to ‘Very unfavorable’.

For Yes/No questions as indicated by ‘’, a ‘Yes’ response was mapped to the same color as ‘Favorable’ and a ‘No’ response to the same color as ‘Unfavorable’. The ‘NA’ rating indicates responses that are ‘not applicable’, ‘cannot rate’ or ‘does not apply’. The McNemar test was used for each evaluation axis, andWe compared multimodal AMIE and PCPs on the patient-centric quality metrics that were collected upon conclusion of the consultations.

Patient-actors rated the interactions with multimodal AMIE as similar to or higher than those with board-certified PCPs across the assessed dimensions in Fig.for full ratings). This included aspects such as being polite, listening, explaining conditions, involving patients in decisions, appearing honest/trustworthy, building rapport, showing empathy and managing patient concerns and Practical Assessment of Clinical Examination Skills criteria,).

The two questions, rated from 1 to 5, asked how well the doctor addressed the patients’ questions about image artifacts and how well the doctor explained the findings from the artifacts. FigureA group of 18 specialists in dermatology, cardiology and general practice evaluated multimodal AMIE and PCPs in terms of numerous attributes of effective diagnostic conversations, including multimodal reasoning, history-taking, diagnostic accuracy, management reasoning, communication skills and empathy. Aggregated results are presented in Fig.).

The conversations held by multimodal AMIE were consistently rated more highly by specialists overall and across three disciplines. This is reflected in all items in our assessment, across diagnosis and management and quality of history-taking.

The positive assessments of the diagnostic dialogues that multimodal AMIE conducted were also supported by results in the MUH component of our assessment: specialists assigned higher ratings on average to multimodal AMIE’s interpretation of the multimodal artifacts, its reasoning about these artifacts and the way it handled patient questions and concerns about the multimodal artifacts. This result is also reflected in all three disciplines studied.

Multimodal AMIE provided more appropriate diagnoses and management plans, and patients were more happy to return for a second visit with multimodal AMIE in our three fields . Similarly, in all three disciplines, multimodal AMIE performed more accurate interpretation of artifacts and reasoning about those artifacts, with no increase in hallucination or mis-reporting , although the superiority was statistically significant only for clinical documents, potentially due to the smaller sample size.

In Supplementary Section, we provide qualitative examples from the three disciplines, of dialogues between multimodal AMIE and a patient-actor and between a PCP and a patient-actor where multimodal AMIE has correctly identified the most probable diagnosis. For each discipline, we use the same patient scenario for the dialogue with multimodal AMIE and with the PCP, to facilitate a qualitative comparison.

Beyond the OSCE study, this section details automated evaluations of multimodal AMIE using simulated dialogues and auto-rating : Comprehensive patient profiles are created, including condition, demographics, symptoms and medical history, using information derived from web searches and real-world datasets—for example, PTB-XL and SCIN .

These profiles are then used to generate detailed patient scenarios, outlining the patient’s presentation and expectations. Step 2 : A doctor agent and a patient agent engage in a text-based consultation with multimodal artifact upload. The doctor agent is instructed to provide empathetic and clinically accurate responses while the patient agent responds truthfully based on the generated scenario.

Step 3 : An auto-rater agent evaluates the simulated dialogue based on predefined criteria, including management appropriateness, information-gathering effectiveness and the presence of hallucinations. Qualitative feedback is also provided by the auto-rater to explain the scores. Mx, management. To further understand the contributions of different system components to overall performance, we conducted several ablation studies using our automated evaluation framework.

This section details experiments comparing the full multimodal AMIE system against simpler baselines and evaluating the specific impact of test-time reasoning, history-taking and robustness to variations in patient presentation on conversation quality, as assessed by the auto-rater. One of the key elements of our system is the state-aware dialogue phase transition framework . To assess its contribution, we compared its diagnostic performance and dialogue quality against a baseline model lacking this explicit reasoning structure.

We hypothesized that the dynamic phase transitions and uncertainty-driven questioning employed by our framework would lead to more accurate diagnoses compared to a simpler approach. Using our simulation environment and auto-rater, we evaluated conversations generated for synthetic scenarios derived from PAD-UFES-20, SCIN, PTB-XL and ClinicalDoc-QA datasets.

We compared two conditions: multimodal AMIE with the full test-time reasoning framework and a ‘vanilla’ baseline using the same Gemini 2.0 Flash model but without the explicit state-aware reasoning framework, relying only on domain-specific instructions and lacking the dynamic phase transitions and uncertainty-based questioning components. ) , multimodal AMIE with reasoning consistently outperformed the vanilla baseline on key diagnostic metrics.

Specifically, Extended Data Table) shows that multimodal AMIE with reasoning achieved higher DDx accuracy. Notable top-1 accuracy gains occurred on Clinical Documents , PAD-UFES-20 and PTB-XL . Reasoning also improved information-gathering scores . Hallucination rates remained negligible in both settings although was sometimes higher for state-aware reasoning.

Management plan appropriateness also improved . Fig. 5: Auto-rater evaluation to quantify the value of state-aware reasoning and dialogue-based interaction. , State-aware reasoning ablation compares the performance of multimodal AMIE with its state-aware reasoning against a baseline without it, using the dialogue simulation environment across four datasets: skin photographs , quality of information gathering and non-hallucination rate .

DDx accuracy of the baseline LLM that infers the differentials directly from the image artifacts against that of the multimodal AMIE system, which leverages the combination of provided images and conversational history ’). Error bars center on the mean and represent 95% confidence intervals derived from 10To understand the additional performance gained with history-taking versus analyzing multimodal artifacts alone, we conducted an ablation comparing two settings: ‘Images-only’, where the base model diagnosed from images without dialogue, omitting history-taking, and ‘Images + Dialogue’, where multimodal AMIE used images, full dialogues and its state-aware reasoning at inference..

The ‘Images-only’ setting yielded markedly lower accuracy; for instance, top-1 accuracy on PAD-UFES-20 dropped significantly without dialogue. Conversely, the ‘Images + Dialogue’ setting consistently improved performance across all datasets . To evaluate system robustness against variations in patient presentation, we used LLM-driven augmentation to introduce variations to the original patient scenarios across three axes: personality styles , demographics and semantic changes to background/symptoms .

The results, detailed in Extended Data Fig.), demonstrate that the performance of multimodal AMIE across key auto-rater metrics—including diagnostic accuracy, information gathering, hallucination rate and management plan appropriateness—remained highly consistent between the original and augmented scenarios. In addition to inference-time strategies, we also explored the potential of training-time adaptations through domain-specific supervised fine-tuning to enhance model performance. We conducted experiments comparing our base model to a version fine-tuned on specialized medical dialogue and Q&A datasets.

Although these experiments revealed that SFT could yield statistically significant improvements in diagnostic accuracy for highly specialized tasks such as ECG analysis, this came at the cost of a notable decrease in performance on other crucial aspects of the consultation, including management plan appropriateness , a key challenge is effectively weaving artifact interpretation into the flow of clinical dialogue without degrading the interaction.

Multimodal AMIE achieved this via our state-aware reasoning framework , enabling the system to handle multimodal data mid-interaction while preserving clinically sound, empathetic dialogue. This capability is essential for real-world telehealth where multimodal data exchange is routine. The modalities tested reflect the expanding breadth of data present in telehealth; for instance, clinical photographs and reports are readily accessible to consumers.

Additionally, ECGs are common in primary care, but they are increasingly available on consumer-grade wearable devices, which enable recording of both single-lead and, with guidance, multi-lead configurationsIllustrated here is multimodal AMIE’s state-aware dialogue phase transition framework, which structures the diagnostic conversation. The system progresses through three distinct phases, each with a specific goal: History-taking , Diagnosis & management and Answer follow-up questions .

Within each phase, multimodal AMIE maintains an internal state—its dynamic understanding of the patient’s situation, evolving diagnoses and knowledge gaps—derived from the dialogue history and any multimodal inputs. This state guides specific actions, such as asking targeted questions, requesting multimodal data , generating internal summaries and differential diagnoses or providing explanations.

Transitions between phases are triggered automatically when the system assesses that the objectives of the current phase have been met, based on its internal state evaluation. This mechanism enforces a structured yet flexible dialogue flow, inspired by the methodical approach of experienced clinicians. Notably, multimodal AMIE demonstrated proficiency in reasoning about patient-provided multimodal artifacts, as indicated by specialist ratings.

Assessed blinded using our MUH rubric, specialist physicians gave multimodal AMIE higher scores compared to PCPs on axes encompassing artifact engagement, interpretation and artifact-grounded reasoning. Moreover, patient-actors rated multimodal AMIE significantly higher than PCPs in addressing questions regarding artifacts and explaining findings derived from visual evidence. Although caution is warranted when generalizing to broader remote care—given that clinician performance might be constrained by the text chat modality versus video—multimodal AMIE was rated higher in these communicative tasks.

This suggests that patients valued multimodal AMIE’s systematic approach to verbalizing visual findings, potentially more than the implicit assessments offered by clinicians. In one such case, after the patient sends images, multimodal AMIE summarizes the images and relates findings to the conversation history. By contrast, the PCP sometimes did not directly address the shared artifacts .

Explicitly addressing images, as multimodal AMIE did, may provide reassurance that the images were transmitted correctly and that the doctor reached a conclusion based on specific evidence. We regard this explicit acknowledgment as vital for multimodal interaction, as patients expect explanations referring to the visual evidence provided. The scenarios were designed to ensure that successful diagnosis required appropriate use of multimodal information. Subgroup analyses supported the validity of this design.

Both multimodal AMIE and PCP diagnostic accuracy dropped in cases where image quality was deemed inferior; if performance remained high irrespective of image quality, it would suggest that the consultation did not depend on the visual upload. Instead, the parallel drop in performance confirms that multimodal inputs were critical to the diagnostic process, validating these scenarios as effective tests of multimodal reasoning.

Further support was seen in how diagnostic accuracy improved for both multimodal AMIE and PCPs when they were rated as having appropriately used visual information. These findings lend credence to multimodal AMIE’s potential robustness to lower-quality images and ability to overcome intermediate misinterpretations. Ensuring robustness, reliability and equity is essential for responsible diagnostic AI, although achieving this requires extensive further study.

Our study includes extensive dialogue-level error analysis through our rubrics , but we acknowledge that a large-scale, fine-grained analysis of the model’s internal reasoning states is a limitation of our present work. Our evaluations also undertook initial steps to explore robustness. Automated evaluations using LLM-driven augmentations showed that the core metrics of multimodal AMIE remained largely stable with respect to variations in simulated patient personality, demographics and minor semantic changes.

Specialist evaluations further examined factors influencing accuracy. Notably, multimodal AMIE demonstrated greater robustness in diagnostic accuracy against specialist-rated image quality than PCPs. Although these preliminary findings regarding stability are encouraging, further research is necessary to ensure reliable performance across diverse populations.

Finally, we acknowledge that although the clinical documents and in-house generated ECG trace images were private, the skin images and underlying raw ECG signals were publicly available and may have been encountered during pretraining. We sought to mitigate this risk by integrating these artifacts into unique, synthetic patient scenarios; notably, our perception test results suggest that the influence of data leakage on performance is limited.

Nevertheless, validating on fully private or prospectively collected datasets is one of the key directions for future research.. However, this requires substantial resources and carries the risk of catastrophic forgetting, potentially degrading general conversational competence. To investigate this, we compared fine-tuned Gemini 2.0 Flash with the original model. Although experiments displayed potential gains on specific tasks, they highlighted performance degradation on other aspects, such as management plan appropriateness.

Consequently, we focused on leveraging a strong, general-purpose base model enhanced with domain-specific inference-time strategy. We selected Gemini 2.0 Flash for its higher out-of-the-box performance in multimodal understanding and conversational fluency compared to other variants. This strong baseline provided a robust foundation for the state-aware reasoning framework . Although training-time specialization remains an area for future investigation, our findings highlight the effectiveness of inference-time enhancements, a strategy also proven successful for management reasoning in related work.

This aligns with clinician feedback preferring thorough information gathering to avoid premature assessments. However, a limitation of this structured approach is potential rigidity when critical information emerges late in the dialogue, suggesting a need for more fluid state transitions. Regarding the limitations of chat-based interactions, our evaluation used a synchronous chat interface, mirroring ubiquitous instant messaging often used for medical conversation. Although accessible, this modality presents limitations compared to video or in-person visits.

Chat restricts non-verbal cues, limits dynamic visual assessments and precludes physical examinations, all of which provide crucial diagnostic information. Furthermore, video calls may foster stronger rapport and trust.

Consequently, the scope of conditions and depth of assessment in this format are constrained. Additionally, the potential for unblinding due to systematic stylistic differences between multimodal AMIE and PCPs poses a challenge for comparative studies. These limitations must be considered when interpreting findings. Highlighting the importance of real-world validation and future directions, although our findings demonstrate the potential of multimodal AMIE in simulated consultations, considerable research is essential before clinical translation.

Our study is not a randomized clinical trial but, rather, represents an exploratory investigation into the properties of multimodal diagnostic dialogue. Before real-world clinical deployment, it will be important in future work to conduct a full randomized clinical trial, conforming with full CONSORT criteria, including prespecified endpoints and preregistered statistical analysis conducted by an independent contract research organization.

Future studies must rigorously evaluate the performance, safety and reliability of multimodal AMIE under the complexities of actual healthcare delivery, including its impact on workflows and health equity. Furthermore, we acknowledge that comprehensive clinical care often relies on multimodal clinical data beyond those currently covered in this work .

Extending this framework to support additional modalities, such as radiology or pathology imaging, is a valuable direction for future research to broaden the system’s diagnostic scope. We emphasize that multimodal AMIE remains an evolving research system.

In conclusion, this work advances AMIE to dynamically integrate multimodal reasoning within diagnostic conversations. Achieved through a novel state-aware inference-time reasoning technique using Gemini 2.0 Flash, the performance of multimodal AMIE within the OSCE study matched or surpassed PCPs across diagnostic accuracy and consultation quality as rated by specialists and patient-actors.

Although further research is essential before real-world translation, these findings demonstrate progress toward AI systems that are capable of comprehensive clinical interactions.by integrating multimodal perception, enabling it to conduct diagnostic conversations that are more clinically realistic and that incorporate various forms of medical data beyond text. Our approach centers on a state-aware dialogue phase transition framework, which leverages the multimodal reasoning capabilities of Gemini 2.0 Flash.

This framework guides multimodal AMIE through structured phases of history-taking, diagnosis and management and follow-up, dynamically adapting the conversation based on intermediate model outputs that reflect the evolving patient state and diagnostic hypotheses. To rigorously evaluate the system’s performance during model development and in comparison to human clinicians, we employed a two-pronged evaluation strategy.

First, we developed an automated evaluation pipeline involving perception tests on isolated medical artifacts and simulated dialogues assessed by an auto-rater across key clinical dimensions, such as diagnostic accuracy and information gathering. Second, we conducted an expert evaluation using an OSCE-style methodology to assess the capabilities of multimodal AMIE in realistic, simulated patient encounters involving multimodal data. Real clinical diagnostic dialogues follow a structured yet flexible path.

Clinicians methodically gather information, form potential diagnoses, strategically request and interpret further details , continually update their assessment based on new evidence and eventually formulate a management plan. This process requires a clinician to adapt the line of questioning based on evolving hypotheses and uncertainty while ensuring that all critical information is considered Given the rapid advancements in LLM capabilities and their increasing proficiency in following complex instructions, one might achieve considerable progress toward emulating this process using a sophisticated system prompt alone.

However, we hypothesize that, for a safety-critical and highly dynamic task like multimodal diagnosis, building an explicit state-aware reasoning system layered on top of the LLM offers critical advantages. Such a system provides greater control over the dialogue flow, enables more reliable tracking of the diagnostic state and uncertainty, facilitates more deliberate integration of multimodal inputs and ultimately leads to higher-quality, more dependable clinical reasoning rather than relying solely on complex prompting .

Therefore, the multimodal AMIE system implements this state-aware dialogue phase transition framework to manage the diagnostic process. This framework dynamically controls the progression of multimodal AMIE through three distinct phases: history-taking, diagnosis and management and follow-up. Transitions between phases, and actions within each phase, are driven by intermediate model outputs representing the evolving patient state and diagnostic hypotheses .

Crucially, each phase builds upon the accumulated dialogue history, which contextually incorporates various forms of patient data, including text, images and clinical documents . This state-aware approach allows multimodal AMIE to emulate the structured yet adaptive reasoning process of clinicians ; positive and negative symptoms; past medical, family and social/travel histories; medications; other relevant details; and a prioritized list of knowledge gaps. Initially, this profile may contain minimal information. An internal, evolving DDx is generated. This DDx is not initially presented to the patient.

DDx generation begins after an initial interaction period, allowing baseline information collection. The frequency of DDx updates is configurable. A key decision point is whether to continue gathering history or transition to presenting a diagnosis. This decision uses a set of criteria, including assessing if sufficient information exists to formulate a reasonable differential diagnosis.

A decision module, querying Gemini 2.0 Flash, determines if current information is sufficient to proceed or if more targeted questions are needed. This module considers the dialogue history and preliminary DDx. If history-taking continues, the system generates focused questions to address information gaps identified in the patient profile and DDx uncertainty.

It is important to note that this internal measure of uncertainty, derived from the model’s iterative DDx generation, is used as a pragmatic heuristic to guide the history-taking dialogue toward more relevant lines of questioning. It is not presented to the user nor is it a formally calibrated diagnostic probability intended for direct clinical interpretation. Crucially, the system is designed to recognize when multimodal data are necessary and to strategically request them.

For instance, based on reported symptoms such as a rash, multimodal AMIE will prompt the user to upload skin photographs. Its reasoning extends to requesting additional views if needed .

Similarly, for reported cardiac symptoms, it might request an ECG tracing if available. Upon receiving an artifact, multimodal AMIE elicits detailed descriptions adhering to instructions for interpreting the artifacts and their key aspects for determining salient findings . This explicit prompting for descriptive details ensures that the model extracts salient features from the artifact to inform the ongoing conversation and diagnostic reasoning.

The generated question or request is presented to the patient. The process is iterative. New information from patient responses and data uploads is incorporated into the dialogue history, updating the patient profile and internal DDx. The patient summary is periodically refreshed to reflect the current understanding.

DDx validation : The system enters a DDx validation phase. Focused questions are generated to gather specific evidence that supports or refutes potential diagnoses within the internal DDx. A decision module determines when the DDx is sufficiently validated to present to the patient. The system presents a ranked DDx .

Crucially, the explanation for each diagnosis is grounded in evidence from the entire interaction, explicitly referencing and explaining findings from the provided multimodal data . After presenting the DDx, the system formulates a management plan based on dialogue history, patient profile and the presented DDx.

The plan includes recommendations for investigations, testing and/or treatment. The system delivers the management plan, potentially iteratively across several turns, allowing for clarification and addressing patient concerns. Plan communication and question answering: The system addresses remaining patient questions, ensuring that the patient understands the proposed management plan, potentially referencing the multimodal artifacts again to clarify points. Responses are guided by the dialogue history, patient profile, presented DDx and management plan.

Dialogue continues until patient questions are addressed, the management plan is communicated and a natural conclusion is reached. Once the dialogue reaches a natural conclusion , the system automatically generates a structured post-questionnaire. This process leverages the full multimodal dialogue history and the final internal state to produce a comprehensive summary suitable for clinical review.

Specifically, multimodal AMIE is prompted to:Based on the entire interaction, including all textual exchanges and interpretations of multimodal artifacts, generate a final, ranked DDx listing the most probable condition and several plausible alternatives. Leveraging both the established DDx and a constrained web search process for grounding in current medical knowledge and guidelines, detail the recommended management plan.

This includes proposed in-visit and ordered investigations, specific actions or recommendations for the patient and necessary escalation level with justification and follow-up requirements . This retrieval-based approach is used exclusively during this offline step to ensure that recommendations are informed by up-to-date information.

The specific methodology is detailed in Supplementary SectionIdentify and list the key clinical findings observed in any provided images that were relevant to the diagnostic and management reasoning. The structured post-questionnaire serves as a standardized record of multimodal AMIE’s clinical assessment, reasoning and recommendations based on the completed multimodal consultation; it is used both as an artifact rated by clinicians in the OSCE evaluation and as the standardized output scored by the auto-rater in simulated dialogues.

We established an automated evaluation framework to support rapid iteration and rigorous assessment of the multimodal AMIE system. This included evaluating perception capabilities on isolated medical artifacts and using a simulation environment with auto-raters to assess complete diagnostic multimodal dialogues and perform component ablations and determine optimal model selection. A critical precursor to effective multimodal diagnostic conversation is ensuring that our underlying models possess a fundamental capability: accurate perception of diverse medical artifacts.

Although LLMs have demonstrated remarkable progress in understanding and generating text, their ability to reliably ‘see’ and interpret medical images and documents— akin to a clinician’s initial visual assessment—remains less explored in the context of conversational DDx. Without robust perceptual grounding, even the most sophisticated conversational framework would be limited in its ability to meaningfully integrate multimodal data into history-taking and diagnostic reasoning.

Therefore, we first conducted a suite of perception tests designed to isolate and evaluate the visual understanding of our base models when presented with common medical artifacts: smartphone-captured skin images, ECG tracings and clinical documents. The objective was not to achieve state-of-the-art performance on complex diagnostic tasks per se but, rather, to establish a baseline confidence in whether our models could reliably discern key visual features and clinical information from these modalities in isolation.

, prompting the model to provide probable diagnoses or answer expert-validated questions based on ECG images generated from raw signals. Finally, for clinical documents, we created the ClinicalDoc-QA dataset that consists of question-answering tasks based on a large collection of deidentified clinical notes and patient records generated by physicians to evaluate the model’s comprehension of medical information from clinical documents.

Additional dataset details can be found in Supplementary Section To ensure accurate processing of artifacts, we verified the perceptual capabilities of our base model, Gemini 2.0 Flash. Our objective was not to achieve state-of-the-art performance on isolated benchmarks but, rather, to confirm a sufficient perceptual foundation for our state-aware reasoning framework. Overall, this verification demonstrated that Gemini 2.0 Flash possessed the necessary foundational perceptual grounding across our target modalities.

This provided the confidence required to build multimodal AMIE’s advanced reasoning and dialogue functions upon this base model. This is quantitatively validated by our ablation studies encompasses patient profile and scenario generation, which is then used in step for the turn-by-turn multimodal dialogue generation process used to mimic real-world telemedicine interactions, and, finally, the simulated dialogue is used in step by the auto-rater for automated assessment of multimodal AMIE’s conversational abilities. Generating realistic synthetic dialogues requires detailed patient profiles and corresponding clinical scenarios. Our methodology involves compiling comprehensive patient metadata, including symptoms, demographics and medical history, tailored to specific medical domains.

We used the SCIN and PAD-UES-20 datasets, which provide rich, preexisting patient attributes , requiring no synthetic imputation. We started with the PTB-XL ECG dataset. As it includes only age and sex, we employed Gemini 2.0 Flash, using its web search tool, to impute clinically plausible symptoms, social/family/medical history and cardiovascular risk factors associated with the ECG’s indicated condition.

This grounded the synthetic profiles in the model’s medical knowledge and real-world examples from the web. For this domain, we collaborated with the same external clinical partners who developed our OSCE scenarios. They created realistic patient profiles, document artifacts and corresponding metadata specifically for our simulation needs. Once comprehensive metadata were established for all domains, we used Gemini 2.0 Flash to generate detailed clinical scenarios.

These scenarios provide context for the patient agent in the simulation, outlining the initial presentation, details to share only upon questioning, patient expectations and concerns, desired outcomes and potential questions for the doctor. We guided Gemini using few-shot examples, ensuring scenario diversity and clinical realism. Although the raw images are from public datasets, we augmented them with unique, synthetically generated scenarios.

We designed this approach to mitigate potential data leakage by ensuring that the model encounters these images within novel clinical contexts. Crucially, the scenarios deliberately omit the final diagnosis, requiring the AI agent to deduce it through the simulated conversation and analysis of any provided multimodal artifacts. We simulate dialogues turn by turn between a doctor agent and a patient agent, mimicking a telemedical interaction.

The doctor agent is instructed to be empathetic and clinically accurate. It uses the state-aware dialogue phase transition framework to navigate history-taking, diagnosis and management and follow-up phases. A key capability is its multimodal nature: it can strategically request and analyze relevant medical artifacts based on patient-reported symptoms during the conversation.

The patient agent simulates a patient strictly following a predefined scenario from step 1, which details their profile and clinical scenario. It is prompted to respond truthfully using concise, casual language appropriate for an online consultation while adhering to specific rules regarding politeness and pacing the release of information to avoid overwhelming the doctor agent. In each dialogue turn, the doctor agent generates a question or statement based on the conversation history and its internal state.

The patient agent then formulates a response according to its scenario instructions. If the scenario dictates and the doctor agent requests it, the patient agent can upload a relevant medical artifact . Multimodal AMIE then analyzes this artifact, integrating the findings into its ongoing diagnostic reasoning and subsequent dialogue turns. This turn-by-turn exchange continues until the conversation reaches a natural conclusion or hits a predefined maximum turn limit.

After the exchange is concluded, the doctor agent generates the structured post-questionnaire. The output is a complete simulated dialogue transcript, capturing the dynamic exchange of text and multimodal data. Auto-rating is crucial for the iterative development and safety assessment of conversational models. Although human evaluation is valuable, it is often constrained by cost, time and scalability.

Auto-raters provide a mechanism for rapid, scalable and consistent performance assessment across essential characteristics. We employ an auto-rater based on Gemini 2.0 Flash to evaluate the simulated dialogues. The auto-rater, which is given access to the ground truth condition, assesses various aspects, from quantitative metrics such as diagnostic accuracy to qualitative dimensions such as information gathering and safety.

The specific criteria and scoring methods used by our auto-rater are detailed below . Evaluates the suitability of the proposed management plan relative to best medical practices and the patient’s specific situation. Rated on a five-point Likert scale . Detects instances of fabricated information stated by the model .

This safety-critical metric excludes assessments of the diagnostic or management reasoning itself and is evaluated with a binary score .by generating multimodal patient scenarios grounded in real-world datasets and simulating realistic patient−clinician interactions involving these multimodal medical data. As detailed in Supplementary SectionTo compare multimodal AMIE’s capabilities in undertaking multimodal diagnostic conversations to those of PCPs, we extended the remote OSCE study design introduced by Tu et al..

The quality of dialogues was assessed using a set of rubrics and metrics reflecting the perspective of patients and specialist physicians. We also introduced a new evaluation rubric specifically to assess the ability to use multimodal medical data effectively in the context of clinical consultations. The OSCE is a standardized practical assessment widely used in healthcare education to objectively evaluate clinical skills and competencies by simulating real-world practice.

Unlike traditional knowledge-based examinations, the OSCE assesses practical skills of real-world clinical encounters, typically involving candidates rotating through a series of timed stations where they encounter a trained patient-actor portraying a specific clinical scenario. Test-takers perform designated tasks such as taking a medical history, conducting a physical examination, interpreting results or counseling the patient. Examiners observe these interactions and score the test-taker’s performance against detailed, predefined checklists that assess crucial skills such as history-taking, examination technique, clinical reasoning and communication.

In this work, we designed and conducted a virtual analogue of the OSCE adapted for multimodal text chat. In this setting, patient-actors engaged in blinded, synchronous chat conversations with either the multimodal AMIE system or PCPs, conducted through a chat interface that allowed exchange of both text and uploaded images as is now commonplace for mobile chat applications .

Within the virtual consultations, patient-actors were instructed to upload images such as skin photographs, laboratory tests or ECG tracings, emulating how popular text chat platforms have been reportedly used as a means for remote consultation. In collaboration with two organizations that routinely perform OSCE assessments in Canada and India, we developed 105 case scenarios.

The scenarios were centered around three types of image artifacts that are commonly available to patients in telemedical primary care: smartphone-captured skin images, ECG tracings and clinical documents. We selected these modalities as the most likely to occur in primary telecare: patients can readily photograph their skin, scan documents they received or report with ECGs collected via consumer toolsprovides an overview and examples of image artifacts.

Scenarios contained different representations of these image artifacts to represent varying levels of image quality, including photographs of a given skin concern from different angles as well as screenshots and smartphone photographs of ECG tracings and clinical documents.. SCIN contains representative skin images along with rich metadata crowdsourced from real internet users with skin concerns. ECG tracings were taken from PTB-XL, the largest publicly available dataset for this modality.

Clinical documents were crafted by OSCE laboratories for the purpose of this study. Notably, we included normal cases, too. Scenarios were designed such that both requesting and interpreting images as well as taking the patient’s history were required—while either alone was insufficient—in order to form a confident diagnosis. To this end, for both skin photographs and ECG-based scenarios, we selected challenging images with high annotation ambiguity in their diagnosis labels as detailed in Extended Data Table.

Furthermore, dermatologists, cardiologists and internists ensured that the accompanying text-based component of scenarios was complementary in the sense that both image and textual information were required to arrive at an accurate diagnosis. Nevertheless, we note that, although the scenarios are consistent with the artifacts and the diagnosis, there is no guarantee that they reflect the true case history, as they were created post hoc.

Scenarios were crafted to match case metadata provided in the SCIN and PTB-XL datasets when available. For Clinical Documents, we selected conditions for which diagnosis formation can be guided based on the results of laboratory reports.

Lastly, to simulate variability in image quality in real-world care settings, for half of the ECG and clinical document cases, we used smartphone photographs of a computer screen showing the artifacts , in a blinded and randomized order.

The chat interface supported text-based communication and sharing of images from the scenarios while displaying the scenario pack information to patient-actors throughout the conversation. After each consultation, patient-actors completed a questionnaire to rate their experience to represent the patient-actor perspective. Separately, a different questionnaire was completed by both multimodal AMIE and PCPs to summarize key clinical findings and next steps.

In particular, the questionnaire asked for a DDx list , a management plan and a description of salient image findings. Lastly, a group of specialist physicians assessed the performance of multimodal AMIE and PCPs in a blinded fashion .

The questions in this rubric were derived from consideration of authoritative assessment schemes of clinical interactions such as the PACES, used by the Royal College of Physicians in the UK for examining history-taking skills The MUH rubric was designed to assess competence in handling and interpreting multimodal artifacts in the context of clinical consultations, including the ability to understand medical image artifacts; to use that understanding to guide the conversation and inform a clinically accurate assessment; and to communicate the salient findings and address the patients’ questions in an appropriate manner.

Details are provided in Extended Data TableThis study involved 19 board-certified PCPs and 25 validated patient-actors, with participants split between India and Canada . Informed consent was obtained from each participant before their participation. To ensure consistent and high-quality interactions, all patient-actors and PCPs completed a standardized training prior to participation.

This training included an interactive workshop to familiarize them with the chat interface and the specifics of the OSCE interaction format based on detailed instructional guides. The PCPs had a median post-residency experience of 6 years with an interquartile range of 3.5−11.5 years.

For quality assessment of the multimodal AMIE/PCP consultations and their post-questionnaire responses, we recruited 18 independent specialist physicians from India and North America over three medical specialities to ensure diverse clinical viewpoints as well as requisite expertise. These specialists were independent of the study team and the patient-actor cohort, had a median post-residency experience of 5 years and were assigned evaluation tasks matching the medical specialty of the scenarios .

For each of 105 scenarios, each assigned patient-actor performed two sessions, one with a PCP and another with multimodal AMIE in a randomized order, yielding a total of 210 consultations. Each of these conversations was evaluated by three independent specialists. We analyzed the results from the remote OSCE study using a mixed-effect approach to account for differences in the scenarios using random effects.

Most rating choice options for patient-actors and specialist physicians were ordinal , and we model them using a cumulative ordinal model with logit link function. For binary ratings , we use a Bernoulli model with logit link. For diagnostic accuracy, we again use a Bernoulli model for ‘Correct/Incorrect’.

In all models, the scenario is used as a random intercept to model differences in the quality of the patient vignettes or the difficulty of the diagnostic task. Both of these factors can affect the flow of conversation and should, therefore, be considered when modeling ratings of these conversations. The experimental arm is modeled as a fixed effect.

For diagnostic accuracy, we also model the fixed effect of the differential size rather than designating a single primary endpoint.values were corrected for FDR using the Benjamini–Hochberg method as implemented in the Python library ‘statsmodels’. These corrections were applied to all statistical estimates of the experimental arm on either patient-actor or expert ratings, both when estimating parametric ordinal models and when estimating preference .

In all figures, we display confidence intervals that show the 2.5th and 97.5th percentiles of bootstrapped distributions of the displayed quantity. Many of the real-world datasets used in the development of multimodal AMIE are open source and can be downloaded upon completion of the required training. Further inquiries about our benchmarking procedures and data analysis may be addressed to the corresponding authors with a maximum response time of 2 weeks.

Additional scenario packs used in the study will be made available upon reasonable request. Our system uses Gemini 2.0 Flash as its base foundation model. Base Gemini models, including Gemini 2.0 Flash, are generally available via Google Cloud APIs. The core techniques, particularly the state-aware dialogue phase transition framework detailed in the, provide the necessary details for reproducing our approach.

However, the specific implementation relies on internal Google infrastructure and tooling. Due to this and, more importantly, the safety implications associated with the unmonitored deployment of AI systems in medical contexts, we are not open-sourcing the codebase and the specific prompts employed in our work at this time. In the interest of responsible innovation, we will be working with research partners, regulators and healthcare providers to further validate and explore safe onward uses of our medical models.

Russo, G., Perelman, J., Zapata, T. & Šantrić-Milićević, M. The layered crisis of the primary care medical workforce in the european region: what evidence do we need to identify causes and solutions? Lee, J. & Kontopantelis, E. A systematic review exploring the factors that contribute to increased primary care physician turnover in socio-economically deprived areas. Zeltzer, D. et al. Comparison of initial artificial intelligence and final physician recommendations in AI-assisted virtual urgent care visits.

Hong, G., Smith, M. & Lin, S. The AI will see you now: feasibility and acceptability of a conversational AI medical interviewing system. Campanozzi, L. L. et al. The role of digital literacy in achieving health equity in the third millennium society: a literature review. Mohamed, I. N. & Elseed, M. A. Utility of WhatsApp in healthcare provision and sharing of medical information with caregivers of children with neurodisabilties: experience from Sudan.

Li, K., Elgalad, A., Cardoso, C. & Perin, E. C. Using the Apple Watch to record multiple-lead electrocardiograms in detecting myocardial infarction: where are we now? Ouyang, L. et al. Training language models to follow instructions with human feedback. InRitunga, I., Claramita, M., Widaty, S. & Soebono, H. Challenges and recommendations in the implementation of audiovisual telemedicine communication: a systematic review.

Google ScholarPacheco, A. G. C. et al. PAD-UFES-20: a skin lesion dataset composed of patient data and clinical images collected from smartphones. Sloan, D. A., Donnelly, M. B., Schwartz, R. W. & Strodel, W. E. The Objective Structured Clinical Examination. The new gold standard for evaluating postgraduate clinical performance.

Phillips, N. A. et al. CheXphoto: 10,000+ smartphone photos and synthetic photographic transformations of chest X-rays for benchmarking deep learning robustness. InDacre, J., Besser, M. & White, P. MRCP PART 2 Clinical Examination : a review of the first four examination sessions . This study was funded by Google LLC and/or a subsidiary thereof.

The project was an extensive collaboration among many teams at Google Research and Google DeepMind. We thank D. Racz, S. Mohamed, C. Wright, R. Sico, J. Wilson, B. Gabriel and J. Sturgeon for their comprehensive review and detailed feedback on the manuscript. We also thank J. Grondie, K. Moore and J. Guilyard for their contributions to the animations and visuals.

We would like to thank S. Man, B. Hatfield and G. Turner for supporting the OSCE study; GoodLabs Studio Inc., Intel Medical Inc. and C. Smith for their partnership in conducting the OSCE study in North America; and JSS Academy of Higher Education and Research and V. Patil for their partnership in conducting the OSCE study in India. We are also grateful to E. Dominowska, K. Kulkarni, S. Garg, R. Wong, A. Wang, R. Ruparel and E. Wulczyn for their support during the course of this project.

Lastly, we would like to thank R. Hadsell, Z. Ghahramani, O. Vinyals and K. Kavukcuoglu for their support of this work. These authors contributed equally: Khaled Saab, Chunjong Park, Tim Strother, Jan Freyberg, Ryutaro Tanno.

Khaled Saab, Chunjong Park, Tim Strother, Jan Freyberg, David G. T. Barrett, Yong Cheng, Wei-Hung Weng, David Stutz, Nenad Tomasev, Valentin Liévin, Elahe Vedadi, Geoff Brown, Yang Gao, Sean Li, S. Sara Mahdavi, Avinatan Hassidim, Joëlle Barral, S. M. Ali Eslami, Pushmeet Kohli, Vivek Natarajan, Tao Tu, Alan Karthikesalingam & Ryutaro TannoAnil Palepu, Yash Sharma, Roma Ruparel, Abdullah Ahmed, Kimberly Kanada, Cian Hughes, Yun Liu, James Manyika, Katherine Chou, Yossi Matias, Dale R. Webster, Adam Rodman & Mike Schaekermann R.T. , A.K.

, T.T and V.N. contributed to the conception and design of the work. R.T. and A.K. contributed to the data acquisition and curation. K.S. , C.P.

, T.S. , J.F. , D.G. T.B.

, Y.C. , W.-H.W. , D.S. , N.T.

, A.P. , V.L. , Y.S. , R.R.

, A.A. , T.T. and R.T. contributed to the technical implementation. J.F. , R.T.

, A.K. , S.L. and M.S. contributed to the expert evaluation framework used in the study. A.K. , A.R. and K.K. provided clinical inputs to the study.

S.S. M., J.M. , K.C. , Y.M.

, A.H. , D.R. W., J.B. , S.M.

A.E. and P.K. contributed to the ideation and execution of the work. All authors contributed to the drafting and revising of the manuscript. K.S. , C.P.

, T.S. , J.F. , D.G. T.B.

, Y.C. , W.-H.W. , D.S. , N.T.

, A.P. , V.L. , Y.S. , R.R.

, A.A. , E.V. , K.K. , C.H.

, Y.L. , G.B. , Y.G. , S.L.

, S.S. M., J.M. , K.C. , Y.M.

, A.H. , D.R. W., J.B. , S.M.

A.E. , P.K. , A.R. , V.N.

, M.S. , T.T. , A.K. and R.T. are employees of Google and may own stock as part of the standard compensation package. K.S. was a Google employee when he conducted this work and may also own stock.thanks Majid Afshar, Martin Härter and the other, anonymous, reviewer for their contribution to the peer review of this work.

Primary Handling Editors: Lorenzo Righetto and Michael Basson, in collaboration with theFor illustration purposes, all responses from five-point rating scales were mapped to a generic five-point scale ranging from ‘Very favorable’ to ‘Very unfavorable’. For Yes/No questions, a ‘Yes’ response was mapped to the same color as ‘Favor- able’ and a ‘No’ response to the same color as ‘Unfavorable’.

Rating scales were adapted from the General Medical Council Patient Questionnaire , the Practical Assessment of Clinical Examination Skills , and a narrative review about Patient-Centered Communication Best Practice . The McNemar test was used for each evaluation axis, and p-values were adjusted using the false discovery rate correction. Asterisks represent statistical significance .

Multimodal conversation and reasoning qualities as assessed by specialist physicians. Specialists rated conversations on ordinal scales. For illustration purposes, all responses from five-point rating scales were mapped to a generic five-point scale ranging from ‘Very favorable’ to ‘Very unfavorable’. The only four-point scale was mapped to the same scale, ignoring the ‘Neither favorable nor unfavorable’ option.

For Yes/No questions, a ‘Yes’ response was mapped to the same color as ‘Favorable’ and a ‘No’ response to the same color as ‘Unfavorable’. Rating scales were adapted from the Practical Assessment of Clinical Examination Skills , a narrative review about Patient-Centered Communication Best Practice , and other sources. The McNemar test was used for each evaluation axis, and p-values were adjusted using the false discovery rate correction.

Asterisks represent statistical significance . Auto-rater performance of multimodal AMIE compared between original synthetic patient scenarios and scenarios with augmentations . Augmentations tested multimodal AMIE’s robustness across Demographic, Personality, and Semantic axes, while maintaining clinical conclusions.

Error bars center on the mean and represent 95% confidence intervals . Significance markers denote p-values obtained from a two-sided Mann-Whitney U test comparing each augmentation condition against the baseline within each metric. Asterisks represent statistical significance .

Consistent performance across augmentations demonstrates multimodal AMIE’s robustness to non-clinically significant variations in patient presentations. Extended Data Fig. 4 Gemini 2.0 vs. SFT performance on top-3 diagnostic accuracy and management plan appropriateness. This figure compares Gemini 2.0 and a supervised fine-tuned version on Top-3 differential diagnosis accuracy and Mx plan appropriateness across SCIN , PTB-XL , and Clinical Documents tasks using synthetic dialogues.

Fine-tuning notably enhances performance on the PTB-XL ECG task for both Top-3 accuracy and plan appropriateness. SFT also yields improvements, particularly in plan appropriateness, on the Clinical Documents task, while showing modest gains on SCIN. Evaluation uses agent-based simulations, distinct from human-actor evaluations. Error bars center on the mean and represent 95% confidence intervals .

Significance markers denote p-values obtained from a two- sided Mann-Whitney U test comparing the two conditions within each panel/metric. Asterisks represent statistical significance . We assess foundational visual understanding across diverse medical data types: SCIN , PTB-XL , ECG-QA , and ClinicalDoc-QA .

We evaluate top-k accuracy for PTB-XL and SCIN, reflecting diagnostic classification ability. For ECG-QA and ClinicalDoc-QA, we report exact match accuracy, indicating the model’s correctness on question answering. Three multimodal LLMs are evaluated: Gemini 1.5 Flash, Gem- ini 1.5 Pro, and Gemini 2.0 Flash. Results indicate that Gemini 2.0 Flash generally exhibits robust perceptual capabilities.

Error bars center on the mean represent 95% confidence intervals derived from 104 bootstrap samples. Step 1: Patient actors simulate cases based on multimodal scenarios containing history, symptoms, and visual artefacts and conduct synchronous chat consultations with both a PCP and multimodal AMIE in a randomized, blinded order. Step 2 : After consultations, patient actors assess interaction quality by filling out a questionnaire, providing patient-centric quality measures.

In parallel, both PCPs and multimodal AMIE document their findings by generating answers to a separate post-questionnaire detailing differential diagnosis, management plans and also salient image findings. Subsequently, specialist physicians evaluate the performance of PCPs and multimodal AMIE based on the dialogue transcript, post-questionnaire answers, and scenario ground truth across multiple criteria.

This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it.

The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit

GooglePlease follow us on Google to support us
We have summarized this news so that you can read it quickly. If you are interested in the news, you can read the full text here. Read more:

NatureMedicine /  🏆 451. in US

 

United States Latest News, United States Headlines

Similar News:You can also read news stories similar to this one that we have collected from other news sources.

Exclusive: Freedom Caucus Cheers House Advancing Privacy Protections Against Data Brokers in Appropriations BillExclusive: Freedom Caucus Cheers House Advancing Privacy Protections Against Data Brokers in Appropriations BillFreedom Caucus lawmakers cheered the House Appropriations Committee advancing an amendment to a major funding bill that would protect Americans' privacy against data brokers.
Read more »

3 Positions Where Jaguars Made the Most Significant Improvements This Offseason3 Positions Where Jaguars Made the Most Significant Improvements This OffseasonWhich position groups made the biggest leaps for the Jacksonville Jaguars in 2026?
Read more »

Company hatches first chicks from artificial egg, advancing avian embryo development, de-extinctionCompany hatches first chicks from artificial egg, advancing avian embryo development, de-extinctionA biotech company that aims to resurrect lost creatures says it has hatched live chicks in an artificial environment. Colossal Biosciences says 26 baby chickens were born from a 3D printed lattice structure that mimics an eggshell.
Read more »

Government Officials Plan to Snoop on People's Home Improvements for Labour's Mansion TaxGovernment Officials Plan to Snoop on People's Home Improvements for Labour's Mansion TaxGovernment officials are to snoop on people's extensions and home improvements in a bid to drag more people into Labour's mansion tax. The Valuation Office will record changes to properties such as extensions and renovations to monitor whether they hit the threshold for paying the controversial levy.
Read more »



Render Time: 2026-07-17 02:28:37