Advancing AMIE towards expert-level audio-visual clinical consultations

原文擷取

When a physician meets a patient, the consultation extends far beyond the words exchanged. The physician observes the patient’s gait, registers visible signs of discomfort, notes their breathing, and guides the patient through physical examination maneuvers. This continuous stream of visual and auditory information is seamlessly integrated with the spoken clinical history. These non-verbal visual and auditory cues are central to effective diagnosis, patient trust, and clinical communication.

AI systems capable of clinical reasoning and dialogue have the potential to dramatically increase access to medical expertise and care, fostering a future where physicians can focus their time on the most meaningful aspects of patient interactions. In early work, the Articulate Medical Intelligence Explorer (AMIE), Google’s research AI system for clinical reasoning and dialogue, demonstrated expert-level performance in text-based diagnostic dialogue and proved effective as a differential diagnosis aid for clinicians. Recently, Google advanced AMIE’s capabilities beyond diagnosis towards treating and managing disease over time.

Google has also extended AMIE’s capabilities towards specialist-level evaluations in oncology, cardiology and ophthalmology, and multimodal diagnostic reasoning over images and clinical documents, in simulated settings with patient actors. In parallel, Google has begun translating these research advances towards clinical practice through a framework for physician-centered oversight, a clinical feasibility study with Beth Israel Deaconess Medical Center, and an ongoing nationwide randomized study in partnership with Included Health.

Text-based interfaces discard the visual and auditory dimensions of clinical practice. Patients must translate complex physical symptoms into written descriptions, which can discard diagnostic information and negatively affect patients with limited digital or health literacy. Text-only systems cannot independently observe the visual and auditory cues that inform clinical reasoning, nor can they guide patients through the physical examination maneuvers that shape differential diagnosis.

Google presents AMIE in a real-time video configuration, AMIE (Video), to address these limitations. Built on Gemini and Project Astra, AMIE (Video) conducts synchronous clinical video consultations, perceives non-verbal clinical cues, guides patient actors through virtual physical examinations, and reasons diagnostically in real time. In a multi-arm randomized study with 100 scenarios, 300 live consultations, and a group of 30 board-certified primary care physicians (PCPs), Google presents what it describes as the first demonstration of an AI system exhibiting expert-level performance in real-time clinical video consultations.

AMIE (Video): An asynchronous multi-agent architecture

Conducting an effective clinical conversation over video requires balancing competing demands: the system must respond to patients at natural conversational speed while simultaneously performing careful clinical reasoning and continuously processing visual and auditory streams. A single agent cannot currently satisfy all these requirements. Deep reasoning takes time, but conversational pauses erode patient trust and rapport.

AMIE (Video) therefore uses an asynchronous multi-agent architecture that divides labor across three specialized agents working continuously in parallel:

  • Talker agent:The patient-facing agent drives responsive, low-latency spoken interaction. It maintains natural conversational flow while incorporating guidance from the other agents.
  • Planner agent:Operating in the background, this agent continuously refines the system’s clinical reasoning, updates differential diagnoses and management plans, identifies information gaps, and re-prioritizes diverse clinical goals.
  • Perception agent:This agent continuously reviews the audio and visual streams, identifies clinically relevant non-verbal cues such as visible signs of distress, physical findings or auditory signals, and contextualizes these observations within the ongoing conversation.

This decoupled design allows AMIE (Video) to maintain natural conversational latency while performing diagnostic reasoning and audio-visual perception that would otherwise introduce unacceptable delays. Automated evaluations reported by Google confirm that each agent in this three-agent architecture contributes to improvements on clinical metrics, including history-taking, clinical reasoning and treatment recommendations, as well as dialogue-quality metrics such as patient-centered communication and response latency.

Guiding development with automated evaluation

A key challenge in building audio-visual medical AI is characterizing a system’s perceptual and reasoning capabilities at scale. To guide development, Google derived a taxonomy of clinical audio-visual competencies relevant to telehealth from the medical literature. The taxonomy covers non-verbal visual cues, auditory signals and physical examination maneuvers, and supports an automated evaluation suite.

The evaluation framework combines targeted single-turn audio-visual assessments with multi-turn simulated audio consultations. Single-turn assessments test specific instances of clinical perception and reasoning, such as correctly identifying anatomical laterality or recognizing signs of respiratory distress. Multi-turn simulated audio consultations assess end-to-end conversational performance while injecting visual cues as textual descriptions into the simulation. For example, an AI patient simulator for a Parkinson’s scenario may inject a verbal description such as [holding up paper to camera showing cramped, tiny script] when the patient is prompted to show their handwriting.

Together, these complementary evaluations enabled rapid iteration on system design and characterized AMIE (Video)‘s capabilities and failure modes before human evaluation.

Evaluation through a randomized video study

To evaluate clinical competence in the more challenging and realistic setting of an end-to-end audio-visual clinical consultation, Google conducted a large-scale, randomized Objective Structured Clinical Examination (OSCE) study with a synchronous video consultation interface.

The study covered 100 clinical scenarios across five body systems: cardiopulmonary; abdominal; head, eyes, ears, nose and throat (HEENT); neurological and psychiatric; and musculoskeletal conditions. Fifteen trained patient actors carried out 300 standardized consultations across three study arms:

  • AMIE (Video):The video configuration of AMIE conducting real-time video consultations.
  • AMIE (Text):A text-only version serving as a baseline to isolate the contribution of audio-visual capabilities.
  • PCP (Video):Ten board-certified PCPs consulting through the same video interface.

An independent panel of 20 experienced primary care physicians evaluated all consultations using established clinical rubrics, including general clinical competency scales and detailed case-specific scoring criteria tailored to each scenario.

Key results

  • Expert-level clinical performance:Across core clinical competencies, history-taking thoroughness, diagnostic accuracy, management appropriateness and communication quality, clinical evaluators rated AMIE (Video) on par with PCPs. AMIE (Video) also matched or exceeded AMIE (Text) on these dimensions.
  • Strength in physical observation and examination:AMIE (Video) was rated significantly higher, on average, than both PCPs and AMIE (Text) at eliciting physical signs and proactively guiding patient actors through virtual examination maneuvers. This advantage was also reflected in case-specific perception and examination rubric scores.
  • Patient actors preferred the video experience:Patient actors strongly preferred the synchronous video interface over text-based chat, rating it as significantly easier to use and more effective for communicating health concerns. They also rated AMIE (Video) favorably on empathy, rapport and confidence in care compared with both PCPs and AMIE (Text).

Limitations and responsible development

The study was conducted entirely with professional patient actors in simulated clinical settings, not with real patients presenting with their own health conditions. Patient actors cannot fully replicate the complexity and unpredictability of real clinical encounters, and the scenarios were limited to conditions that can be authentically portrayed through acting. They omitted important clinical presentations where audio-visual perception would be diagnostically consequential.

Targeted automated evaluations also revealed occasional perceptual and reasoning errors despite overall high-quality conversation and diagnostic accuracy. The system still exhibits intermittent technical issues that can disrupt conversational naturalness. Given the prototype nature of Project Astra, some technical considerations may need to be addressed at a system level beyond the specific medical application explored here.

Assessing these findings in studies with real patients and real clinical conditions is an essential next step before drawing conclusions about real-world utility.

Looking ahead

Google describes this work as demonstrating that the transition from text-based to audio-visual clinical AI is achievable at expert-level quality. AMIE (Video) observes non-verbal cues, guides physical examination and converses through spoken dialogue, more closely approximating a telehealth video encounter.

Important questions remain on the path towards responsible real-world evidence. The findings need validation with real patients, expansion to clinical presentations that cannot be enacted, and robust safety frameworks. Google points to the real-world feasibility study with Beth Israel Deaconess Medical Center and the ongoing nationwide randomized study with Included Health as early steps. These studies are intended to inform how audio-visual capabilities might be responsibly integrated into clinical practice; the current study itself does not establish real-world clinical utility.

Acknowledgements

The research is joint work across many teams at Google Research and Google DeepMind. The named co-authors and contributors include Mahvish Nagda, Jihyeon Lee, Matthew Thompson, CJ Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemenway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno and Cameron Chen.

Source notes

  • This raw capture preserves the Google Research title, publication metadata, AMIE (Video) architecture, automated evaluation framework, randomized OSCE design, reported results and responsible-development limitations.
  • The source is an official Google Research report of a research prototype. Claims about expert-level performance, evaluator ratings, patient-actor preference and contributions of the three-agent architecture remain source-attributed; the page is not an independent replication or clinical approval.
  • The randomized study used professional patient actors and simulated scenarios. Real-patient validation, clinical safety, regulatory status and deployment utility require separate evidence.