AI in Clinical Medicine, ISSN 2819-7437 online, Open Access
Article copyright, the authors; Journal compilation copyright, AI Clin Med and Elmer Press Inc
Journal website https://aicm.elmerpub.com

Review

Volume 2, August 2026, e35


Optimizing the Use of AI Agents in Electronic Health Record Workflows in Healthcare: Clinical Integration and Human Factors Review

Supreet Kaura, c, Andrew Blanfordb

aSection of Biomedical Engineering Department, Wayne State University, Detroit, MI 48202, USA
bSection of Biomedical Engineering Department, Wright State University, Fairborn, OH, USA
cCorresponding Author: Supreet Kaur, Section of Biomedical Engineering Department, Wayne State University, Detroit, MI 48202, USA

Manuscript submitted July 29, 2026, accepted August 13, 2026, published online August 25, 2026
Short title: Optimizing AI Agents in EHR Workflows in Healthcare
doi: https://doi.org/10.14740/aicm35

Abstract▴Top 

Background: Electronic Health Record (EHR) systems impose substantial administrative burden on clinicians and are strongly associated with cognitive overload and professional burnout. In parallel, autonomous and semi-autonomous artificial intelligence (AI) agents—such as ambient documentation scribes, Microsoft Copilot, ChatGPT-based tools, and clinical chatbots—provide transformative potential to re-engineer EHR-mediated workflows in healthcare. If poorly integrated, however, these agents can introduce verification bottlenecks, hallucination risks, and new sources of cognitive load. The study aimed to evaluate clinical operational transformations driven by EHR-integrated AI agents, synthesize human factors engineering (HFE) design frameworks, and establish practical guidelines for integrating conversational AI into hospital IT infrastructure.

Methods: A PRISMA 2020–guided systematic review was conducted across PubMed/MEDLINE, Embase, IEEE Xplore, and Scopus to identify empirical studies evaluating cognitive load, AI agent usability, workflow throughput, and error profiles in AI agent–assisted health IT environments. Eligible studies enrolled clinicians using EHRs in inpatient or outpatient care, deployed autonomous or semi-autonomous AI agents integrated with EHR workflows, and reported human-centered outcomes such as documentation time, administrative time, NASA-TLX cognitive load, System Usability Scale (SUS) scores, and hallucination or omission rates.

Results: Across ambient AI scribe implementations and AI-assisted documentation workflows, AI agent deployment reduced per-patient documentation time by approximately 75%, daily administrative overhead by about 60%, and after-hours “pajama time” by up to 70%. Quantitative measures of cognitive workload and usability generally improved, but new burdens emerged around verification, auditing, and management of AI-generated content. Three core HFE integration patterns—visual provenance, progressive disclosure, and attestation interlocks—were identified as effective strategies to mitigate verification fatigue and reduce hallucination risk when embedded into EHR-native interfaces.

Conclusions: Clinical engineers and healthcare IT leaders need to prioritize HFE-driven integration architectures when deploying AI agents into EHR workflows. Without robust visual grounding and verification workflows, gains in generative efficiency are offset by downstream cognitive overload. Embedding AI agents via SMART on FHIR within native EHR environments, coupled with human-in-the-loop safety interlocks, can safely translate generative AI gains into improved clinical decision-making, operational throughput, and provider well-being.

Keywords: Artificial intelligence; Electronic Health Records; Ambient AI scribes; Microsoft Copilot; Human factors engineering; Cognitive load

Introduction: Electronic Health Record (EHR) Systems Have Become Foundational Infrastructure for Modern Clinical Practice▴Top 

EHR continues to impose substantial administrative and cognitive burden on clinicians. Time-consuming documentation, complex tab navigation, fragmented data display, and high interruption rates contribute to cognitive overload and burnout, particularly in ambulatory and high-volume settings [1]. Rapid advances in artificial intelligence (AI)—especially large language models (LLMs) and ambient clinical documentation technologies—have created new opportunities to re-engineer EHR-mediated workflows [2, 3]. Ambient digital scribes, enterprise copilots (e.g., Microsoft Copilot for Healthcare, Nuance DAX Copilot), and interactive clinician chatbots now act as autonomous and semi-autonomous AI agents across multiple operational vectors as shown in Figure 1, including clinical documentation, chart review synthesis, inbox message triage, and revenue cycle workflows.


Click for large image
Figure 1. Multimodal data fusion architecture with AI agents (Microsoft Copilot, ChatGPT, & Clinical Chatbots).

Despite their promise, AI agents can shift cognitive load from generative writing to verification auditing, especially when interfaces lack visual grounding, clear provenance, or efficient review mechanisms [47].

Clinical engineers—responsible for optimizing health technology—require evidence-based design frameworks to integrate AI agents safely and effectively into EHR environments.

Methodology▴Top 

We conducted a systematic review of empirical studies evaluating AI agents integrated into EHR workflows, following the PRISMA 2020 statement as shown in Figure 2. The review focused on operational, cognitive, and usability outcomes associated with ambient documentation systems, Microsoft Copilot for Healthcare, ChatGPT-based clinical documentation, and specialized clinician chatbots.


Click for large image
Figure 2. PRISMA 2020 flow diagram illustrating identification, screening, eligibility, and inclusion of studies evaluating AI agents (ambient documentation systems, Microsoft Copilot, ChatGPT, and clinical chatbots) in clinical workflows.
Data Sources and Search Strategy▴Top 

Comprehensive database searches were executed across PubMed/MEDLINE, Embase, IEEE Xplore, and Scopus. Search strategies combined keywords and controlled vocabulary terms related to “ambient artificial intelligence,” “digital scribes,” “clinical chatbots,” “Microsoft Copilot,” “large language models,” “electronic health records,” “cognitive workload,” “NASA-TLX,” “System Usability Scale,” and “documentation time.” Reference lists of included articles, relevant reviews, and industry white papers were hand-searched to identify additional studies [811]. Searches were limited to peer-reviewed articles in English assessing clinician-facing AI systems deployed in real or simulated EHR environments.

Exclusion criteria comprised studies focusing solely on algorithmic accuracy metrics without human usability evaluation, non-clinical populations, standalone non-EHR tools, editorials, opinion pieces, white papers, and trade publications. Titles and abstracts were screened independently by reviewers to exclude clearly ineligible studies [1216]. Full-text articles were then assessed against PICOS criteria. Data extraction captured: clinical setting, EHR platform, AI agent type and integration modality, study design, primary and secondary outcomes, and reported usability, cognitive load, and error metrics. Where available, effect sizes on documentation time, administrative time, and burnout-related measures were recorded.

Results▴Top 

Study characteristics

Included studies spanned ambulatory primary care, subspecialty clinics (e.g., oncology, cardiology), and mixed inpatient–outpatient settings. EHR platforms predominantly include Epic and Cerner, with AI agents integrated via native workflows, SMART on FHIR apps, or proprietary ambient documentation pipelines. Sample sizes ranged from small pilot implementations to multi-site pre–post studies. The technical and operational comparison of clinical AI modalities was conducted as shown in Table 1.

Table 1.
Click to view
Table 1. Technical and Operational Comparison of Clinical AI Modalities
 

Operational workflow transformations

Across settings, ambient AI scribes and copilots transformed documentation workflows from manual typing or traditional dictation to real-time ambient capture and AI-generated SOAP notes. Clinicians reported substantial reductions in in-visit typing and post-visit completion time, with gains in perceived focus on patient interaction [17, 18].

For chart review, clinician-facing chatbots and Copilot-style agents enabled natural language queries such as “summarize this patient’s last three cardiology visits” or “show relevant labs and imaging for heart failure management,” replacing tab-by-tab navigation with guided, conversational exploration, (as shown in Table 2. These AI agents synthesized longitudinal timelines, care gaps, and risk factors into single-page summaries as shown in operational transformation matrix across EHR dimensions. The workflow operational transformational matrix across the EHR Dimensions was analyzed as shown in Table 2.

Table 2.
Click to view
Table 2. Operational Workflow Transformation Matrix Across EHR Dimensions
 

Inbox triage agents restructured in-basket workflows by automatically categorizing patient portal messages by urgent clinical, routine clinical, administrative, flagging high-risk messages, and pre-drafting responses for clinician review and signature [19, 20]. This shift reduced manual triage burden and improved prioritization of clinically critical messages.

Revenue cycle AI agents used Microsoft Copilot and clinical chatbots to auto-extract medical necessity evidence, match documentation to payer rules, and pre-fill prior authorizations and coding fields. These workflows reduced back-office chart review time and improved alignment between documentation and billing requirements.

Quantitative impact on time and workload

Across ambient AI documentation and triage implementations, studies were reported as shown in Table 3.

Table 3.
Click to view
Table 3. Quantitative Workflow and Efficiency Impact Metrics
 
  • Per-patient documentation time reductions from approximately 10–12 to 2–4 min (≈75% reduction).
  • Decreases in daily administrative time from about 2.5 h to around 1 h per full-time equivalent (≈60% reduction).
  • Reductions in weekly after-hours (weekly) tasks from roughly 8–10 to 2–3 h (≈70% reduction) as shown in Table 3.

Human factors engineering (HFE) frameworks

Visual provenance and bi-directional grounding

HFE analysis across included studies highlighted visual provenance and bi-directional grounding as critical mechanisms for safe AI integration. Interfaces visually differentiate AI-generated text from clinician-authored content and allow users to trace generated statements back to source audio or chart elements, thus reducing verification friction [21, 22].

Bi-directional grounding techniques—such as hover-to-highlight corresponding transcript segments, audio timestamps, or EHR data fields—allow clinicians to rapidly validate AI-generated content. These patterns mitigate hallucination risk by making evidence pathways visible and actionable.

Progressive disclosure and chunked output

Chunking AI agent output into semantically distinct visual cards (e.g., “history,” “assessment,” “plan,” “medications”) and using progressive disclosure avoid overwhelming clinicians with dense, monolithic drafts. High-level summaries can be presented first, with optional expansion into detailed evidence views. These design patterns help clinicians prioritize review, focus on clinically important sections, and reduce time spent scanning long notes. Progressive disclosure also supports safer CDS by surfacing high-risk alerts and recommendations at the appropriate level of detail.

Attestation interlocks and verification workflows

Attestation interlocks—structured checkpoints where clinicians must explicitly review and confirm AI-generated sections—were identified as key tools to manage verification burden. These can be embedded at the note, section, or statement level, with streamlined UI controls (e.g., accept, edit, discard).

Well-designed verification workflows strike a balance between safety and efficiency: they avoid excessive clicks and redundant confirmations while ensuring that high-risk content, such as diagnostic impressions or orders, receives appropriate human review.

Interoperability architecture: SMART on FHIR integration

Several studies and implementation reports emphasized SMART on FHIR-based architectures as enablers of secure, scalable AI agent integration into EHRs. SMART on FHIR apps can access standardized FHIR resources (patient, encounter, observation, medication, document, and reference) via FHIR APIs, enabling AI agents to operate on near real-time clinical data.

This architecture supports multiple AI modalities—ambient documentation, chart intelligence, triage, and CDS—without requiring separate data silos or manual exports. It also facilitates governance by centralizing audit logs, consent management, and identity/authorization (e.g., OAuth2/OpenID Connect) at the EHR layer. For clinical engineers, SMART on FHIR offers a practical blueprint for integrating multiple AI agents while preserving security, privacy, and interoperability [21, 22].

Discussion▴Top 

This study demonstrated that EHR-integrated AI agents can substantially reduce documentation time, administrative burden, and after-hours work, while improving perceived usability and cognitive workload measures. Ambient scribes and copilots reframe documentation as a review-and-attest task; chart intelligence agents transform navigation into question-driven exploration; triage agents reshape message management into risk-stratified queues; and revenue cycle agents automate evidence extraction for prior authorizations and coding. At the same time, the review underscores that naive AI integrations can simply relocate cognitive overload to verification and auditing, particularly if interfaces lack Visual Provenance, Progressive Disclosure, and Attestation Interlocks. HFE must therefore be treated as a first-class design constraint in AI agent deployment.

This study proposed a four-layer “HFE-AI integration loop” to unify operational, human factors, and technical dimensions of EHR-integrated AI agents as shown in Figure 3. The Data & Context Layer aggregates multimodal clinical and operational data via FHIR/SMART or vendor-native APIs. The Agent Reasoning & Action Layer performs planning, action, reflection, and memory to generate clinical and administrative outputs. The Human–AI Interaction & Verification Layer surfaces these outputs through EHR-native interfaces that implement Visual Provenance, Progressive Disclosure, and Attestation Interlocks [22, 23]. Finally, the Organizational Learning & Governance Layer monitors safety, performance, and workflow impact, feeding back into model and workflow iteration. This loop positions AI agents not as standalone tools but as components of a socio-technical system in which human oversight, interface design, and governance determine whether generative efficiency translates into net reductions in cognitive load and burnout.


Click for large image
Figure 3. A four-layer concentric framework for healthcare AI agents, with Data & Context at the core, surrounded by Agent Reasoning & Action, Human-AI Interaction & Verification (featuring the three HFE patterns: Visual Provenance, Progressive Disclosure, and Attestation Interlocks), and Organizational Learning & Governance at the outermost layer.

The figure depicts four concentric layers representing the key components of your framework:

  1. Data & Context (innermost) – EHR data, ambient audio, patient-reported data, operational systems. The innermost layer supplies the raw inputs (EHR data, ambient audio, patient-reported data, operational systems) that agents consume to perform reasoning and execute tasks. Without high-quality, well-integrated data, agents cannot function reliably.
  2. Agent Reasoning & Action – AI agents performing documentation, chart review, inbox triage, etc. Agents generate outputs (documentation, chart reviews, inbox triage) that must be verified and trusted by clinicians. The HFE patterns (visual provenance, progressive disclosure, attestation interlocks) are positioned here to ensure outputs are transparent, appropriately scoped, and require explicit human confirmation before affecting clinical records.
  3. Human-AI Interaction & Verification – The interaction layer produces signals about trust, errors, overrides, and user behavior that feed into the outermost governance layer. This enables continuous monitoring, policy refinement, training updates, and system improvement based on real-world performance.
  4. Organizational Learning & Governance (outermost) – monitoring, policy, training, and continuous improvement. The outermost layer establishes the rules, guardrails, and feedback mechanisms that shape how data are collected, how agents are deployed, and how human-AI interactions are designed. It ensures compliance, safety, equity, and alignment with institutional goals across the entire stack.

Limitations and future directions

Clinicians remain vulnerable to automation bias—the tendency to trust AI outputs uncritically, especially under time pressure and cannot reliably distinguish AI-generated from human-authored clinical notes, creating risk that inaccurate drafts are accepted without adequate scrutiny [22, 23]. The framework’s verification layer assumes active engagement, but real-world workflows may encourage passive “rubber-stamping,” particularly when AI accuracy is high on routine cases.

Data & Context layer inherits well-documented EHR safety challenges: confusing visual displays, inadequate alerting, interoperability gaps, and opaque defaults. AI agents that read from or write to these systems may amplify existing hazards, for example, misinterpreting ambiguous data fields or perpetuating incorrect default values. The framework does not explicitly address how agents should handle conflicting or incomplete data from multiple EHR sources.

The Human–AI Interaction layer introduces additional verification steps that may increase rather than reduce clinician burden. Reading AI-generated notes in full, checking clinical reasoning, verifying medications and doses, and confirming nothing was added or omitted require sustained attention. If verification becomes perfunctory due to time constraints, the safety benefits of the framework erode. Just as EHR alert fatigue leads clinicians to override warnings, repeated exposure to AI-generated content requiring verification may produce “verification fatigue”—a tendency to accept outputs without thorough review.

The framework’s progressive disclosure pattern helps, but does not fully resolve, the tension between transparency and information overload. In high-volume settings (e.g., inbox triage, routine documentation), the cumulative time required for verification could negate efficiency gains or create bottlenecks. The framework lacks explicit guidance when reduced oversight is acceptable versus when full review is mandatory.

Current regulatory oversight for AI in healthcare is fragmented and often limited to safety monitoring rather than effectiveness. The framework’s attestation interlocks place responsibility on clinicians, but legal liability for AI-assisted errors remains unclear. Organizations may face liability if governance structures are deemed insufficient, even when individual clinicians sign off on AI outputs.

Effective deployment requires substantial investment in training, workflow redesign, and cultural change. Poor human–computer interface design or inadequate training can considerably diminish the tool’s effectiveness, even if the underlying AI is sound. The framework’s outermost layer assumes continuous improvement but does not specify mechanisms for capturing and responding to user feedback, error reports, or near-misses.

Clinical engineering implications

For clinical engineers and health IT leaders, these findings translate into several actionable implications:

  • AI agents should be embedded via SMART on FHIR within core EHR workflows, avoiding fragmented, portal-based experiences.
  • HFE framework Visual Provenance, Progressive Disclosure, Bi-Directional Grounding, and Attestation Interlocks, must be designed up front, not retrofitted.
  • Governance structures must include continuous monitoring of hallucination and omission profiles, coupled with feedback loops for model refinement.
  • Audit logs and interaction analytics should be used to identify friction points, measure verification burden, and iteratively optimize workflows.

These implications are tightly aligned with the mission of AI-focused clinical journals that seek evidence-based strategies for responsible, effective AI integration.

Conclusions▴Top 

AI agents—including ambient documentation scribes, Microsoft Copilot for Healthcare, ChatGPT-based systems, and clinician chatbots—have the potential to transform EHRs from passive documentation repositories into active clinical copilots. When integrated via SMART on FHIR and governed by robust HFE frameworks, these agents can significantly reduce documentation and administrative burden while supporting safer, more efficient clinical workflows.

Clinical engineers play a pivotal role in designing the integration architectures, interaction models, and safety interlocks that determine whether AI agent deployment translates into real improvements in clinician experience, operational throughput, and patient care.

Acknowledgments

No specific individuals require acknowledgment.

Financial Disclosure

This study received no external funding.

Conflict of Interest

The authors declare no conflicts of interest related to this work, including no financial relationships with vendors such as Microsoft, Nuance, Abridge, OpenAI, or other AI or EHR companies.

Author Contributions

Supreet Kaur: conceptualization, study design, data extraction, analysis, drafting of the manuscript, critical revision for important intellectual content, HFE framework development, and interpretation of findings. Andrew Blanford: clinical engineering contextualization and manuscript revision.

Data Availability

This review is based on published literature identified through PubMed/MEDLINE, Embase, IEEE Xplore, and Scopus, and on publicly available clinical reporting guidelines (e.g., PRISMA 2020). All data extracted from included studies are available from the corresponding author on reasonable request.


References▴Top 
  1. Abridge, Linked evidence architecture for ambient clinical documentation and verification. Abridge AI Inc. 2024.
  2. Biro E, Kernberg M, Patel R. Accuracy and error distribution in ambient digital scribes during simulated outpatient encounters. Journal of the American Medical Informatics Association. 2025;32:512.
  3. Cai T, Sinsky CA, Melnick ER. Measurement of cognitive workload in electronic health record interactions: a systematic review. Applied Ergonomics. 2023;108:103950.
  4. Dymek C, Atreja A, Ratwani RM. EHR usability and physician burnout: a human factors engineering evaluation of clinical information displays. Journal of Healthcare Informatics Research. 2021;5:289-302.
  5. Gardner RL, Cooper E, Haskell J, Harris DA, Poplau S, Kroth PJ, Linzer M. Physician stress and burnout: the impact of health information technology. J Am Med Inform Assoc. 2019;26(2):106-114.
    doi pubmed
  6. Gellert GA. Medical scribes: symptom or cause of impeded evolution of a transformative artificial intelligence in the electronic health record? Perspect Health Inf Manag. 2023;20(1):1d.
    pubmed
  7. Hudson TJ, Albrecht M, Smith TR, Ator GA, Thompson JA, Shah T, Shanks D. Impact of ambient artificial intelligence documentation on cognitive load. Mayo Clin Proc Digit Health. 2025;3(1):100193.
    doi pubmed
  8. Hundal J, Jain M, McCollom J. Ambient artificial intelligence in health care documentation: A review of tools, integration, and clinical implications. Journal of Oncology Informatics & Digital Health. 2025;4:88-97.
  9. Kernberg M, Biro E, Shah S. Hallucination and omission profiles of large language models in generative clinical documentation. NPJ Digital Medicine. 2024;7:142.
  10. Laiteerapong N, Shah S, Rotenstein LS. The effect of ambient AI scribes on clinician well-being and electronic health record interaction. JAMA Network Open. 2025;8:e2541209.
  11. Leung TI, Coristine AJ, Benis A. AI scribes in health care: balancing transformative potential with responsible integration. JMIR Med Inform. 2025;13:e80898.
    doi pubmed
  12. Melnick ER, Dyrbye LN, Sinsky CA, Trockel M, West CP, Nedelec L, Tutty MA, et al. The association between perceived electronic health record usability and professional burnout among US physicians. Mayo Clin Proc. 2020;95(3):476-487.
    doi pubmed
  13. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, Shamseer L, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.
    doi pubmed
  14. Ratwani RM, Savage E, Will A, Arnold R, Khairat S, Krevat S, Benda N. A usability and safety analysis of electronic health record system overrides. Journal of the American Medical Informatics Association. 2019;26:34-40.
  15. Razaghi M, Hafez A, Farina JM, Scalia IG, Pereyra M, Abdelfattah FE, Sheashaa H, et al. Transforming clinical documentation with ambient artificial intelligence (AI) scribes: a narrative review of technology, impact, and implementation. Cardiovasc Diagn Ther. 2026;16(1):11.
    doi pubmed
  16. Rotenstein LS, Holmgren AJ, Sinsky CA. EHR documentation time and physician burnout: A cross-sectional study of ambulatory primary care. Journal of General Internal Medicine. 2022;37:1420-1427.
  17. Shanafelt TD, Dyrbye LN, Sinsky C, Hasan O, Satele D, Sloan J, West CP. Relationship between clerical burden and characteristics of the electronic environment with physician burnout and professional satisfaction. Mayo Clin Proc. 2016;91(7):836-848.
    doi pubmed
  18. Sinsky C, Colligan L, Li L, Prgomet M, Reynolds S, Goeders L, Westbrook J, et al. Allocation of physician time in ambulatory practice: a time and motion study in 4 specialties. Ann Intern Med. 2016;165(11):753-760.
    doi pubmed
  19. Stults CD, Deng S, Martinez MC, Wilcox J, Szwerinski N, Chen KH, Driscoll S, et al. Evaluation of an ambient artificial intelligence documentation platform for clinicians. JAMA Netw Open. 2025;8(5):e258614.
    doi pubmed
  20. Tierney AA, Gayre G, Hoberman B, Long X, Stearns F. Ambient artificial intelligence scribes to reduce administrative burden and professional burnout in ambulatory care. Journal of General Internal Medicine. 2024;39:1345-1352.
  21. Wesley DB, Blumenthal J, Shah S, Littlejohn RA, Pruitt Z, Dixit R, Hsiao CJ, et al. A novel application of SMART on FHIR architecture for interoperable and scalable integration of patient-reported outcome data with electronic health records. J Am Med Inform Assoc. 2021;28(10):2220-2225.
    doi pubmed
  22. Westbrook JI, Woods A, Rob MI, Dunsmuir WT, Day RO. Association of interruption rates and severity of prescribing errors in hospital inpatient settings. AMA Archives of Internal Medicine. 2010;170:857-863.
  23. Zheng K, Ratwani RM, Adler-Milstein J. Studying clinician interaction with electronic health records using audit log data: Path forward and review of literature. Journal of Medical Internet Research. 2020;22:e21583.


This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, including commercial use, provided the original work is properly cited.


AI in Clinical Medicine is published by Elmer Press Inc.