Writing

Essay

From Emerge to clinical AI: designing information around clinical work

Bringing the record together is useful. Helping a care team recognize what changed, what is missing, and what needs action is the harder design problem.

Contents

Consider a hypothetical review before ICU rounds. The chart shows an acceptable blood pressure, but the bedside nurse reports that the patient needed more medication overnight to maintain it. The physician has to distinguish improvement from stability sustained by greater support. A family member has noticed unusual confusion. A consultant has recommended a treatment change. Another patient needs urgent attention.

The team needs to understand what happened, decide what matters now, and coordinate what comes next. Access to each piece of information is necessary. So is seeing the relationships among them.

An electronic health record can hold the relevant observations while still leaving clinicians to reconstruct the situation. That is the information-design problem that interests me in Emerge, an ICU safety platform developed with clinicians, systems engineers, and human factors engineers. It is also the problem I want us to carry forward as AI enters clinical work.

Before the decision comes reconstruction

Critical care is interdisciplinary. Bedside nurses, physicians, advanced practice providers, pharmacists, respiratory therapists, and other professionals contribute different observations and responsibilities. The record helps them share those contributions, but the location and presentation of information can make that work harder.

In a qualitative study of 25 physicians in one medical ICU, Khairat and colleagues documented difficulties retrieving high-priority information from a commercial EHR. Participants described reconstructing overnight events from scattered documentation and uncertainty about whether a laboratory test had been ordered, collected, or was still being processed. One physician explained how hourly observations could obscure an intervening heart-rate spike whose context appeared elsewhere in nursing documentation. 1

These are examples of a team trying to recover sequence and status before deciding whether to act or wait. A current value, a treatment change, and an unfinished task answer different questions. Putting them in the same record does not necessarily make their relationship visible.

Emerge began with the work of preventing harm

Emerge approached this as a problem involving people, tasks, information, and technology. Its Concept of Operations, or ConOps, described how prevention should work before that account was translated into software requirements.

Romig and colleagues organized the ConOps around four components: risk assessment, appropriate therapies, monitoring and feedback, and communication with patients and families. The team used that framework to describe the tasks and workflows involved in preventing several kinds of ICU harm. 2

That sequence matters to me. It gives the interface an operational purpose. A care-status display can help the team answer whether the relevant preventive work has happened and where attention is still needed. It also forces the designers to decide which observations support that answer.

From the work to the informationFour questions before designing a display
What could harm this patient?
Risk assessmentRelevant observations and their timing
What preventive care is appropriate?
Appropriate therapiesThe plan and evidence of delivery
What happened, and what is still open?
Monitoring and feedbackChanges, incomplete work, and missing information
What do the patient and family understand?
CommunicationGoals, concerns, and the shared plan

Author’s interpretation of the four ConOps components described by Romig et al. (2018). This is a design map, not a clinical protocol or a reproduction of Emerge.

The difficult part is the connection between the display and the work it represents. A blank field might mean an action was not performed, was not documented, or has not reached the display yet. Those possibilities require different responses. A useful overview has to preserve uncertainty well enough for the team to investigate it.

What the studies actually measured

The Emerge pilot provides a useful, bounded test of information retrieval. Eighteen clinicians completed the same tasks using Emerge and the existing EHR, with the order randomized. Testing used a stored, de-identified historical patient dataset in a controlled setting. The study measured retrieval time, clicks, and correct identification of 28 data elements. 3

Controlled retrieval pilot · 18 cliniciansLess searching. More correct information.

Retrieval time

seconds · lower is better

Existing EHR1,098
Emerge481

Scale: 0 to 1,200 seconds

Navigation

clicks · lower is better

Existing EHR96
Emerge38

Scale: 0 to 100 clicks

Accuracy

correct elements · higher is better

Existing EHR22 / 28
Emerge25 / 28

Scale: 0 to 28 correct elements

Medians from Romig et al. (2019), Results. Each measure has its own zero-based scale. Bars show medians only, not variability.

Same tasks and historical patient dataset; order randomized. These are information-retrieval results, not patient outcomes.

Those results support a claim about performance on the tested tasks. They do not establish that the display improved patient outcomes, and they do not tell us how a contemporary AI interface would perform. The small sample and controlled setting also limit generalization. The practical insight is that information presentation can change the effort and accuracy of work even when clinicians already know the underlying record system.

Related work at Mayo Clinic gives the argument another setting. In a randomized crossover simulation with 20 critical-care physicians, Ahmed and colleagues compared a task-focused interface with a standard record interface. Median NASA Task Load Index scores, a measure of perceived workload, were 38.8 with the redesigned interface and 58 with the standard interface. 4

A later pilot stepped-wedge cluster-randomized study evaluated the AWARE viewer during routine work across four ICUs. Median pre-round data-gathering time decreased from 12 to 9 minutes per patient, and participants reported less difficulty and mental demand. The study included 80 observations before implementation and 63 afterward. 5

A controlled retrieval exercise, a simulation, and a study in routine clinical work answer different questions. Taken together, they give us reasons to test task-focused information design. They are not a pooled estimate of patient benefit.

A focused interface has a different design scope

A comprehensive EHR serves many professions and purposes. A specialized ICU display can concentrate on a narrower question, such as whether the patient has received appropriate preventive care. That difference in scope helps explain why an additional view may be useful without replacing the record beneath it.

It would be too broad to infer from these studies that EHR vendors ignored clinicians. Ratwani and colleagues’ study of 11 vendors found variation in user-centered design practices, along with barriers to observing clinical work and recruiting suitable participants. 6

For me, the lesson is to connect observation of work with explicit requirements. Who needs the information? What must they recognize? What can they safely infer from the display, and what must they verify? How will the team know that a task remains unfinished?

Human-centered design becomes more useful when those questions are answered together with the people responsible for the care process.

Clinical AI should carry that lesson forward

The September 1, 2026 announcement of an Epic integration with ChatGPT for Healthcare brought this concern into focus for me. OpenAI describes both bringing authorized EHR context into ChatGPT and integrating ChatGPT into supported EHR layouts. Its announcement also describes pilot-partner work with frontline teams to validate the capabilities in practice. 7

I welcome the direction. In my own development work with Codex and ChatGPT, I have seen the potential to generate interactive charts and visual interfaces. That makes me more interested in how clinical information could be presented, not satisfied merely because it can be summarized.

The text-first demonstration discussed in my original essay gave me pause. Its pre-visit summary distributed symptoms, laboratory results, medications, and follow-up information across a long response. Chart citations offered a route back to supporting information. The clinician still had to read the sections, compare their contents, and work out what deserved attention.

That is my critique of the demonstrated presentation, not a finding from a clinical evaluation or a claim about every supported deployment. Emerge and the demonstration also address different tasks and settings. A fair comparison would require clinicians performing comparable work with the relevant alternatives.

Still, I regard a text-first review that requires repeated prompting and rereading as a step backward from the recognition that made Emerge compelling. Conversation can be useful for explanation and follow-up. A predictable overview should help clinicians find important changes and unfinished work before they have to ask for them.

Make change, evidence, and unfinished work easy to find

The standard I would apply is practical. Across patients and repeated reviews, can clinicians find the same kinds of information in predictable places? Can they distinguish a new finding from an old one? Can they see what remains uncertain or incomplete?

A citation needs more than a clickable marker. The interaction should make clear which observation it supports, when that observation was made, and whether the clinician is looking at the original record or an interpretation. If several laboratory results share a citation, the relevant result and time point should remain obvious when the clinician asks a follow-up question.

Visual structure also needs scrutiny. A compact display can hide context, and a reassuring status can conceal missing data. I would want evaluation to include whether users recognize omissions and conflicting evidence, how long verification takes, and whether another team member can understand what remains to be done. Retrieval speed alone cannot answer those questions.

The design should support both overview and inquiry: a stable view of the patient’s changing situation, with conversation available to explain, investigate, and clarify. That is a design proposal to test with clinicians, rather than an outcome established by the earlier studies.

The standard is better clinical work

Emerge’s most useful lesson is that access to data is one part of a larger responsibility. Systems engineering made the intended work explicit. Human factors engineering helped translate it into information people could use.

For healthcare leaders evaluating AI, I would begin in the same place: define the clinical work and responsibilities, design with the people who perform that work, and test whether the tool helps them understand and act without adding new burdens.

I remain optimistic about AI’s potential. The question that will keep me interested is whether it helps a care team recognize what has changed, understand the evidence, and see what still needs attention. An impressive answer is a starting point. The work around that answer is where the design has to prove useful.

Notes

  1. Khairat S, et al. “Physician experiences of screen-level features in a prominent electronic health record” (2021). Study. Qualitative interviews with 25 physicians at one academic medical ICU; participant experiences are not a prevalence estimate for all EHR users.
  2. Romig M, et al. “Developing a comprehensive model of intensive care unit processes: Concept of operations” (2018). Study.
  3. Romig M, et al. “Intensive care unit providers more quickly and accurately assess risk of multiple harms using an engineered safety display” (2019). Study. The graphic uses medians from the Results section; the abstract reports means. This pilot did not test patient outcomes.
  4. Ahmed A, et al. “The effect of two different electronic health record user interfaces on intensive care provider task load, errors of cognition, and performance” (2011). Study.
  5. Pickering BW, et al. “The implementation of clinician designed, human-centered electronic medical record viewer in the intensive care unit” (2015). Study. Pilot stepped-wedge cluster-randomized trial in four ICUs at an academic referral center.
  6. Ratwani RM, et al. “Electronic health record usability: Analysis of the user-centered design processes of eleven electronic health record vendors” (2015). Study.
  7. OpenAI. “Healthcare organizations can now connect EHR and additional industry data to ChatGPT” (September 1, 2026). Announcement. A first-party account of capabilities and evaluations, not an independent comparison with Emerge. The interface critique concerns the demonstration documented in my original essay.
About the illustration
Layered paper records align along a shared timeline in front of a hospital bed and care team; a terracotta loop marks unfinished work.
AI-generated conceptual illustration of records organized around clinical work; not a clinical interface or patient record.

Back to top