Supervision cockpit for a healthcare AI — where 50% of the product effort stays invisible.

Type

Case study

Timeframe

4 days

Toolkit

Gemini (3 triage agents) · Lovable · Supabase

Year

2026

Problem

Alan rolled out Mo, its AI health assistant, to 600,000 members in March 2026. Behind the assistant, a medical vigilance infrastructure absorbed up to 50% of the team's effort: roughly 15% of conversations get reviewed by an actual doctor in under 15 minutes, up to 2,000 reviews a day at launch. This system has a user nobody shows: the reviewing physician. I designed their tool. At 2,000 reviews a day, reviewing AI conversations can turn into moderation work: fast, tedious, fallible. My hypothesis: medical vigilance can be designed as a medical act, not a ticket queue.

Solution

Triage thinks on the doctor's behalf. Priority = risk × time left before the 15-minute deadline. Always the first case on the list. The sorting rule is spelled out in plain text: invisible prioritization destroys the trust of whoever is subject to it. Signals carry a polarity. The signal that triggers escalation is red and surfaces first, reassuring signals are green. The doctor answers the only question that matters with a single glance: why is this case in front of me? Co-signature instead of correction. The doctor's input doesn't overwrite the AI's answer: it becomes a co-signed message. Correcting a machine 2,000 times a day is menial work; co-signing a response is still medicine. In the background, small decisions against fatigue: single-case focus mode, desaturated badges, a 2-second breathing pause after each review. The interface never pushes toward the next case.

Supervising a healthcare AI, without burning out the people who supervise it.

Role: Product Designer (self-initiated, solo project) — Duration: 4 days — Year: 2026 — Stack: Claude Design · Lovable · Supabase · Gemini (3 triage agents).

The context

Alan rolled out Mo, its AI health assistant, to 600,000 members in March 2026. Behind the assistant, a medical vigilance infrastructure absorbed up to 50% of the team's effort: roughly 15% of conversations get reviewed by an actual doctor in under 15 minutes, up to 2,000 reviews a day at launch.

This system has a user nobody shows: the reviewing physician. I designed their tool.

The problem

At 2,000 reviews a day, reviewing AI conversations can turn into moderation work: fast, tedious, fallible. My hypothesis: medical vigilance can be designed as a medical act, not a ticket queue.

Three decisions

Triage thinks on the doctor's behalf. Priority = risk × time left before the 15-minute deadline. Always the first case on the list. The sorting rule is spelled out in plain text: invisible prioritization destroys the trust of whoever is subject to it.

Signals carry a polarity. The signal that triggers escalation is red and surfaces first, reassuring signals are green. The doctor answers the only question that matters with a single glance: why is this case in front of me?

Co-signature instead of correction. The doctor's input doesn't overwrite the AI's answer: it becomes a co-signed message. Correcting a machine 2,000 times a day is menial work; co-signing a response is still medicine.

In the background, small decisions against fatigue: single-case focus mode, desaturated badges, a 2-second breathing pause after each review. The interface never pushes toward the next case.

The Patient Cards Lab

The third screen tests triage against AI-simulated patients: a persona, a scenario, and a hidden truth the patient only reveals if the assistant asks the right questions. The conversation runs through the real triage pipeline; the verdict compares the outcome obtained against the expected outcome.

What the prototype taught me

Simulated patients overact. The prompts had to be constrained to recover the banality of a real patient message — the kind that makes triage genuinely hard.

Fiction fills in the gaps. Give it an incomplete card, and the model invents plausible, coherent, false details. For an evaluation set, that's silent contamination.

The tool corrected me. On one card, my expected outcome was wrong: the triage had scored the real conversation, I had scored the full scenario. I fixed my calibration set, not the pipeline.

Limitations

Built without a doctor, calibrated against a public figure (15% escalation rate). This is a design exercise, not a product. The hardest question stays open: does a "Nothing to flag" button reachable in 2 seconds encourage vigilance, or reflexive clicking?

Result

Working prototype live: 3 screens, a real triage pipeline (3 Gemini agents), playable simulated patients. Built in 4 days. The full breakdown of decisions and findings is in the Medium article.

Try the demo: care-cockpit.ibnoudiop.com

Read the article on Medium

4 min read

Let’s talk

© 2026 Ibnou Diop