HealthcareB2B SaaSAI0-to-1Design Systems

AI Healthcare SaaS Platform

A 0-to-1 clinical product built on one hard problem: getting a time-starved doctor to treat an AI as a real clinical assistant. Sole owner across eight surfaces, from problem framing to a validated, handed-off design system.

Client
Arogo AI
Role
Product Design Intern (sole designer)
Timeline
3 months (Apr to Jul)
Team
2 founders, 6 engineers, 1 designer (me)
Year
2025

Snapshot

AROGO AI is a B2B SaaS platform that turns a doctor's consultation into structured clinical work automatically. The doctor talks to the patient, and the system uses conversational AI and live transcription to generate the prescription, the patient health card, and a layer of AI insights (risk factors, medication interactions, recommended tests, and more). I joined when the product was an idea and a rough thesis. As the only designer, I owned the full surface area: onboarding, the multi-hospital workspace model, the conversational AI consultation screen, the AI-generated health card, the operations dashboard, the financial layer, and patient communication. The central design problem was not "make it look good." It was "make a time-starved doctor trust an AI with clinical work, and give back the minutes the paperwork currently steals."

I delivered the complete end-to-end design and handed it off to engineering, who carried it into development after I left. Before development, I validated the core experience (the medical health cards, onboarding, and the in-app consultation flow) with 10 practicing doctors through prototype usability testing.


Context: what AROGO was solving

Practicing doctors, especially in India, lose a large share of their day to work that is not medicine. They write notes during and after every consultation, hand-write or re-type prescriptions, chase patients for follow-ups, and reconcile cash, UPI, and card payments at the end of the day. Many of them run clinics across more than one hospital, so they constantly switch contexts: different patient lists, different schedules, different billing.

AROGO's bet was that conversational AI could absorb most of this. If the system listens to the consultation and produces the clinical artifacts, the doctor spends their scarce attention on the patient instead of on the keyboard.

That bet creates the hardest design constraint in the whole product: clinical trust. A prescription or a risk flag is not a low-stakes UI suggestion. A doctor cannot ship something they did not see, and the platform cannot let just anyone create a clinical workspace. Every design decision downstream answered to that constraint.

Primary User01

The Doctor

A practicing physician, often working across multiple hospitals, chronically short on time, personally liable for every clinical output. Needs the system to earn trust before it earns usage.

Secondary User02

Front Desk & Clinical Staff

Administrative and clinical staff managing scheduling, patient intake, billing, and communication. High-volume, process-driven work where errors have real downstream effects.


My role and what I actually owned

I was the sole designer on a 0-to-1 product, so "end-to-end" means I owned the problem framing, the user flows, the information architecture, the interaction design, the visual system, and the handoff. I worked directly with both founders on product direction, with the six-person engineering team on feasibility and handoff, and I owned the usability testing side of the project. Being honest about this matters more than it looks: claiming you single-handedly built an entire product invites skepticism, while owning a clearly-scoped surface and naming your collaborators reads as senior.

What I owned:

01

End-to-end information architecture and navigation model

02

Eight core product surfaces, lo-fi flows to hi-fi screens

03

Trust and verification model for clinical authenticity

04

Prototyping and usability testing with 10 practicing doctors

05

Engineering handoff

06

Design system: type scale, color roles, components and states


How I approached it

I did not run the primary user interviews myself; the founders curated the domain research and the understanding of the doctor's day, which is common at an early-stage company where the founders carry the domain expertise. Where I owned the user signal was validation. I built interactive prototypes and tested them with 10 practicing doctors, focusing on the parts of the product where being wrong was most expensive: the medical health cards, the onboarding and verification flow, and the consultation experience. That testing is what turned founder hypotheses into design decisions I could defend.

My working sequence was roughly: understand the doctor's day and where the time leaks, map the jobs each user is trying to get done, identify the riskiest surface (the AI consultation), and design outward from that risk rather than starting with the easy screens. I treated the conversational AI as the spine of the product and let every other surface support it.

A principle I held throughout: reduce visual and cognitive noise. A doctor between patients has no spare attention. So I consistently chose general, calm, explanatory interfaces over dense, decorated, or over-personalized ones, and I isolated complexity to the moments where it earned its place.

01UnderstandThe doctor's day and where the time leaks
02MapThe jobs each user is trying to get done
03Identify riskThe riskiest surface: the AI consultation
04Design outwardFrom risk, not from the easy screens
Held throughout

Reduce visual and cognitive noise. A doctor between patients has no spare attention.


The eight surfaces, framed as decisions

I will walk through the product the way I designed it, leading with the decision and the tradeoff in each, not just the screen.

01Auth & OnboardingTwo paths. PIN for fast re-entry.
02Workspace & VerificationGate authenticity without killing momentum
03AI ConsultationReview chain, not black box
04AI Health CardHierarchy by clinical urgency
05DashboardWhat needs me right now
06Financial OverviewRevenue as insight, not ledger
07CommunicationPatient contact inside the platform
08Design SystemToken architecture for a solo-to-six handoff

1. Authentication and onboarding: lower the entry cost, raise the security floor

Decision: Keep sign-in and sign-up to two paths, Google Auth or email, and add a numeric PIN as a second layer for returning sessions.

Why: Doctors open the app many times a day between patients. Forcing a full password every time is friction that compounds into abandonment, but a clinical product cannot drop authentication entirely. The PIN resolves the tension. The doctor authenticates fully once, then re-enters with a fast PIN, which keeps the security floor high without taxing the dozens of daily re-entries.

Tradeoff I accepted: a PIN is weaker than a full re-auth on a shared or stolen device. I treated it as the right call for a single-owner professional device and mitigated the risk with a weekly full re-authentication that resets the trusted session, plus device binding so the PIN only works on the doctor's registered device. That combination keeps daily friction near zero while closing the obvious attack surface: a stolen PIN on an unrecognized device gets nowhere, and the weekly reset caps how long any single authenticated session can persist.

2. Workspace creation and verification: gate authenticity without killing momentum

Decision: Require doctors to upload personal identity and degree or license documents during workspace creation, route them to admin for verification, and unlock the service only after approval.

Why: In healthcare, the integrity of the platform depends on every practitioner being real and licensed. Unverified clinical workspaces are a patient-safety and legal liability, not a growth-funnel inconvenience.

The actual design problem: verification takes time, and a blank "pending approval" wall is where new users churn. So I designed the wait deliberately. The doctor completes their full profile setup and explores the product in an empty preview state while verification runs in the background, with honest status copy and an expected timeframe rather than an indefinite spinner. When the admin approves the workspace, a notification tells the doctor they are live and can start seeing patients. The principle: be honest about the delay and let the user keep moving wherever it is safe to do so, instead of locking the entire experience behind one blocking gate.

3. The conversational AI consultation: the spine, and the hardest screen

This is the surface the entire product rises or falls on, so it got the most of my attention.

Decision: Integrate live transcription directly into the consultation screen so that the prescription, health card, and supporting clinical information generate from the conversation in near real time.

The three problems I had to solve at once:

First, attention. Live transcription is useful only if it does not pull the doctor's eyes off the patient. I designed the transcription to be present but quiet, available as confirmation rather than as a thing the doctor must watch. The patient relationship stays primary; the AI works in the periphery.

Second, trust and review. The system cannot output a prescription the doctor did not approve, so I designed the entire consultation as a chain of review gates rather than a black box. As the conversation transcribes in real time, the doctor can read it and edit any line. Crucially, every edit saves as a new version while the raw transcription stays accessible, so the doctor can correct the record without ever losing the original source of truth. The doctor can also add their own notes on top of the transcription. Only once the doctor confirms does the system generate the downstream artifacts: the prescription, and the creation or updating of the patient health card. Each of those remains editable too. The pattern is consistent across the whole flow: the AI proposes, the doctor disposes, and nothing reaches the patient without an explicit confirmation. This is the single most important interaction in the product, because it is where clinical liability and AI convenience meet.

Third, latency. AI generation is not instant, and a doctor staring at a frozen screen is a doctor losing faith in the product. So I split the AI insights into two tiers tuned to two different needs. Quick Insights generate fast, so the doctor gets the immediately useful read of the patient without waiting, in the window where they are still with the patient. Detailed Insights, the full per-section snapshot of the patient, take longer to compute, so I did not make the doctor sit and wait for them. Instead the work continues in the background, and a notification tells the doctor when the detailed report is ready. At that point the doctor can review it and send the information to the patient. This is the same honesty principle as the verification wait: rather than hide the delay behind an endless shimmer, I told the doctor what was being generated and let them keep working, isolating the wait to the section that actually needed it instead of locking the whole screen.

4. The AI health card: segregating signal from a flood of output

Decision: Take everything the AI generates (key insights, risk factors, recommendations, medication interactions, health patterns, health snapshots, and required tests) and organize it by clinical urgency rather than dumping it as one long list.

Why: The AI produces a lot. Undifferentiated, that volume is noise, and noise in a clinical tool is dangerous because the important thing hides among the routine. The card was the surface I tested most heavily with doctors, because it is where information design and patient safety overlap most directly.

The hierarchy. I structured the card around one question: what does a doctor need to see in the first five seconds versus what they pull up on demand? That sorted the seven information types into three tiers by clinical consequence.

The top tier is safety-critical information, the things that prevent immediate harm: risk factors and medication interactions. These lead the card, carry the strongest visual weight, and are never hidden behind a tap when severe. A drug interaction the doctor does not see is the worst failure the product can produce, so it gets the most prominent position on the card and the highest-contrast treatment.

The middle tier is decision-driving information, the things the doctor acts on during this visit: the key insights summary (the one-glance read of what matters about this patient), the recommendations, and the tests required. This tier sits directly below the safety zone and is structured so each item maps to a clear next action.

The bottom tier is supporting context, the longitudinal picture: health patterns and health snapshots. This is reference material the doctor pulls up when they want depth, so it lives lower on the card and uses progressive disclosure, expandable rather than always-open, to keep the default view calm.

How it was structured visually. A few specific decisions held the card together:

First, a one-glance summary band at the top. Before any detail, the card opens with the patient's health snapshot as a compact summary so the doctor orients in a second. Detail unfolds below it.

Second, a dedicated clinical-risk color scale, separate from the interface's destructive-action color. This was a deliberate call. If "delete this record" red and "this medication combination could harm the patient" red look identical, the doctor's eye learns to discount both. So clinical risk got its own severity scale (critical, warning, informational) used consistently across risk factors and interactions, distinct from the UI red reserved for destructive controls. Severity, not decoration, drives color.

Third, scannability over density. I laid the card out so urgency reads top-down and left-first, matching how a hurried eye actually moves, and I used progressive disclosure for the long-tail context so the card never opens as a wall of text. A doctor who sees everything sees nothing, so the default view shows the consequential few and lets the rest expand.

Fourth, visible AI provenance. Because every block is machine-generated, each carried a clear signal that it was AI-generated and reviewable, reinforcing the same trust contract as the consultation flow: the AI surfaces, the doctor decides.

Tradeoff: prioritizing a few items means deprioritizing others. I chose patient-safety information as the top tier and accepted that richer longitudinal context takes a second tap, because in a clinical tool the cost of burying a drug interaction is far higher than the cost of one extra tap to see a trend.

5. The dashboard: the doctor's command center

Decision: Make the dashboard answer "what needs me right now," not "here is all your data."

Why: A dashboard that simply mirrors the database forces the doctor to do the triage the product should do for them. So I designed it as a dynamic dashboard, meaning the surface reorders and adapts to what needs the doctor's attention right now rather than showing a fixed static layout. The ordering encodes a priority judgment: the items that change the doctor's next action sit highest, and reference data sits lower.

The default view reads top to bottom as a deliberate sequence:

First, key operational insights, the numbers that frame the day at a glance: total appointments and the average wait time per patient. These answer "how heavy is today and is my clinic running on time" in the first second.

Second, critical patient alerts, surfaced high because a patient who needs urgent attention cannot be one row in a list the doctor has to scroll to find. This is the highest-stakes block on the screen.

Third, AI insights and the doctor's chats, the contextual layer that supports the day's decisions and patient contact.

I also put an instant consultation entry point on the dashboard, so the doctor can start the core action of the product, the AI-assisted consultation, in one tap from the home surface rather than navigating to find it. The most-used action does not get buried.

The dashboard closes with the financial overview and the patient satisfaction rate. These are the practice-health signals: how the business is doing and how patients feel about the care. They live at the bottom because they are review-cadence information, checked periodically rather than acted on between patients, so they earn a place on the default view without competing with the urgent items above them.

The principle throughout: urgency and frequency-of-use decide vertical position. What is time-sensitive or used constantly leads; what is reviewed occasionally follows.

The multi-hospital reality lives here too. A doctor practicing across hospitals carries a real risk of acting in the wrong context, ordering, prescribing, or billing against the wrong facility, which is one of the more dangerous quiet failures in the product. So the active hospital is always visible and unambiguous, never inferred silently. Switching it is fast and explicit: the doctor changes context deliberately, and the dashboard reloads to that hospital's appointments, patients, and financials so the data on screen always matches the place the doctor is working. Making the switch a clear, confirmed action rather than a passive background state is what keeps context errors from happening in the first place.

6. The financial overview: turn payment data into decisions

Decision: Present revenue as insight, not as a ledger.

Why: The raw facts (consultation fees collected, the split across cash, UPI, and other methods, which patients have pending dues, and the invoices generated) are only useful if they answer a question the doctor or their staff actually has: how am I doing, who owes me, and is the practice growing?

I structured the surface using the standard analytics pattern of overview-first, detail-on-demand, so a glance answers the headline question and a tap reveals the breakdown:

Summary cards at the top carry the headline figures: total revenue, the payment-method split across cash, UPI, and other methods, and outstanding dues. These give the one-glance read. Putting the method split up front matters in the Indian context specifically, where cash and UPI behave very differently for reconciliation, so the doctor or front-desk staff can see at a glance how much of the day's revenue is already settled digitally versus sitting as cash to account for.

Trend visualization sits below the cards to answer the growth question. Income and growth are shown over time so direction, not just the current number, is legible, because "is the practice growing" is a question about slope, not a single value. I kept the charts to the few comparisons that drive a decision rather than decorating the page with every metric the data could produce.

The pending-payments view turns dues into an actionable list: which patients owe, how much, and how long it has been outstanding, so chasing payment becomes a task the staff can work through rather than a number they have to investigate. This is the difference between reporting a problem and giving someone the means to resolve it.

Invoice generation closes the loop, letting the doctor or staff produce an invoice directly from a consultation or a pending balance, so billing lives in the same system as the care rather than in a separate tool.

The summarized view serves the quick daily read; the detailed view serves end-of-day or end-of-month reconciliation. Same data, two depths, matched to the two real moments the staff actually use it.

7. Communication: messaging, video, and follow-up alerts

Decision: Keep patient communication inside the platform rather than scattering it across WhatsApp and phone calls.

Why: Consolidating messaging, video consultations, and automated follow-up alerts keeps the clinical record in one place and gives the doctor one surface for patient contact, instead of scattering it across personal phone numbers and WhatsApp where nothing is logged against the patient's record.

Messaging and video I designed as a continuation of the consultation, not a separate inbox. The doctor reaches a patient from within that patient's context, so a message or a video call carries the clinical thread with it rather than starting from a blank conversation. For a telemedicine interaction, the same review-and-confirm discipline from the in-person flow applies: the consultation still produces the prescription and health card through the doctor's confirmation, so the channel changes but the trust contract does not.

Follow-up alerts were the part I treated most carefully, because automated clinical nudges sit on a knife edge. The well-documented failure mode in clinical software is alert fatigue: when a system fires too many low-value notifications, clinicians learn to dismiss all of them, and the one alert that mattered gets dismissed with the noise. So the design question was not "how do we send reminders" but "how do we send only the alerts that change behavior."

I tied follow-up alerts to genuine continuity-of-care moments rather than blanket scheduling: a follow-up the doctor explicitly set during the consultation, a recommended test that should be completed by a date, or a return visit a treatment plan depends on. The alert is composed from the clinical context of that specific patient, so it carries why the follow-up matters, not just a generic "please return." And because the highest-stakes signals (a critical patient needing attention) surface on the doctor's dashboard rather than as one more dismissible push, the alerting load stays proportional to clinical importance. Severity decides the channel and the prominence, which is the established principle for keeping clinical alerts trusted rather than tuned out.

8. The connective tissue: a design system that held it together

Designing this many surfaces solo without a system produces inconsistency fast, so the system was not a deliverable I made at the end; it was the thing that let one designer hold eight surfaces together and hand them to six engineers without ambiguity.

Color, and why purple was the right call for a clinical product. The platform's primary color was #4C0AD1, a deep saturated violet. That choice does real work in a healthcare interface. The most important rule in clinical color is that status colors must stay reserved: red has to mean danger or critical, amber has to mean caution, and green has to mean stable or resolved, because clinicians read those colors as clinical signal, not decoration. A brand color that competed with that palette would dilute the signal. Purple sits outside the red-amber-green status range entirely, so it can own brand, primary actions, and navigation without ever being confused for a clinical state. That separation is the whole point: the interface chrome and the clinical alerting never speak the same color language.

I also kept color accessible by design, not by accident. I treated WCAG AA contrast as the floor for text and interactive elements, and I never used color as the only carrier of meaning. A risk state is marked by an icon, a label, and position, not hue alone, so a color-blind clinician (a meaningful share of any doctor population) reads the same urgency everyone else does. Relying on color alone is a documented accessibility failure, and in a clinical tool it is a safety failure.

I defined the color roles by function, not by swatch:

  • Primary action (#4C0AD1 and its derived states) for the main path forward on any screen
  • Destructive for irreversible interface actions like delete, kept visually distinct from clinical red
  • Clinical-risk as its own severity scale (critical / warning / informational), the most carefully guarded part of the palette, never reused for anything cosmetic

A two-tier token architecture. I stored color in two levels, which mirrors the modern primitive-then-semantic token pattern used by mature systems. Level 1 was the raw palette: the actual color values, named by what they are. Level 2 was the design tokens that the UI actually consumed, named by what they do, and built on top of the level-1 values. So a button did not reference "violet 600"; it referenced the semantic primary-action token, which in turn pointed at the raw value. The payoff is real and worth stating to an engineer: the entire product can be re-themed, or the brand color shifted, by changing level-1 values once, and every semantic token downstream updates without touching a single component. It also makes the system legible: a developer reading primary-action knows the intent, where a raw hex tells them nothing.

Type and components. A defined type scale gave hierarchy a small, repeatable vocabulary instead of arbitrary sizes per screen, which matters most on dense surfaces like the health card and dashboard where hierarchy is doing the heavy lifting. The component library covered the shared building blocks across all eight surfaces, and critically I defined the full state set for interactive components, hover, active, disabled, loading, empty, and error, not just the happy path. Designing empty and error states explicitly is what separates a system that survives real data from one that only looks good in a portfolio mockup, and in a clinical product the loading and empty states (a health card mid-generation, a dashboard before the day's appointments load) are not edge cases; they are the experience.

The system is what let one designer keep eight surfaces coherent, and it is what made the handoff to a six-engineer team tractable.


The hardest design tension, stated plainly

Every meaningful decision in AROGO traced back to one tension: AI convenience versus clinical trust. The faster the AI acts on its own, the more time it saves and the more risk it introduces. The more the doctor must review, the safer it is and the slower it gets. My job was not to pick a side. It was to design the seam between them: make the AI do the heavy lifting, then put the doctor at the exact decision points where their judgment and liability require them, and nowhere else. That balance is the actual product.

AI Convenience
  • Faster clinical work
  • Auto-generates prescriptions
  • Surfaces risk patterns from data
  • Scales with patient volume
Clinical Trust
  • Doctor reviews every output
  • Nothing ships without confirmation
  • Liability stays with the practitioner
  • AI proposes — doctor disposes
The seam

AI does the heavy lifting. Doctor at exact decision points where judgment and liability require them, and nowhere else.


Outcome and impact

0→1Product

End-to-end design across 8 surfaces in 3 months

10Doctors

Validated core flows via interactive prototype testing

6Engineers

Solo designer handoff that held through development

In three months as the only designer, I took AROGO from an idea to a complete, validated, end-to-end product design across eight surfaces, and handed it off to a six-person engineering team that carried it into development after I left.

The impact I can point to concretely:

I de-risked the product before a line of production code was written. I built interactive prototypes and ran usability testing with 10 practicing doctors, concentrating on the three surfaces where being wrong would have been most expensive to unwind later: the AI-generated medical health cards, the onboarding and verification flow, and the in-app consultation experience. Validating the hardest, highest-stakes parts with real clinicians at the prototype stage is the cheapest possible time to find problems, and it meant engineering built against a design that doctors had already used, not against assumptions.

I produced the system that made a solo-to-six handoff work. The two-tier token architecture and the full component-and-state library were not decoration; they were the mechanism that let one designer's intent survive into six engineers' implementation without a design owner in the room afterward. A handoff that holds up after the designer has left is a real measure of whether the system was built right.

I owned the decisions that define the product's character: the review-and-confirm trust model on the consultation, the two-tier insight latency strategy, the verification preview state, and the clinical-risk color discipline. These are the parts a clinician feels, and they were design calls, not features handed to me.

(Out of respect for the company's confidentiality, I have described the thinking and the decisions rather than showing the proprietary screens or internal data.)


What I learned

01

Trust model is the product

Review-and-confirm is not a safety afterthought. It is the feature that unlocks AI value.

02

Honesty beats illusion

An honest wait with clear expectations builds more confidence than a fake-instant interface.

03

Design system is leverage

A system is how a designer's judgment scales beyond their own presence.

04

Constraints clarify

Domain constraints sharpened the design rather than limiting it.

In a clinical product, the trust model is the product. I came in thinking the hard part would be making AI output look good. The real work was designing the seams where a doctor's judgment overrides the machine: the editable transcription with preserved raw versions, the confirmation gate before any prescription generates, the explicit review on every AI block. I learned that in a high-stakes domain, the review-and-confirm interaction is not a safety afterthought bolted onto a smart feature; it is the feature. The AI's value is only unlocked by the doctor trusting it, and trust is built in those small moments of control.

Honesty beats illusion when the system is slow. The latency problem taught me something I now apply everywhere. Faced with AI generation that genuinely takes time, the tempting move is a shimmer that pretends nothing is loading. The better move was to tell the doctor the truth, split insights into a fast tier and a deferred tier, and notify them when the slow part was ready. An honest wait with a clear expectation builds more confidence than a fake-instant interface that eventually stalls. I would rather a user trust the system than be briefly impressed by it.

A design system is leverage, not polish, and the proof is what happens after you leave. Building the two-tier token structure and full state coverage early felt like overhead at the time. Its value only became fully clear at handoff: I left, and the design still had to function as the single source of truth for six engineers with no designer to ask. Consistency does not survive eight surfaces on willpower. The system was what made my three months keep paying out after I was gone, and that reframed how I think about systems: they are how a designer's judgment scales beyond their own presence.

Constraints clarify, they do not limit. Reserving the status palette for clinical signal forced the brand into purple, and that constraint produced a cleaner result than an unconstrained choice would have. I learned to treat domain constraints (clinical color rules, accessibility floors, liability requirements) as the thing that sharpens the design rather than the thing that gets in its way.


What I would do differently

Two things, both honest.

If I did it again — 01

Own discovery, not just validation

Sit in on primary conversations. Designing from a problem you heard firsthand beats one you were handed.

If I did it again — 02

Pressure-test context switching

Validate that doctors always knew which hospital they were in. It is exactly the kind of quiet failure a designer cannot catch from their own desk.

I would own discovery, not just validation. On this project the founders curated the user understanding and I owned the testing of solutions against it. That division is normal early-stage reality, and validating with 10 doctors was real rigor. But the closer a designer sits to the original problem definition, the better the solutions, so given the chance again I would push to sit in on the primary conversations with doctors myself rather than inheriting the framing. Designing from a problem you heard firsthand beats designing from a problem you were handed.

I would have put the multi-hospital context switching in front of doctors. My usability testing concentrated on the health cards, onboarding, and consultation, the right priorities given the stakes and the time. But acting in the wrong hospital's context is one of the most dangerous quiet failures the product can produce, and I validated the switching model through reasoning rather than testing. With more runway I would have specifically pressure-tested whether doctors always knew which hospital they were operating in, because that is exactly the kind of error a designer cannot catch from their own desk.