Case Study

Explainable AI

Designed an explainability layer for (clinical) ML models that translated opaque risk scores into readable reasoning, significantly increasing willingness to act on model outputs.

Explainable AI
Explainable AI
01

Context

We had a collaboration with the NHS Trust in Birmingham and Solihull to predict mental health crisis in patients. After a first round of deployment, we found that the value the system delivered was not in line with what the technology made possible.

02

Problem

The mental health crisis prediction model was technically strong but not being acted on by clinicians. NHS staff were trained to reason from evidence — they wanted to understand why a risk score was generated, not just what it said. A model that output a score without explaining its basis would either be ignored or followed without appropriate clinical judgement. This was both an adoption problem and a regulatory risk.

In practice, clinicians went through the list of model outputs and manually scored each patient and situation against their own set of understood criteria, because the reasoning behind why someone had been flagged at risk was not clear to them. This created additional manual work and placed extra responsibility on the clinicians — the opposite of the intended outcome. It got in the way of saving them time and prevented them from trusting the system. Based on clinician feedback, we redesigned how results were presented so that clinicians could understand the reasoning of the ML model in their own language and within their existing clinical protocols.

03

My role

I owned the product and UX design of the explainability layer. I personally led the user research with clinicians to understand their reasoning process, and translated the AI team's technical approaches into interfaces calibrated to clinical decision-making. The AI team's selection of the Trepan method was delegated; my job was to design around it.

04

Approach

I ran user research with psychiatrists and nurses to understand how they reason about risk, what level of detail was useful versus overwhelming, and what would make them willing to act on a model output. I designed the interface to surface the top contributing factors in plain language (e.g., "appointment non-attendance in the last 6 weeks", "recent prescription change") alongside the risk score, so clinicians could weigh model reasoning against their own knowledge.

05

Outcome

In user research with clinicians, we found that those given access to predictions plus explanations were significantly more willing to act on model outputs than those who received a risk score alone. The explainability layer was a condition of clinical adoption and directly enabled the rollout of the crisis prediction service in NHS Birmingham. The project reinforced a key finding: AI adoption in clinical settings is primarily a trust problem, not a technical one.

Next Project

Perspectives

Perspectives