Artificial intelligence is already being used to spot signs of disease in scans, predict patient outcomes, and support decisions that affect people’s lives. But there’s a persistent problem: many AI systems don’t really explain how they reach their conclusions. Even systems that look transparent might secretly be hiding information from the doctors and experts relying on them. Now, researchers at King’s College London have built a mathematical framework to catch exactly this kind of hidden behaviour.
The study, published in the Journal of Machine Learning Research, offers a way to test whether AI systems that appear to explain their decisions really are being honest about them. This matters more than ever, as healthcare and other high-stakes fields come under pressure to make AI more open to scrutiny.
The “Black Box” Problem
AI systems can be impressively accurate. They can read medical images, predict which patients are at risk of complications, and much more. But many work as so-called black boxes: they spit out an answer without showing their working.
For a clinician, that’s a serious issue. If an AI says a scan looks suspicious, the doctor needs to know why. Is the system picking up on something real, or has it latched onto an irrelevant detail? Without that insight, it’s hard to spot mistakes, correct them, or push back when something looks off. And with new regulations around the world pushing for greater transparency and human oversight of AI, this isn’t just a technical issue, it’s a legal and ethical one too.
Concept-Based AI: A Partial Solution
One attempt to make AI more understandable is to build “concept-based” models. Instead of just crunching raw data, these systems base their decisions on clearly meaningful ideas, such as blood pressure, the presence of a fever, tumour size, or specific abnormalities in a medical image.
In theory, this should let a clinician look at the AI’s output and see which concepts it relied on, making it easier to review and challenge decisions. But there’s a catch, called “information leakage.” Sometimes, those concepts quietly carry extra information that doesn’t show up on the surface. The AI appears to be reasoning in a transparent way, but behind the scenes, it may still be relying on data the user can’t see or examine. In other words, it can look like an open book while really still being a black box in disguise.
A Way to Measure the Hidden Information
The team at King’s College together with colleagues at the Alan Turing Institute and with support from the Turing-Roche Strategic Partnership set out to build a mathematical framework that can detect and measure this leakage.
Their framework defines two things to look for:
- Concepts-task leakage (CTL): hidden information linked to the final prediction
- Interconcept leakage (ICL): hidden information shared between the different concepts
They tested the framework across several datasets and found it could reliably detect leakage and predict how the models would behave when their concepts were deliberately tweaked. That’s an important sign that the tool really does capture what it claims to.
Why This Matters for Patients
The research gives developers practical guidance on how to design concept-based models with less leakage, so that their transparency is real and not just skin-deep.
Dr Christopher Banerji, AI+ Senior Fellow (Clinical-Academic) and senior author of the paper, said: “Rather than putting the cart before the horse, while the field of AI is moving quickly to apply models to real-world problems, we have taken a step back to consider what is needed to make these systems safe and reliable. Our work has focused on understanding limitations and addressing them, so that we can move towards deploying these models in clinical practice in a safer way.”
Dr Enrico Parisini, Senior Research Fellow in Machine Learning and first author of the paper, said: “Concept-based AI has the potential to make AI systems more transparent, but our research shows that models can appear interpretable while still relying on information that is hidden from the person using them. By identifying and measuring this hidden information, we can take steps towards developing AI systems that are more transparent.”
Looking Ahead
The framework provides researchers, developers, and regulators with a clear way to assess whether an AI system is transparent enough for people to meaningfully oversee it, a question that will only become more pressing as AI is embedded into hospitals, clinics, and other high-stakes settings. The team is now working to apply these approaches to real clinical problems.
If AI is going to help make life-changing medical decisions, trusting its explanations isn’t optional. This research helps make sure those explanations hold up to scrutiny.
Parisini, E., Chakraborti, T., Harbron, C., MacArthur, B.D., & Banerji, C.R.S. (2026). Leakage and Interpretability in Concept-Based Models. Journal of Machine Learning Research, 27(174), 1–40.