
24th Nov, 2025
3 min read
AI Safety
AI is now even more prevalent in healthcare, from diagnoses to risk assessment, lab predictions and even treatment recommendations. As AI integrates further in Clinical Decision Support Systems (CDSS) [2], it’s becoming critical to evaluate the explainability requirements for AI’s safe and usable deployment.
When we begin training AI, human evaluators go through a Supervised Fine Training (SFT) process[1] to lay the basic groundwork of AI’s thought process and alignment with human values. It’s safe to say that in a mission-critical application such as healthcare, this stage of training is more rigorous and regulated to ensure there are no gaps or misalignments.

At this stage of training, a physician may train AI in several aspects of medicine, inclusive of the correct groundwork for diagnoses, symptom checks, treatment recommendations and how to identify patterns in lab results, which can be further fine-tuned into better predictions via Feedback.
AI during Training — What’s missing?
During training, AI focuses primarily on output alignment with human values. In healthcare, general treatment and guidance require patient collaboration, where a patient’s values and preferences directly correlate to the treatment they would opt for, and the trust they place in their physician guides their decision.
AI systems are inherently solution-oriented, so they offer recommendations often on a shallow depth or surface level, which in several cases may be erroneous or misguided, as opposed to the real world, where recommendations are provided based on experience, diagnoses, characteristics and features from lab results and any possible respective underlying conditions
UX Lens — What’s different?
As a UX Practitioner, we make it a point to understand the pain points of our end users and often ask probing, exploratory questions to get to the root cause of the problem. AI, being result-driven, is known to overlook this step as it skips to the solutioning stage.

To ensure we have a holistic approach in AI development, Physicians and potential patients could be brought on board as human evaluators during SFT training, so a dialogue between them could serve as a grounding base for providing more context and scaffolding to AI during training. Additionally, we could also explore programming AI with a curiosity to explore causality, which could further help it in producing more efficient, focused and accurate predictions.
Harmonising AI with healthcare should not just be about accurate predictions; it should also integrate patient collaboration and exploration, which further pave the road towards trust and reliable guidance.
References:
1. DiSorbo, M. D., Ju, H., & Aral, S. (2025). Teaching AI to Handle Exceptions: Supervised Fine-Tuning with Human-Aligned Judgment. ArXiv. https://arxiv.org/abs/2503.02976
2. Amann, J., Blasimme, A., Vayena, E. et al. Explainability for artificial intelligence in healthcare: a multidisciplinary perspective. BMC Med Inform Decis Mak 20, 310 (2020). https://doi.org/10.1186/s12911-020-01332-6


