Building a Trustworthy AI research lab for healthcare on a budget

· Medium ·

4 min read Original article ↗

Gbenga Awodokun

This a first in a series of posts on building generative AI models for healthcare

I have been involved in building AI models or experimenting with them for the last 8 years but recently I started working on a generative AI model for healthcare use cases. These models have privacy (remember HIPAA) requirements and has a set of unique challenges. My background is in media and entertainment (YouTube, Roku), where you can use your access to APIs from frontier models like Gemini 3.8 or OpenAI GPT 5.6, try out a few things to establish some benchmarks then go to hugging face to find an open-source models e.g, Llama model 3.1. Using the benchmarks from your early iterations, you post-train (or fine-tune) the model with some of your data, host the models, then do offline evals for your benchmarks before experimenting with users. Everything good? feature is launched 🚀

Let say, we want to build a new video clip generation models to generate clips from existing videos content, we can use GPT-5.6 from openAI. Once, we test the idea with our data, we head to hugging face then download a Qwen3-VL model, and post-train for the specific use case, do offline evals, host and then run A/B tests on performance. In healthcare, not so easy!

AIML Models Flywheel: From ideas to production launch

Why building a model for healthcare can be difficult?

  1. Evals is specific and has a high-bar. For example, different clinicians implement various versions of the SOAP standard (subjective, objective, assessment, plan) notes and evals must meet the strictest requirements i.e., AI must know what is good and get it right everytime from start.
  2. Data is private. There is no IMDB (movies database) of healthcare. Whatever dataset you find will be heavily annoymised, redacted and privileged. If you find any, it is probably proprietary. No free lunch!
  3. Lastly, Offline evaluation requires expensive domain experts (clinicians) who are very busy and A/B testing with candidate models is difficult becuase you often lack sufficient samples for hypothesis testing
Model evals can be challenging in healthcare

Why evals is even harder?

More on evals, AI models in healthcare has subjected to the highest standards and required to be sovereign — no sending patients data to OpenAI GPT APIs. So AI applications are required to pass the bar for:

a. Safety: AI models and agents in healthcare must have consistent behavior and operate within guardrails. Guardrails are often established by clinicians and they have to be rigorously adhered to. Clinical errors such as an omission due to AI token budget limits, accessing unauthorized records by agents make the application usable in actual clinical practice at scale

b. Accuracy: Hallucinations can cost lives literally ☠️, for example — if AI models picked lisinopril instead of acetaminophen, we have a loss of trust scenario or even worse, we have a 911 ER visit. See ChatGPT health lawsuit

c. Completeness: Generative AI models often require some context for their output. It could be an instruction prompt, clinical data (especially in RAG applications) and so on. However, most models are limited by context windows (how much memory it has available). And context can be very long, for example, a doctor consultation visits could be up to 30 minutes or more. Any generative output must ensure the entire context is considered (except where intentional left out ) before generating an output

How we might build an AI models/agents for healthcare?

Here is the good news — there’s hope. In my subsequent posts, I walk through an example building AI models like an AI scribe (e.g. Newcare, Abridge, Heidi and so on). Let’s walkthrough how we can achieve this:

  1. Find clinicians and create a set of evals. Most clinicians have high frequency and recall of “what good looks like” in a clinical settings. As an example, coming up with a treatment plan for some clinical diagnosis. Walk through with them what “good” looks like. Often, they can offer access to redacted clinicial notes that can be used in creating good benchmarks. Speciality varies so you may need more than one.
  2. Begin with open-source models. Host open-weight models to begin evaluating the examples provided against the benchmarks using your model. Also, it is an opportunity for prompt optimization (e.g., APO). Once you are consistenly meeting and exceeding those benchmarks, it is time to start thinking about the next steps. If needed, you may also do some post-training (fine-tuning open weights) at this stage. Generally, avoid sending private health data to public models like OpenAI GPT
  3. Offline evaluations with gaurdrails. We want to ensure clinical benchmarks are met and exceeded by working with clinicians to evaluate the outputs with real-life traces. We evaluate for three benchmarks I mentioned earlier in this article.

So our flywheel look differently as below. In the next posts, we walk the code to accomplish all of these steps using some open-weight models.

Healthcare AI Model Flywheel