The Validation Boundary Inside the House

I first published my overarching framework for AI in regulated life sciences, the “House of AI Trust,” in March.

The five-layered framework stacks from the ground up:

  • Data governance and regulatory awareness (Layer 1)

  • AI governance (Layer 2)

  • Control layer/validation assurance (Layer 3)

  • Domain/process expertise (Layer 4)

  • Business ROI (Layer 5)



Within Layer 3 sit the Probabilistic Validation Lifecycle and the Probabilistic Failure Mode Taxonomy. The term probabilistic, in the context of this piece, as well as in my scope of practice, refers to the system’s (CoU x model x harness x HITL) capacity for output variation. This is distinguished from the term probabilistic as it refers more broadly to models that report a probability, but return a determined output.

The Probabilistic Shift

I first articulated the conceptual shift underlying my practice in October of 2025: “Validation for LLMS: An Interdisciplinary Perspective.” While my methodological thinking has evolved since then, the principles that grounded the launch of my practice have been consistent: large language model validation requires a risk-based, multi-disciplinary, and domain-dependent approach.

Large language models do not operate on the same assumptions that underpin the computerized system validation (CSV) discipline in GxP environments:



  • Test case with an expected result (same input = same output)

  • Traceability from requirement to identical test evidence

  • Deviation investigation and root cause

  • Requalification triggers (adaptive systems share this property)

  • Supplier-inherited evidence (adaptive systems share this property)



The last two are shared with adaptive systems and are substantively covered by the GAMP 5, 2nd Edition¹, Appendix D11 static/dynamic axis. Dynamic systems are explicitly carved out as requiring unique monitoring considerations. I argue that the fundamental reason the first three assumptions no longer hold lies in a probabilistic system’s inherent property: computing a distribution over an unbounded output space which cannot be enumerated in advance. In essence, there is no one correct output to test against for any given input.

That matters for GxP use cases. The model is only one source of the probabilism; the elements that live around the model, the harness and the human-in-the-loop (HITL) , are additional origination points for the same property. This is why my work centers around the methodological adaptations that bridge CSV practices with probabilistic systems. The principles hold (ICH Q9(R1), GAMP 5 2nd Ed.¹, GAMP AI Guide²) and the methods evolve to ensure the evidence continues to serve its purpose: proving a particular system is fit-for-intended use. GAMP 5¹ explicitly calls out "combining technical, procedural, and behavioral controls" when assessing the risk of a computerized system.

The Boundary Inside the House: AI Governance (Layer 2) and Domain Expertise (Layer 4)

This makes the case for revisiting the floors of the house in the context of the system boundary (CoU x model x harness x HITL). In Layer 1, data governance exists as a distinct discipline of its own, and regulatory awareness has been interwoven across several pieces of my recent work.

Although I’ve introduced Layer 2 (AI Governance), I have yet to operationalize it in my public-facing work. Inventory, governance policy, and accountability live here. Inventory is the first component of Layer 2, and it allows an organization to visualize the AI systems that are already in use. This matters: not every system carries the same level of risk, and GAMP guidance is clear that the bulk of validation efforts should be focused on systems with direct impact on product quality and patient safety.

The question then arises: what operationalizes that determination? The same model can be used in combination with different contexts-of-use, different harnesses, and different humans in the loop. Take System A, System B, and System C, pictured below, each utilizing the same base model:

In order to initiate the validation assurance lifecycle, a clear visualization of the system in question is required, to understand the level of risk carried by each component within the validated boundary. This exercise directly feeds the validation assurance lifecycle activities in Layer 3 covered by the Probabilistic Validation Lifecycle, a probabilistic-specific lifecycle mapping to GAMP 5¹/GAMP AI² Guide principles.

Layer 4 is dependent on domain and process-specific expertise. Explainability and communication (originally cross-cutting threads) live here. Only an internal resource with the appropriate context and domain expertise can evaluate whether a particular confidence interval or model output is appropriate for the defined CoU.

The domain/process expert's role and the interface's obligation (Communication) meet at exactly one point: the moment a probabilistic output has to become a human decision. Only the domain expert can make the contextual determinations that acceptance criteria rest on: at what threshold is it necessary for an output to route to a human for review? What does a pharmacovigilance adverse-events triage model’s determination that it is 85% certain of its decision mean for its CoU?

These are domain questions, not statistical ones, and they require domain-specific justifications. The failure mode error taxonomy's Class 3 (Contextual Misapplication) and Class 6 (Population Bias) are the two classes only a domain expert can catch; no guardrail, no ontology check, and no confidence score detects "this answer is correct for the wrong population."

This is all inherited upwards in Layer 5; achieving business ROI through the framework depends upon framing AI not as a model problem, but as a system one.

The Cross-Cutting Layers

Two threads cut vertically across all five layers of the house:

The boundary remains relevant here as well; with vendor systems, parts of the boundary sit outside the sponsor’s walls, and surfacing which elements of the boundary sit where and who is accountable for each is the foundation for qualification and security considerations. This will be explored in greater depth in a piece on the inheritance boundary.

In Summary

The probabilistic system definition is relevant outside of layer 3 (the validation lifecycle itself). If a probabilistic system is CoU x model x harness x HITL, the inventory unit is the system, the risk determination is per system, and the acceptance criteria are domain judgments rather than statistical ones. Ingrid Witherell made a sharp point to this effect in the comments of my recent piece “The Last Deterministic Thing”: "Where the human sits in the [workflow] sequence is a domain decision.”




¹ ISPE. GAMP® 5: A Risk-Based Approach to Compliant GxP Computerized Systems, 2nd ed. International Society for Pharmaceutical Engineering, 2022.

² ISPE. ISPE GAMP® Guide: Artificial Intelligence. International Society for Pharmaceutical Engineering, 2025.

Next
Next

The Version Bump Nobody Change-Controlled