The Inheritance Boundary Inside the House
Inheritance, Ownership, and the EMA Annex 22 Workshop Report
Two weeks ago, I published “The Validation Boundary Inside the House,” which covered the expanded validation boundary (CoU x model x harness x HITL) in the context of the probabilistic system, and revisited Layer 2 (AI governance) and Layer 4 (Domain and Process Context) in greater depth.
The overarching thesis of the piece rests on the concept that if the probabilistic system includes the model, harness, and human-in-the-loop within its CoU, every layer of the house must evaluate the entire system, not just the model. This includes regulatory guidance interpretation, data governance, AI governance, validation assurance evidence, domain/process expertise, and business ROI.
On October 2nd, 2026, the EMA released its report from the June 30-July 1 workshop on Annex 22 (EMA/156789/2026), with final guidance expected to land sometime in Q4. The original draft Annex 22 was structured by both technology class and system behavior: generative AI, large language models (LLM), models with a probabilistic output, dynamic and adaptive models are deemed out of scope and “should not be used” in critical applications. Note that draft Annex 22’s only defines a deterministic model (one with the same output for the same input), implying that a probabilistic model is one where the same input may produce varying outputs, while this piece defines probabilism outright as a system-level property (given that even probabilistic models wrapped in a deterministic layer may retain probabilistic properties at the system level).
The industry consensus captured in the 2 October EMA/156789/2026 report suggests a shift toward a risk-proportionate approach; the conclusion states that “subject to appropriate controls and a robust risk-based framework, there may be viable pathways for the use of adaptive, probabilistic and generative AI technologies in GMP-regulated applications while continuing to protect product quality and patient safety.”
Technology vs. System Classification
Given that technology continues to evolve, any technology classification-based guidance may become obsolete, meaning that enduring guidance likely has to remain at the system level. This necessitates a structured framework for classifying “AI” systems by property, as this enables the identification of appropriate failure mode grouping, governance and control strategies.
Generative AI, large language models, and agentic AI describe classes of current technology, rather than types of system behavior. Dynamic, adaptive, and probabilistic labels describe properties a system has, and are enduring classifications. Creating taxonomies and frameworks around system behavior, rather than current technology, ensures their durability over time.
Guidance exists for validation and oversight strategies by static vs. dynamic behavior of computerized systems, including AI. However, probabilistic systems, defined as systems that may produce varying outputs given the same input, have a distinct behavioral profile that currently lacks a structured approach. In the 2 October 2026 report (EMA/156789/2026), regulators explicitly requested “a quality risk management approach to the entire lifecycle of such AI system models that may be used in GMP high risk areas where there may be a low detectability of deficiencies but a high impact on the patient and taking into account ICH Q9R1 principles.”
The House of AI Trust was built to create this structure. The five-layered framework stacks from the ground up:
Data governance and regulatory awareness (Layer 1)
AI governance (Layer 2)
Control layer (Layer 3), including the Probabilistic Validation Lifecycle and Probabilistic Failure Mode Taxonomy
Domain and process expertise (Layer 4)
Business ROI (Layer 5)
The Inheritance Boundary
Where “The Validation Boundary Inside the House” examined the house from a horizontal perspective, this piece examines the house from a vertical one. The question this piece answers is: as we walk up the layers of the house, what evidence can be inherited, and what cannot be delegated?
The House is structured such that the evidence that may be inherited narrows, while what is handed up accumulates, as you walk up the house from the ground floor (data governance and regulatory awareness) to the peak (business ROI).
Inherited: Repurposed from an internal or external source
Owned: Must be constructed for its purpose
Handed up: Handed from one layer to the one above
Layer 1: Data Governance and Regulatory Awareness
Data governance is a mature discipline, and the vast majority of regulated life science companies have some form of implementation. While stored data can be inherited from external sources, the interpretation of data (context) the system retrieves at runtime should not be delegated. Data should be interpreted based on the sponsor’s policies and systems in a structured manner (e.g., via an ontology or knowledge graph).
Regulatory awareness can be inherited, as regulatory guidance is external by definition. This may include approved enforceable regulation, guidance, and industry best practice documents. The interpretation and application of guidance into policy occurs at the governance level (Layer 2).
Layer 2: AI Governance
Layer 2 is where the elements of Layer 1 are handed up to translate into policy and oversight. AI governance includes:
Inventory and Classification
Governance policy
Accountability
Explainability
Inventory is the prerequisite for policy, accountability, and explainability: you can’t govern what you can’t see. Ideally, the organization should maintain a living representation of all AI systems and their contexts-of-use (not just the models, but the harness layer - the components around the model - RAG systems, confidence scoring, ontology, deterministic guardrails, and where the human lives in the loop). Behavior is governed by the system, not the model, and visibility should extend in parallel. Computerized systems with AI components should be explicitly classified by system-level properties (i.e. adaptive, dynamic/static, probabilistic) in order to determine which elements of the validation strategy require a tailored approach.
Governance policy may be inherited from the existing quality system (with AI-specific extensions as required); AI systems are computerized systems, and regulated pharma already has established policies around governing those. The organizational chart can be utilized to determine which roles and responsibilities lie with which teams.
Accountability cannot be inherited. Each AI system currently in use (or set to be put into production) must have a named owner who is responsible for monitoring the system in production and ensuring it remains within its operational boundaries.
Explainability means that the system behavior and outputs should be interpretable to the people who act on them, monitor them, and defend them to an auditor. This cannot be inherited; it must be architected into the system intentionally.
The inventory should ultimately include all AI-enabled computerized systems, each with its own system-level property classification, governance policy, ownership (accountability), and explainability requirements. The system inventory profile then directly informs the validation assurance evidence and controls required in Layer 3.
Layer 3: Control Layer (Validation Assurance Evidence)
Industry consensus captured in the 2 October 2026 report suggests that “AI requires that established controls and newer AI-specific controls operate together as an integrated framework.” However, the challenge raised during Q&A was how a manufacturer could demonstrate to inspectors that the combined layered controls provide reliable output for a critical decision.
The Probabilistic Validation Lifecycle and Probabilistic Failure Mode Taxonomy were developed to adapt GAMP 5 2nd Edition and GAMP AI Guidance lifecycle computerized system validation to probabilistic systems. Validating systems that may contain variability is not new to pharma; as Brian Drapeau notes in the August 2026 PharmTech article “Europe Tried to Ban Generative AI from Critical GMP. The Ban May Not Survive. It Does Not Matter,” pharma already validates analytical methods, biological assays, and manufacturing processes.
The distinction is that these systems have an established referent, and their behavior is therefore more predictable.
The 2 October 2026 report Q&A session for Topic 1: Regulatory Pathways for Adaptive/probabilistic AI GMP focused on whether the framework addresses features of generative AI and LLMs, including hallucination, fabrication, sycophancy, or epistemic overconfidence. The response was that these behaviors should be identified and evaluated during development and validation, and that if the manufacturer cannot demonstrate fitness for intended use in view of those failure modes, the use cannot be justified.
However, you cannot control, or CAPA, what you cannot classify. The Probabilistic Failure Mode Taxonomy is intended to be utilized in a FMEA-style risk assessment approach, during stage 2 (risk assessment) of the PVL, as well as during CAPA and root cause analysis. It classifies probabilistic system errors by error type and origination point, as successfully applying controls requires identifying where the failure mode originates.
Each system handed up from the inventory profile in Layer 2 requires assurance evidence. The type and depth of evidence is determined by the context-of-use, risk assessment, and system properties.
Some evidence may be inherited, including supplier test evidence and model benchmarks and/or model cards, and included in the validation evidence package. GAMP 5 2nd Ed. and GAMP AI Guide encourage leveraging supplier evidence where appropriate. However, this evidence does not supplant demonstrated fitness for the sponsor’s specified CoU across the full system boundary.
Given the boundary between supplier and sponsor, and the need for continuous monitoring of systems with certain properties (particularly probabilistic and/or adaptive), a Unified Capture Layer may provide ongoing visibility into whether the various components of the system (including the model, the harness layers, and the human-in-the-loop) remain within the established validated boundary.
Layer 4: Domain and Process Context
Domain and process experts play a key role in the AI implementation and oversight process. When inputs may produce varying outputs (probabilistic system), domain expertise input assists in determining appropriate acceptance criteria, and in detecting domain-specific failure modes such as population bias or a model trained on one site’s data. This documentation is a critical part of the validation package, and cannot be delegated. The justification is the referent, specific to each context-of-use for a probabilistic system.
The explainability requirements established in Layer 2 are operationalized here. It must be demonstrated via evidence that all parties responsible for validating, evaluating, and monitoring the AI system understand the system’s capacities and limitations to the degree the role requires. This cannot be inherited from the supplier or model creator.
This is also where it must be demonstrated, with evidence, that the human(s)-in-the-loop are qualified for the delegated task. Several mechanisms may be employed to accomplish this goal, including seeded challenges, review task time, and percentage of outputs accepted vs. challenged. The mechanism(s) chosen and the rigor of evidence required is tied to the context-of-use and risk assessment conducted in Layer 3.
Layer 5: Business ROI
When Layers 1 through 4 are successfully implemented with a robust inheritance boundary, business ROI becomes defensible. Business ROI is the evidence that the non-delegable judgment held across all four vertical layers. It must be earned through evidence, not asserted.
Deloitte’s 2026 Life Sciences Outlook found that only 22% of executives surveyed have scaled AI, and even fewer (9%) report significant returns. An EY survey of 975 C-suite executives found that firms with real-time AI monitoring and oversight committees are more likely to report improvements in revenue growth, employee satisfaction and cost savings.
Many of the commonly cited causes of poor ROI are not model problems, but governance and data problems. Governance is what turns claimed value into provable value.
The Cross-Cutting Threads
Two threads cut vertically across all five layers of the house:
Security (”Your Frozen Architecture May Have a Backdoor,” “Safety Paradox”): Cybersecurity breaches reached a record global high in 2026 (IBM 2026) and regulated life sciences companies own sensitive data, including patient information and intellectual property, whose compromise may produce negative business consequences including manufacturing delays, reputational backlash, and lawsuits. Verizon’s 2026 Data Breach Investigations Report identified that vulnerability exploitation accounted for 31% of breaches (November 2024 through October 2025) up from 20% a year prior.
Supplier Qualification (VALID Trust, Harness Evidence Worksheet for AI Supplier Assessment): When the model and parts of the harness sit on the vendor’s side, the sponsor can inherit the supplier’s evidence, but not is accountability. Third-party involvement in breaches rose 60% year over year in the 2026 Verizon DBIR, and standard certifications such as SOC 2 attest to a vendor’s security controls, not to whether its AI components stay within the sponsor’s validated boundary.
Conclusion: Walking One System Up the House
As established in “The Validation Boundary Inside the House,” the probabilistic system includes the CoU × model × harness × HITL. For a deviation triage assistant the system includes the following components:
CoU: drafts deviation investigation summaries and proposes root-cause categories from batch records and event data. It does not decide disposition or impact.
Model: a vendor-hosted foundation LLM. The supplier controls versioning.
Harness: retrieval over SOPs and historical deviations, prompt templates, required citations to source records, and deterministic checks (e.g., mandatory fields, an approved root-cause list).
HITL: a QA investigator reviews, edits, and approves every draft.
Layer 1: Data governance and regulatory awareness
Inherited: regulation and guidance (Part 11, Annex 11, draft Annex 22, GAMP AI Guide, GAMP 5 2nd Edition) and the vendor’s statements about training data.
Owned: what goes into the retrieval corpus and how it’s interpreted (superseded SOP versions, site-specific terminology, and historical deviations that were misclassified at the time).
Handed up: lineage for the retrieval corpus
Layer 2: AI governance
Inherited: existing CSV and QMS policy; roles taken from the organizational chart.
Owned:
an inventory entry that captures the full system, not just “uses [vendor] LLM”;
a named system owner;
the system-level property classifications: probabilistic, since the same deviation record can yield different drafts, and dynamic, since the historical-deviation corpus grows over time.
the explainability requirements, e.g., every claim in a draft must trace to a source record.
Handed up: the inventory profile.
Layer 3: Control layer
Inherited: the vendor’s model card, test evidence, and certifications.
Owned:
Risk assessment using the Probabilistic Failure Mode Taxonomy, by origin point. Examples:
a fabricated citation (model);
retrieval of a superseded SOP (retrieval);
anchoring to the most frequent historical root cause (model and corpus);
a reviewer approving the draft without changes (human)
Probabilistic Validation Lifecycle execution with CoU-specific acceptance criteria.
Characterizing output variation by running the same inputs repeatedly;
A trigger for vendor model updates;
An envelope of anticipated changes (e.g., “pre-agreed change types that don’t trigger revalidation)
Monitoring through the UCL: model versions, retrieval logs, and reviewer actions.
Handed up: the validated boundary and the known failure modes.
Layer 4: Domain and Process Context
Inherited: essentially nothing. Vendor claims about domain capability don't transfer.
Owned:
SMEs define what a defensible root-cause rationale looks like;
domain-specific failure modes, e.g., downgrading an aseptic event;
investigators trained on the system's known limitations;
HPQ evidence: seeded-error challenges, review time, and accepted-versus-challenged rates.
Handed up: evidence that the human is an effective control, not a nominal one.
Layer 5: Business ROI
Inherited: nothing. Vendor time-saving claims and benchmarks are assertions.
Owned:
investigation cycle time measured net of review and rework;
quality indicators such as reopened deviations, recurrence, and CAPA effectiveness;
proof the savings didn't come from rubber-stamping, by comparing the acceptance rate with the seeded-error catch rate.
Cross-cutting threads
Supplier qualification (VALID Trust / Harness Evidence Worksheet): which boundary components sit at the vendor (the model, possibly hosted retrieval), and what the contract requires for notifying model changes.
Security: prompt injection through free-text deviation descriptions, and where batch-record data goes.
This walkthrough is a response to the workshop Q&A call for demonstrating that layered controls produce reliable output for the context-of-use. The same structure scales to critical uses with more rigor at Layers 3 and 4. What’s inherited narrows, and each layer hands up evidence the next one depends on. The structure is meant to provide a crosswalk from regulation, to governance, to fit-for-purpose evidence, to ultimately, return on the business investment. An AI investment must be justified, not assumed.
Citations:
Depa, J. (2025, October 8). How can responsible AI bridge the gap between investment and impact? EY. https://www.ey.com/en_gl/insights/ai/how-can-responsible-ai-bridge-the-gap-between-investment-and-impact
Drapeau, B. (2026, August 13). Europe tried to ban generative AI from critical GMP. The ban may not survive. It does not matter. Pharmaceutical Technology. https://www.pharmtech.com/view/europe-tried-to-ban-generative-ai-from-critical-gmp-the-ban-may-not-survive-it-does-not-matter-
European Commission. (2025, July 7). Annex 22: Artificial intelligence [Consultation draft]. EudraLex Volume 4: Good manufacturing practice guidelines. https://health.ec.europa.eu/document/download/5f38a92d-bb8e-4264-8898-ea076e926db6_en
European Medicines Agency. (2026). Good manufacturing practice (GMP): Multistakeholder workshop on expert contributions to artificial intelligence guidance development (Annex 22) [Workshop report] (EMA/156789/2026). https://www.ema.europa.eu/en/documents/report/report-multistakeholder-workshop-expert-contributions-artificial-intelligence-guidance-development-annex-22_en.pdf
Gartner. (2024, July 29). Gartner predicts 30% of generative AI projects will be abandoned after proof of concept by end of 2025 [Press release]. https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025
IBM. (2026). Cost of a data breach report 2026. https://www.ibm.com/reports/data-breach
ISPE. GAMP® 5: A Risk-Based Approach to Compliant GxP Computerized Systems, 2nd ed. International Society for Pharmaceutical Engineering, 2022.
ISPE. ISPE GAMP® Guide: Artificial Intelligence. International Society for Pharmaceutical Engineering, 2025.
Lyons, P., Konersmann, T., Jacobson, S., Kleyn, N., Rekhraj, K., & Gosalia, D. (2025, December 9). 2026 life sciences outlook. Deloitte Insights. https://www.deloitte.com/global/en/insights/industry/health-care/life-sciences-and-health-care-industry-outlooks/2026-life-sciences-executive-outlook.html
Verizon Business. (2026). 2026 data breach investigations report. https://www.verizon.com/business/resources/T161/reports/2026-dbir-data-breach-investigations-report.pdf