Services · AI Governance for Regulated Science

Make LLMs and agentic AI inspectable, defensible, and audit-ready.

Every engagement follows the same lifecycle. The depth scales with your risk — and with how close the system sits to patient safety, product quality, and data integrity.

Context of Use Risk Tiering Evaluation Design Acceptance Criteria Oversight Model Monitoring Change Control

Find out where you stand

Before a regulator or an auditor does. Both of these are free.

Free download

The House of AI Trust™

32-page guide · no email required
  • Five-layer governance architecture, layer by layer
  • 25-point readiness checklist with a five-tier maturity model
  • Seven-step Probabilistic Validation Lifecycle with named deliverables
  • Twelve inspector questions and the evidence that answers each
  • Two worked examples, end to end
Free

20-Minute Fit Check

No cost · no pitch

I'll walk through your current setup, tell you where you sit on the validation lifecycle, and recommend the smallest next step that creates the most clarity. If that step isn't me, I'll say so.

Where you stand against what's landing

The first paid step, and the one most teams should start with.

Entry point

The AI Exposure Screen

$4,500 fixed fee · up to 3 systems · report in 5 business days

Where your deployed AI stands against the rules about to land.

Who it's for

Teams that have already deployed or piloted something. If you built an agentic workflow or an LLM-assisted process in the last eighteen months and have started to wonder, quietly, whether it would hold up under inspection — this is that answer, in writing.

Scope

Up to three AI, LLM, or agentic systems currently in production, in pilot, or committed on the roadmap. Additional systems $900 each, to a maximum of five.

Three is deliberate. A screen that gives each system twenty minutes is a survey, not evidence.

How it runs

  • Async intake — system inventory form plus a short document request. Roughly 45 minutes of your time, most of it forwarding artifacts you already have: architecture notes, prompt or configuration records, whatever validation documentation exists.
  • One 90-minute working session, with IT or data science and a Quality counterpart both in the room. Because the intake carries the discovery, this session is confirmation and challenge — not you explaining your stack from scratch.
  • Exposure Report delivered within 5 business days.
  • 30-minute readout call.
Free re-score against the final EU GMP Annex 22 text once it publishes, for 12 months from delivery. When the final language lands, you get updated scoring without a new engagement.
Fixed fee, invoiced by card or ACH · Two-page services agreement · No MSA and no PO required

What you're scored against

  • EU GMP Annex 22 (draft). The static-model restriction, and the exclusion of dynamic models, generative AI, and LLMs from critical applications — the clauses least likely to move in the final text. Critically, this includes the criticality determination itself: whether your classification of a system as non-critical is documented, justified, and defensible, which is where most generative-AI use in a GMP environment actually lives.
  • EU GMP Annex 11 (draft revision). Validation, audit trails, supplier oversight, identity and access management, and cybersecurity. This is the framework that reaches AI-enabled systems outside GMP manufacturing — the MLR support tool, the PV triage assistant, the submission QC step — where Annex 22 may not apply at all.
  • 21 CFR Part 11. Where electronic records and signatures are in scope for the system under review.
  • The FDA seven-step credibility framework. Is there a defined context of use, a model risk assessment across model influence and decision consequence, and a credibility assessment plan.
  • The ten FDA–EMA Guiding Principles, published 14 January 2026.
  • HITL / HOTL adequacy for each context of use.
  • The Two-Dimensional Error Taxonomy — our proprietary two-axis failure-mode lens, applied consistently across systems. The full rubric ships as an appendix so the scoring can be checked, argued with, and reused internally.

What you get

  • A scored exposure table across your systems
  • The two or three findings that would be hardest to defend under the direction the drafts have set
  • A recommended remediation sequence, ordered by exposure and effort
  • The scoring rubric as an appendix — what makes this evidence rather than opinion
  • A one-page executive summary written to be forwarded on its own

Three to four pages plus the summary. Short enough to read, structured enough to send to your VP of Quality without a cover note explaining it.

A note on timing. Annex 22 and the revised Annex 11 are not yet in force, and there is no enforceable US regulation specific to AI. The Screen assesses your position against the direction regulators have made explicit in draft and in joint principles — while there is still a grace period to act in, rather than after the final text lands.

What comes next

The Screen ends with a sequence, not a plan. Neither of these is a condition of the Screen — the report stands on its own and is yours to execute however you like.

Portfolio

Prioritized Use-Case Map

from $9,500 · up to 10 use cases · 2 weeks

Where the Screen goes deep on three systems, the Map goes wide across the portfolio. Most teams don't have a validation problem yet — they have a prioritization problem.

  • Full inventory of your AI, LLM, and agentic workflows — live, piloted, and planned
  • Each use case risk-tiered by model influence × decision consequence
  • A written, prioritized map: what to validate first, why, and what the first artifact must cover
  • An on-ramp to a Build engagement — or a defensible plan you can run yourself
Organization

AI Governance Framework

Scoped per engagement

Layer 2 of the House, built into your QMS. For teams that need the governance layer standing before individual systems can be validated against anything.

  • Model inventory and shadow-AI discovery process
  • Acceptable use policy — permitted, restricted, prohibited
  • Accountability structure and escalation pathways
  • Risk-tiering SOP and explainability requirements scaled to tier

Built against the House of AI Trust™ and adapted to how your organization actually approves things.

Turn "we use it" into "we can defend it"

Scoped engagements that produce the evidence. Three tiers, matched to system complexity and regulatory exposure.

Tier 1

R&D Fit-for-Purpose Sprint

from $18,000 · 2–4 weeks

Choose this if you have one AI tool or LLM deployment in R&D and would rather start with one done right than boil the ocean.

  • Context of Use and intended-use definition
  • Risk-tiered testing protocol and acceptance criteria
  • Biomedical error taxonomy for the deployment
  • QA-ready evidence package
Tier 2

Agentic Evidence Package

from $28,000 · 4–8 weeks

Choose this if your system chains multiple models, tools, or agents — where one step's output becomes the next step's input, and deterministic assumptions break.

  • Frozen architecture design — locked models, prompts, tool versions
  • Multi-agent error-propagation analysis
  • Human-on-the-loop oversight framework
  • Transparency and traceability architecture
Tier 3

GxP Validate-Launch

from $55,000 · full lifecycle

Choose this if an AI system is entering a GxP environment and needs to survive inspection — not just internal QA review.

  • Full-lifecycle validation across all seven stages
  • Complete audit-ready evidence package
  • Oversight, monitoring, and drift-detection design
  • Aligned to your QMS and change control

The Head of AI Quality role, before you hire one

Retainer

Fractional AI Quality Lead

Engaged monthly · scoped per organization

An embedded, ongoing partnership across your AI portfolio — governance, oversight, and inspection readiness held by someone who has done the validation work and the model work. Scope depends on portfolio size and how much translation sits between your data science, quality, and regulatory functions, so it's priced per engagement rather than off a rate card.

  • Ongoing validation oversight across your AI portfolio
  • AI governance framework development and maintenance
  • Cross-functional translation (data science ↔ QA ↔ regulatory)
  • Workforce training on AI interaction, skill narration, and oversight models
  • Inspection readiness and regulatory horizon scanning
Built for

Who this work is for

QA / Validation Digital / IT R&D Informatics Data Science Regulatory Affairs CMC / Manufacturing Pharmacovigilance C-Suite / Board

Pharma and biotech primarily, plus CROs, CDMOs, software vendors, and MedTech teams where regulated thinking adds value.

Not sure where to start?

If you've already deployed something and want to know what an inspector would find, start with the Exposure Screen. If you're not sure yet, read the guide and score yourself against the 25-point checklist — then book the free fit check and I'll tell you which layer to fix first.