Coherence ≠ Compliance: The Capability Paradox, 10 Months Later
“We often assume that as LLMs get smarter, the validation burden decreases. The opposite is true, especially for R&D and CMC workflows.”
This is the opening of my piece “The Capability Paradox: Why Soaring LLM Benchmarks Demand Stricter Validation,” published in December of 2025, three weeks after filing the LLC for Britt Biocomputing.
Intelligence does not automatically indicate compliance. The methodology developed since rests on the thesis that the controls primarily lie in the architecture around the model, not only within it. And the controls include not only the system components, but the humans who interact with them.
Since December, I laid the foundations with my roadmap piece ”The Guidance Says What” in February and my umbrella governing framework, “House of AI Trust” in March. In both pieces, I laid out the roadmap for stress testing the architecture against increasingly complex probabilistic systems (agentic and multimodal AI, case studies), to demonstrate that it can bear weight. I also committed to publishing my methodological work:
The Ledger: What’s Shipped, What’s Shipping, and What’s Still Open
What’s Shipped
About half of the roadmap has shipped. The methodological work became named, public artifacts, licensed under CC BY 4.0: the Probabilistic Validation Lifecycle (published in October 2025, predating “The Capability Paradox”), VALID Trust, and the Probabilistic Failure Mode Taxonomy. VALID Trust has a public supplier worksheet, with a Zenodo DOI (10.5281/zenodo.23089779). Both VALID Trust and the Probabilistic Failure Mode Taxonomy have licensable methodologies that extend the public artifacts.
The examination of harnesses as a control substrate (”Beyond the Weights: Harness Monitoring as the Validation Leash”) was the seed for “The Last Deterministic Thing: The Unified Capture Layer.” Where the probabilistic system is the model x harness x human-in-the-loop within the defined context of use, the harness layer acts as a “Swiss cheese” of guardrails around the model. However, the guardrails, while a control on the model’s behavior and output, may be themselves probabilistic and/or dynamic. This raises the question: where can determinism live in a probabilistic system?
The Unified Capture Layer is one architectural answer to that question. It logs the event path, producing a deterministic trace. It can also distinguish a check that failed from a check that was triggered but never ran.
“The Inheritance Boundary” concludes with a worked example walkthrough of a deviation intake triage assistant up the floors of the House of AI Trust.
What’s Shipping
Two items are shipping this week, either in this piece or via a subscriber deliverable. This piece uses agentic structure to show why the unit of validation is the system. A full worked example of an agentic system using the House methodology will ship to subscribers of the AI Validation Boundary: Extended Edition on Friday (10/9/26). Signing up is free. All that’s needed is your company e-mail address.
What’s Still Open
Multimodal systems create additional validation and failure mode classification challenges that do not integrate cleanly within the existing frameworks. While both automated visual inspection and text LLMs used in GxP individually have emerging or established validation precedents, I am not aware of a precedent for the combination of both, fused at different points within the system pipeline.
Some multimodal failure modes exist in the interaction between the modalities, not only within each individual modality channel itself. The harness, the UCL, and the human reviewer all bear more weight in multimodal systems.
The research on this methodology is ongoing and will be submitted for peer review. Human performance qualification is an ongoing area of research and collaboration.
Beyond the Benchmarks
My December piece emphasized a need for more than generic benchmarks to justify an LLM’s use in a GxP setting. System-level validation, tailored to the context of use, is required. The December suggestion went entirely to the system, but the human who reviews the model’s output is another potential failure mode. Even if the system-level validation is rigorous, the human needs to be demonstrated to be qualified for the assigned review task. The burden does not land entirely on the validation of the computerized system.
Coherence ≠ Compliance
As model capabilities improve, outputs become more fluent and coherent. They don’t always become more accurate. This creates a challenge for humans reviewing outputs of these systems in a regulated setting.
Kim et al. (2024) found that when the model reports its uncertainty, humans are more likely to catch errors. However, if the output is presented as fluent, internally coherent text, humans are less likely to flag errors.
Humans are subject to automation bias, and LLM outputs create additional challenges with fluency, confidence, and internal coherence. The reviewer does not fail independently of the probabilistic system, and may, in fact, fail in step with it.
The Capability Baseline
On OpenAI’s FrontierScience benchmark, GPT-5.2 was the highest-performing model at the time of its release, and it scored 77% on the Olympiad set. The FrontierScience benchmark was developed to measure LLMs’ scientific reasoning capability; it was one of the first benchmarks to grade open-ended research reasoning rather than exam-style answers. The Olympiad set is a structured benchmark where the model must generate a constrained, short answer or expression.
The Research track, however, uses open-ended tasks written by scientists with PhDs, and graded against a 10-point rubric. These tasks require extended reasoning. GPT-5.2 only scored 25% on the Research set. On this benchmark, the tasks that are most difficult for a human reviewer to assess tend to be the types of tasks where the model’s performance is the worst (Wang et al., 2026).
A newer generation of benchmarks, such as Raycaster’s Biopharma Bench, more closely measure models’ capability at real-world workflows in various domains of expertise using artifacts from reconstructed work environments. However, there is still a need to test whether an agent is capable of the actual task in its designated context of use. This is the gap between a generic benchmark and defensible evidence that can be leveraged during qualification activities in GxP-regulated environments.
The Compounding Error
The December piece argued that if an agentic workflow delivers 79x efficiency, it is capable of increasing its error volume per unit time at a similar pace. Where that piece argued on the basis of speed, this piece argues on the basis of agentic structure. Each additional action in the agentic system’s path produces an additional potential point of error. The errors can propagate and compound.
If the reviewer only sees the agent’s final output, there is no reliable mechanism to detect errors that occurred earlier in the agentic path. A deterministic trace — which the Unified Capture Layer (UCL) provides — is what makes intermediate steps inspectable at all.
Some organizations default to an assigned SOP in the Learning Management System to deem human reviewers competent for the task. Organizations use this as evidence of sufficient human review of generative AI outputs. As stated in “The Inheritance Boundary” earlier this week, “the mechanism(s) chosen and the rigor of evidence required [to deem a human qualified for the delegated task] is tied to the context-of-use and risk assessment conducted in Layer 3.”
Reading an SOP does not qualify a human to review fluent, coherent, but incorrect outputs. Parasuraman & Manzey (2010) found that automation complacency occurs under conditions of multiple-task load, and occurs in both naive (beginner-level) and expert participants. Being an SME is also not an automatic qualification for reviewing generative outputs as a human in the loop.
Who Qualifies the Overseer?
I’ve argued what is insufficient for human performance qualification, but not what “good” actually looks like.
EU AI Act Art. 26(2) requires (for high-risk AI systems) and the GAMP AI Guide recommends overseer competence, but harmonization and standards are still developing.
This is not a GxP-specific challenge; other regulated life sciences are grappling with this same question. The FDA’s recent discussion paper Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback explores a competency-based approach to premarket evaluation of GenAI-enabled devices, somewhat analogous to how human clinicians are evaluated and credentialed through medical training. It then suggests sample-based clinician review as one postmarket control (alongside periodic benchmarking and degradation monitoring).
Automation bias has been found to extend to clinician review of AI output as well. Dratsch et al. (2023) found that radiologists’ performance at reading mammograms declined in cases where the AI suggested an incorrect BI-RADS category. The performance decline occurred across all expertise levels; however, very experienced readers rated significantly more mammograms correctly when the AI predictions were incorrect (compared with moderately experienced and inexperienced readers).
Across the regulated life sciences, regulatory thinking is converging on human review monitoring as an ongoing control mechanism. The ultimate challenge is the reference point itself: what qualifies the overseer?
Citations:
Dratsch, T., Chen, X., Rezazade Mehrizi, M., Kloeckner, R., Mähringer-Kunz, A., Püsken, M., Baeßler, B., Sauer, S., Maintz, D., & Pinto dos Santos, D. (2023). Automation bias in mammography: The impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology, 307(4), e222176. https://doi.org/10.1148/radiol.222176
GAMP AI Guide. International Society for Pharmaceutical Engineering. (2025). ISPE GAMP® guide: Artificial intelligence. ISPE.
Kim, S. S. Y., Liao, Q. V., Vorvoreanu, M., Ballard, S., & Vaughan, J. W. (2024). "I'm Not Sure, But…": Examining the impact of large language models' uncertainty expression on user reliance and trust. FAccT '24. arXiv:2405.00623.
Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). (2024). Official Journal of the European Union, L, 2024/1689. http://data.europa.eu/eli/reg/2024/1689/oj
U.S. Food and Drug Administration, Center for Devices and Radiological Health, Digital Health Center of Excellence. (2026, August 19). Considerations for the regulation of generative AI-enabled medical devices: Discussion paper and request for feedback (Docket No. FDA-2026-N-7874). https://www.fda.gov/medical-devices/digital-health-center-excellence/considerations-regulation-generative-ai-enabled-medical-devices-discussion-paper-and-request
Wang, M., Lin, R., Hu, K., Jiao, J., Chowdhury, N., Chang, E., & Patwardhan, T. (2026). FrontierScience: Evaluating AI's ability to perform expert-level scientific tasks. arXiv:2601.21165. https://doi.org/10.48550/arXiv.2601.21165