The Burden of Proof Moved
Nobody is asking whether AI works in pharma anymore — they're asking what would count as showing it. EMA's draft Annex 22 bars generative AI from critical GMP work, and four of its own controls are unenforceable against a foundation model. Plus: Schrödinger trims headcount in the quiet repricing of platforms that actually have revenue, Nature Medicine publishes the digital-organism blueprint with every author holding equity, Insilico's CEO argues deal value is the wrong scoreboard, and the Onion Desk.
This is the full August 18 briefing, reproduced here in its entirety. This week's theme: the burden of proof moved. Four stories, one shift. European regulators are asking AI vendors for evidence about foundation models that those vendors structurally cannot produce. The public markets are trimming headcount at a physics-and-ML platform that has been selling software into pharma longer than almost anyone. A Nature Medicine Perspective lays out a decade-scale vision for simulating biology, authored entirely by people with equity in the company that would build it. And the CEO of the one company with a peer-reviewed AI-discovered clinical readout spent the week arguing that deal value is the wrong scoreboard. Nobody is asking “does AI work in pharma” anymore. They are asking a much harder question: what would count as showing it?
Europe’s Draft GMP Rules Bar Generative AI From Critical Manufacturing — and the Four Controls That Make It Unenforceable
RegulatoryWhat’s new — EMA’s draft GMP Annex 22 states plainly that it does not apply to generative AI or large language models, and that such models should not be used in critical GMP applications — but sustained industry pushback has EMA now gathering evidence on how to let them in anyway.
What happened — The six-page draft annex to EudraLex Volume 4 was written by EMA’s GMP/GDP Inspectors Working Group with PIC/S, with FDA and the UK’s MHRA sitting in as observers. It supplements Annex 11 rather than replacing it. After the consultation closed in October 2025, EMA reported stakeholder support for enabling generative AI and convened a two-day multistakeholder workshop on June 30–July 1, 2026 specifically to gather expert input on guardrails and mitigation measures. As of August 11, no final text exists in EudraLex, PIC/S still lists Annex 22 under draft guidelines, and EMA’s working plan targets Q4 2026 for delivering final text to the European Commission. The July 2025 exclusion therefore remains the published position — this is analysis reported by Pharmaceutical Technology on August 13.
How it works — Annex 22’s scope covers machine learning models that are static (parameters frozen, no adaptation during use) and deterministic (identical inputs produce identical outputs). Everything downstream of that scope definition is conventional validation machinery: characterize intended use including rare inputs and bias risk; set test metrics and acceptance criteria before testing begins; wall test data off rigorously from training data; provide explainability via feature attribution such as SHAP or LIME; log confidence scores against thresholds that force an “undecided” flag rather than an unreliable answer; apply change control and drift monitoring.
That machinery is sound. It is also built for a category of model that excludes what most people now mean by AI.
Key insight — Per the Pharmaceutical Technology analysis, four of Annex 22’s own controls become difficult or impossible to apply to a foundation model — not because the annex is badly drafted, but because it assumes properties these models don’t have.
Test-set independence is unprovable. Section 6 requires demonstrating that test data never touched training. Against a web-scale corpus you didn’t assemble and can’t inspect, you cannot produce that evidence. This is the difference between measuring generalization and measuring recall.
Explainability becomes narration. A model asked to explain its reasoning produces more model output. Whether that account faithfully describes the computation is a separate question — and a plausible narrative is worse than no explanation if it manufactures confidence the evidence doesn’t support.
Confidence scores aren’t confidence. Token probability is not an estimate of factual correctness. The “undecided” flag depends on the system reliably knowing when it doesn’t know, which is precisely the capability in question.
The frozen state may not be yours to freeze. With a vendor-hosted model behind an API, weights, system prompts, and safety layers can change on the provider’s schedule. A change-control trigger you cannot observe is not a control.
There is a fifth trap that applies to every model: Section 4.3 requires acceptance criteria at least as high as the process being replaced. Almost no organization has ever measured the error rate of the human process it is now automating.
Behind the news — FDA went a different direction. Its January 2025 draft guidance proposes a seven-step, risk-based credibility assessment for a defined context of use — credibility for a context, not validation of a system. But FDA has already moved from guidance to enforcement: in Warning Letter 320-26-58, issued April 2, 2026, the agency gave inappropriate use of AI in pharmaceutical manufacturing its own subsection, citing 21 CFR 211.22(c) after AI-generated specifications and master production records entered use without adequate quality-unit review. That is the first named AI enforcement action of its kind in this space, and it did not wait for anyone’s final guidance.
Why it matters — If you export into the EU, you meet EU GMP wherever you operate — this is not a European problem to monitor from a distance. More immediately: the draft never uses the word “agent” across its six pages. Agentic systems get sorted by the properties the draft does define — generative, dynamic, probabilistic — meaning the category the industry is deploying fastest is governed entirely by implication. If you have an agent anywhere near a batch record, you inherit an assurance problem nobody has named yet.
We’re thinking — Everyone is watching whether the exclusion survives. That’s the wrong thing to watch, because both outcomes land in the same place. If the exclusion holds, generative AI stays in non-critical work where the draft already assigns responsibility to qualified personnel. If it softens, any pathway into critical use will still require establishing who is qualified, for which decisions, under what conditions — because no technical control absorbs GMP accountability. Work distributes between a person and a model. A signature does not.
The tradeoff nobody is pricing: the less assurance you can establish around the model, the more assurance you need around the human review layer — and reviewing model output is a harder qualification problem than the interventions operators are currently qualified for. An operator performing an aseptic intervention can observe the process. A reviewer checking model output is assessing a finished artifact, with no faithful trace of how it was produced, from a system that is confidently wrong in the same register it is confidently right. Automation bias makes this worse, not better.
Our prediction on second-order effects: the winners here won’t be the organizations with the best models. They’ll be the ones who figured out early that the unit requiring qualification was never the model — it was person × system × version × use case, and that a “deviation management” use case is not a use case at all. “Retrieve and summarize related closed deviations for human investigation” and “determine probable root cause from historical deviation records” are the same model, the same records, and completely different evidentiary burdens. Vendors will keep selling the second while their documentation only supports the first. Expect that gap to show up in warning letters before it shows up in guidance.
Sources: EMA · PIC/S · Pharmaceutical Technology · FDA
Schrödinger Trims Headcount — the Quiet Repricing of AI Platforms That Actually Have Revenue
Industry dataWhat’s new — BioSpace’s layoff tracker logged workforce reductions at Schrödinger and Aura Biosciences on August 11, putting one of the field’s longest-running physics-plus-ML drug discovery platforms on a list it has largely avoided.
Context — Aura’s cuts were disclosed in detail: 20% of staff and a C-suite overhaul as the Boston biotech narrows to ocular oncology under a CEO installed in April, per Fierce Biotech. Schrödinger’s numbers were not broken out in the trackers as of publication; we’re flagging the entry rather than the magnitude. Broader context from BioSpace: through August 11, eight biopharmas were making or projecting reductions affecting 161 employees, against 11 companies and 562 employees over the same stretch in August 2025.
Why it matters — Schrödinger is the closest thing this sector has to a control group. It sells software with recurring revenue, publishes its methods, and has been doing computational chemistry for pharma since long before “TechBio” was a word. When the tools-as-software model trims while pre-clinical platforms with no disclosed candidates raise billions, the market is not making a judgment about whether computation works in drug discovery. It is making a judgment about which business model captures the value.
We’re thinking — There are roughly three viable models in this field: tools-as-software, pipeline-with-platform, and partnership-only. Capital is currently flowing hardest toward the model with the longest feedback loop and the least near-term revenue accountability. That is not irrational — a decade-long development cycle rewards optionality — but it does mean the segment with the most legible economics is the one under pressure. Watch whether software-first platforms respond by adding pipeline exposure. That would be the tell that pure tooling can’t sustain a public-market multiple in this cycle.
Sources: BioSpace · Fierce Biotech
Nature Medicine Publishes the Blueprint for an “AI-Driven Digital Organism”
PapersWhat’s new — A Perspective published August 13 in Nature Medicine proposes constructing an AI-driven digital organism (AIDO) — a modular, connectable system of integrated multiscale foundation models spanning molecules to cells to whole individuals — as a safe, affordable, high-throughput alternative to manipulating biology in the physical world.
Disclosure and source quality — This is a Perspective, not a results paper: a vision statement with a three-stage roadmap, not benchmarks. All three authors — Le Song, Eran Segal, and Eric Xing — declare a financial interest in GenBio AI, the company positioned to build the thing being described. Affiliations span GenBio AI, MBZUAI, the Weizmann Institute, and Carnegie Mellon. Peer reviewers are named (Weidi Xie, James Heath), which is more transparency than most such pieces carry. Read it as a well-credentialed prospectus rather than as evidence.
How it works — The architecture argues that biology is inherently multiscale and that single-modality models — a protein language model here, a single-cell transcriptomics model there — cannot capture cross-scale causation. AIDO proposes connecting DNA, RNA, protein, cell, and tissue-level foundation models through shared representations, so that a perturbation modeled at one scale propagates to predictions at another. The claimed payoff is better-guided wet-lab experimentation rather than replacement of it.
Why it matters — The virtual cell has become the field’s most-funded unproven idea, and this is now the most citable statement of it in a top-tier clinical journal. Expect it in pitch decks by September. That’s precisely why the disclosure matters: a Nature Medicine citation confers evidentiary weight that a vision paper hasn’t earned. When a partner cites this to justify a platform valuation, the correct follow-up question is which of the three roadmap stages they’ve completed, not whether the vision is compelling.
Sources: Nature Medicine
Insilico’s CEO Wants to Be Judged on the Pipeline, Not the Biobucks
OpinionWhat’s new — In an interview with BioSpace published August 11, Insilico Medicine CEO Alex Zhavoronkov pushed back on headline deal values and made two claims that cut against most of what his sector says publicly.
What he said — Insilico has signed multi-billion-dollar partnerships this year with Eli Lilly, Takeda, Exelixis, and Boehringer Ingelheim among others, running roughly 40 programs on a partner-at-preclinical model with no intention of becoming a clinical or commercial company. Zhavoronkov said his current objective is sustainable profitability and that the company is “reasonably close,” while declining to give guidance. He was blunter on the clinical side: any startup claiming it is better than pharma at running trials with AI is, in his framing, talking nonsense.
Why it matters — Insilico is the only company on anyone’s list with a peer-reviewed Phase IIa readout for a fully AI-discovered, AI-designed asset. That gives this particular scepticism weight the same words wouldn’t carry from a pre-clinical platform. And the second claim is the more consequential one for BD: he is drawing a hard line at discovery. AI compresses target-to-candidate. It does not compress Phase III, and it does not change the underlying biology of efficacy and safety.
We’re thinking — The most useful thing here is the business-model argument hiding inside the modesty. Partnering at preclinical means Insilico converts platform output into revenue before the expensive failure modes arrive — which is exactly how you reach profitability in a field where nobody has an approved product. It also means the deal totals everyone quotes are overwhelmingly back-loaded milestones that will mostly never be paid. When you see a nine- or ten-figure AI partnership headline, the number that tells you something is the upfront and the near-term milestones; that’s the portion the pharma’s technical diligence actually underwrote. Everything after that is an option, priced accordingly.
Sources: BioSpace
TCS launches ADD AgentHub for clinical development and pharmacovigilance
Product launchAnnounced August 17, the platform embeds role-based AI agents into ICSR intake, data entry, coding, review, and literature analysis, with TCS claiming up to 40% efficiency gains in clinical data management and 30% cost savings in end-to-end safety case processing. — Editorially: these are vendor-supplied figures with no independent verification and no named reference customers — and note the timing, an “audit-ready” agentic platform for pharmacovigilance landing while EMA is still deciding whether generative models belong anywhere near critical GMP. The word “audit-ready” is doing an enormous amount of unearned work.
Open-source OpenDDE beats AlphaFold 3 on antibody–antigen — and loses to it on protein–ligand
Model releasePer the technical report, OpenDDE reaches 0.700 on antigen–antibody complexes against AlphaFold 3’s 0.488 and ESMFold2’s 0.581, with independent replication on Tamarind Bio confirming the FoldBench v1 result rather than debunking it — but it scores 0.601 on protein–ligand co-folding, below AlphaFold 3 at 0.649. — Editorially: the independent confirmation is the story, given how many self-reported AF3-beating claims have collapsed under scrutiny. The split result is the more useful signal — frontier structure prediction is fragmenting by task rather than consolidating into one leader, which means “best model” is no longer a coherent procurement question. Ask which interface class you actually need.
Aureka closes $100M Series B for biological foundation models with a lab-in-the-loop architecture
FundingThe August 10 raise was funded in a first tranche exclusively by Granite Asia with a subsequent tranche led by an undisclosed strategic investor, bringing the Laguna Hills– and Shanghai-based company to nearly $200M total since founding in 2023. — Editorially: the architectural claim is the interesting part — Aureka puts its experimental platform inside the model’s training pipeline rather than downstream as a validation checkpoint. That’s a real bet that data generation, not model architecture, is the binding constraint. The capital trajectory is a different question: two years to $200M is several turns behind what Xaira raised at launch in the same category.
Consultancy Announces AI Zone at Conference, Immediately Classifies It as Critical GMP Infrastructure
GENEVA — Global life sciences advisory firm Cartwright Pemberton unveiled its 2026 flagship offering this week: a dedicated AI Zone at an industry exhibition, which the firm has internally designated a mission-critical component of the pharmaceutical quality system.
“The AI Zone is 400 square meters with a coffee station,” said a spokesperson for the firm, who requested anonymity on the grounds that the AI Zone’s confidence score for that statement had returned undecided. “We’ve applied full change control. If anyone moves the standing desks, we log it.”
The announcement arrives alongside Cartwright Pemberton’s newest market forecast, which projects the AI-in-drug-discovery sector growing from a number to a larger number over a period of years, at a compound annual growth rate the firm described as “the one that made the chart go up nicely.” The methodology section consists of a QR code linking to the firm’s booking page.
The forecast’s headline finding — that lead optimization represents approximately half the market — was reportedly derived by asking a large language model what percentage sounded plausible, then rounding.
Elsewhere in the report, Cartwright Pemberton identified a critical emerging risk in pharmaceutical manufacturing: coding agents quietly modifying equipment software under permissions nobody is monitoring. The firm noted that this risk is severe, systemic, and best addressed by purchasing the firm’s new AI Governance Readiness Accelerator, a 12-week engagement in which consultants deploy coding agents to modify equipment software under permissions nobody is monitoring.
Asked whether any of its own AI deployments had been qualified against a defined use case, a defined system version, and a named accountable person, the firm confirmed that all staff had completed the mandatory 40-minute e-learning module and were therefore, in its terminology, “AI certified.”
The certificate does not specify for what.
Cartwright Pemberton also announced it will publish a report next quarter on why 95% of enterprise AI pilots fail to deliver measurable business impact. The report will be produced by an AI pilot.
What we’re watching
The specific thing we expect: at least one additional FDA warning letter or Form 483 citing AI-generated content in GMP documentation before the end of Q1 2027, and — more consequentially — at least one naming an agentic or LLM-based system by function rather than treating “AI” as an undifferentiated category. The early confirming signal is narrower and should be visible sooner: watch FDA’s warning letter database for a citation under 21 CFR 211.22(c) or 211.100 where the inspectional observation describes who reviewed the output and against what criteria, rather than simply noting that AI was used. That phrasing shift would indicate investigators have moved from “AI was involved” to “the human control layer was inadequate for this specific delegation” — which is the framework the industry will then have to build against, regardless of what EMA publishes in Q4.
What would prove us wrong: EMA finalizes Annex 22 on schedule with a workable risk-based pathway for generative models, FDA harmonizes to it, and enforcement stays quiet through Q1 2027 — which would suggest the April 2026 letter was an isolated quality-unit failure at a small manufacturer rather than the leading edge of a pattern. We’d also count ourselves wrong if the next citations land in discovery or clinical rather than manufacturing; our whole thesis is that GMP is where AI accountability gets tested first, because it’s the only part of the value chain with a fifty-year-old apparatus for assigning a signature to a variable process.
Also this week
- Novo Nordisk names AWS its preferred cloud and AI partner, opens London innovation hub — Fierce Biotech
- With tech backing, AI biotechs force biopharma into “fail-fast” drug development — BioSpace
- Stanford and Arc Institute publish the first generative-AI-designed functional viral genomes — Science / Arc Institute
- Pharma’s AI ROI problem isn’t a tech problem — it’s a qualification problem — Pharmaceutical Technology
- The real AI risk in manufacturing is coding agents, not generative drug discovery — Pharmaceutical Technology
- ICON inks Anthropic partnership to deploy Claude into clinical trials — Fierce Biotech
- Assessing shared failures with AI is key to cell and gene therapy’s evolution — BioSpace
- AI in drug discovery market projected to reach $25.0B by 2035 — ResearchAndMarkets (paid market research)
- CPHI Milan adds AI, cold chain and labeling zones to 2026 agenda — Pharmaceutical Technology
Disclosure
Anthropic is referenced in this issue’s link roundup. Claude, Anthropic’s AI assistant, is used in the research and drafting of this newsletter.