Sources & Evidence
A companion to “Drug Discovery Has No Magic Wands”
This companion lays out the sources behind the quantitative claims in Deep Phenotype's first article, Drug Discovery Has No Magic Wands, so that anyone who wants to check the argument can trace each figure to its origin. It states where the evidence is strong, where it is suggestive, and where a number is an order-of-magnitude estimate rather than a measured quantity. Where a figure is our own estimate rather than a published result, it is labeled as such and the assumptions behind it are shown in full.
1. The scale of unmet need
The essay opens with the observation that the large majority of recognized human diseases have no approved therapy. Any figure of this kind depends on how diseases are individuated. The most direct primary source is the medical knowledge graph assembled by Huang et al. (2024), which catalogs 17,080 diseases together with their known drug indications and contraindications: 92% of those diseases have no approved drug indication at all. A different estimate is Every Cure’s figure that fewer than 22% of the world’s recognized diseases have an FDA-approved treatment, computed against a universe of roughly 18,000 recognized conditions. This is why the essay states the figure as a range, from around a quarter of diseases down to single digits. Both estimates support the same conclusion: among the diseases medicine can name and catalog, the overwhelming majority have no approved treatment.
2. Why drugs fail
Taxonomy of program failures
The essay’s headline figure, that more than 90% of drugs entering clinical trials fail, is well established. Wong, Siah, and Lo (2019), analyzing over 21,000 compounds, estimate the overall probability of success from Phase I to approval at 13.8%. The more informative question is why they fail. Three independent lines of evidence establish that when drugs fail in the clinic, the dominant cause is inadequate efficacy, rather than a failure of the molecule’s construction or of safety.
Hwang et al. (2016) tracked 640 novel therapeutics that entered pivotal trials between 1998 and 2008, with follow-up through 2015. Of these, 54% failed in clinical development. Among the failures, 57% were due to inadequate efficacy, 17% to safety, 22% to commercial or funding reasons, and 5% for reasons not recorded.
Arrowsmith (2011), analyzing attrition by phase, found the same ordering: in Phase 3 (2007–2010), 66% of failures were for efficacy, 21% safety, and 7% commercial; in Phase 2, 51% were for efficacy.
The AstraZeneca “5R” analysis is the most direct evidence that wrong biology is what remains once the other causes are controlled. Cook et al. (2014) analyzed the company’s 2005–2010 pipeline, the period before the framework was applied, and identified target validation, drug exposure, and patient selection as the recurring failure points. After AstraZeneca imposed explicit discipline on those dimensions, its success rate from candidate nomination to Phase 3 completion rose from 4% (2005–2010) to 19% (2012–2016) (Morgan et al. 2018). The claim about what remained is qualitative, and should be read as AstraZeneca’s own interpretation rather than a statistical result; neither paper partitions the residual failures. But the company’s account of the reformed pipeline is unambiguous: its projects then rarely failed for safety reasons or for missing proof of mechanism, and the major remaining cause of failure was a scientific hypothesis that turned out to be incorrect (Pangalos & Rees 2018).
Together these establish the essay’s core empirical claim: the binding chemistry, whether a molecule engages its intended target, is seldom what sinks a program. Most programs fail because of lack of efficacy.
A fourth line of evidence points the same way from the opposite direction: when the biology is right, success rates rise. Drug programs whose target has human genetic support, a naturally occurring variant linking the gene to the disease, are roughly twice as likely to succeed in the clinic as those without it (Nelson et al. 2015; King et al. 2019). The effect has held up as the datasets have grown, with a relative success of 2.6 overall rising to 3.7 for Mendelian evidence (Minikel et al. 2024), and about two-thirds of the drugs approved by the FDA in 2021 carried such support (Ochoa et al. 2022). What most improves the odds is stronger evidence that the target genuinely drives the disease. Why genetic evidence carries this predictive weight, and how to generate causal evidence of this kind at scale, is the subject of Part 2; we note it here as independent confirmation that the mechanism governs clinical success.
It is worth being precise about what an efficacy failure is, because the category is broader than “wrong mechanism.” A drug can fail to show efficacy for at least three distinct reasons:
the target was wrong: the drug engaged it, but the mechanism does not govern the disease, the “scientific hypothesis that turned out to be incorrect” that Pangalos & Rees (2018) discuss;
the drug did not adequately engage the target, or could not reach it at a tolerable dose, a molecule and tissue-exposure problem (Sun et al. 2022);
the trial itself was flawed: underpowered, mis-dosed, or poorly enrolled (Fogel 2018).
Below we discuss the latter two.
Failures due to insufficient exposure
Kola and Landis (2004) documented that inadequate PK and bioavailability caused roughly 40% of attrition in 1991 but only about 10% by 2000, the payoff of introducing early, high-throughput ADME screening. Sun et al. (2022) report the same trajectory for drug-like properties more broadly, from 30–40% of failures in the 1990s to 10–15% today. Efficacy, the failure mode that depends on getting the biology right, did not improve over the same period.
Sun et al. (2022) also argue that a meaningful share of efficacy and toxicity failures reflect the second cause: the right mechanism pursued with a molecule that cannot achieve adequate exposure in the diseased tissue relative to healthy tissue. They do not quantify how much of the efficacy bucket is attributable to tissue exposure versus wrong mechanism. They offer illustrative case studies rather than a portfolio-level fraction, and note that the relevant tissue-exposure data are not systematically collected.
This failure mode has persisted for the same reason mechanism prediction has not improved. Binding affinity is the ideal computational target: decades of high-throughput assays have produced abundant structured data, and the in vitro readout is fast and cheap. In vitro ADME properties also lend themselves to rapid feedback. Tissue exposure and selectivity in the right human tissue, which Sun et al. argue also affects the efficacy and toxicity balance, behave the opposite way on both counts. There is little training data because it is rarely measured, and there is no fast feedback loop, because the ground truth lives in human tissue and emerges only late and expensively. Here too, computational and AI effort has concentrated where the data and the scorecard already exist, and the harder property remains under-modeled.
Molecules do matter. But progress tracks the availability of fast, cheap, quantitative feedback, and the hardest problems in drug discovery, where the least AI progress has been made, are the ones that lack it.
Trial failure versus program failure
Fogel (2018) documents the third cause in detail: recruitment, enrollment, retention, eligibility criteria, and trial design. Its figures are a useful reminder of how much trial outcome depends on execution. Of 114 trials funded by two UK agencies and recruiting between 1994 and 2002, only 31% met their original recruitment target (McDonald et al. 2006, reported in Fogel 2018), and slow accrual is consistently the single most common reason a trial is terminated early (Fogel 2018).
A trial terminated for under-enrollment usually does not kill the asset; the program is redesigned and re-run. When a drug program is actually abandoned, the dominant causes are efficacy and safety, the biology, as the failure-attribution data above establish (Hwang et al. 2016; Arrowsmith 2011; Wong et al. 2019).
The two places where operational choices most affect outcome, the endpoint and the enrolled population, reinforce the point above. When a trial fails on the wrong endpoint, or enrolls patients who do not carry the target mechanism, that is usually a symptom of insufficient disease understanding. Knowing which readout will reveal benefit, and which patients will respond, is itself a product of understanding the mechanism.
3. Crowding, and the retreat from novel biology
The essay’s claim that the industry concentrates on a narrow set of familiar mechanisms is drawn from L.E.K. Consulting’s 2025 analysis of the global R&D pipeline (Mancuso et al. 2025). Its findings, all for the 2024 pipeline:
Of roughly 13,600 unique drug–target pairs in the preclinical and clinical pipeline, about one quarter are concentrated on just 37 biological targets, about 2% of active targets, each associated with 50 or more drugs. The number of novel biological targets entering the pipeline per year fell from around 100 pre-pandemic to just 30 in 2024.
This is not a shrinking pipeline. The overall pipeline nearly doubled, from about 11,000 active programs in 2015 to about 21,000 by the end of 2024. The retreat from novel targets happened while total activity grew.
L.E.K. also reported that roughly 350 novel targets did enter the pipeline between 2020 and 2024, concentrated in oncology, immunology, metabolism, and neuroscience, with about 70% still in preclinical development. Novel biology is not being ignored entirely. The essay’s point is one of balance: the ratio of effort has tilted heavily toward crowding around known mechanisms.
4. The data chasm
Cellular data scale
The essay’s central chart compares the scale of biological data available today with the scale a causal model of human cell biology would require. The chart is drawn in bytes, which lets the human cell-biology corpus be set against other large datasets on a common axis, including the success story most often invoked, the protein-structure corpus behind AlphaFold. This section gives the published evidence on both sides of that comparison, in two units: bytes, to match the chart and to enable the cross-domain comparison; and cells, perturbations, and cell types, the scientifically meaningful measures of coverage that reveal where the true shortfall lies.
Bytes and coverage do not move together. The single-cell field has amassed a large byte-volume of data drawn from a tiny slice of the biological condition space, so a corpus can look large in bytes while sampling almost none of the perturbations and cell contexts a causal model must cover. Bytes match the chart and enable the AlphaFold comparison; coverage is what the disease-understanding argument turns on. We give both, and note where they diverge.
What has actually been measured, in cells.
Descriptive atlases. CZ CELLxGENE Discover, the largest single aggregation of human single-cell data, held 169.3 million cells across more than 1,550 datasets as of October 2024, of which 93.6 million are unique once duplicated submissions are removed. The Human Cell Atlas consortium reported more than 100 million cells from more than 10,000 donors at its late-2024 milestone. “Hundreds of millions of cells” is right for the aggregate corpus; the deduplicated count is still under one hundred million, which is worth knowing if anyone presses on the phrase.
The largest genetic perturbation screens. Replogle et al. (2022) ran genome-scale Perturb-seq, 2.5 million cells across 9,866 genes, in two cell lines (K562 and RPE1). Xaira Therapeutics has since surpassed this by more than an order of magnitude. Its X-Atlas/Orion (2025) profiled about 8 million cells targeting all protein-coding genes across two cell lines (HCT116 and HEK293T), sequenced deeply to over 16,000 UMIs per cell (several times the depth of prior atlases), and released as a 520 GB download. X-Atlas/Pisces (2026) reached 25.6 million perturbed single-cell transcriptomes across sixteen biologically diverse contexts spanning cell lines, iPSCs, and differentiating iPSCs, the largest genome-wide CRISPRi Perturb-seq compendium reported to date. Together these trained X-Cell, scaled to a 4.9-billion-parameter model (X-Cell-Ultra).
The largest chemical perturbation screen. Tahoe-100M reports more than 100 million cells across roughly 1,100 small-molecule treatments in 50 cancer cell lines.
The largest virtual-cell training set. Arc Institute’s State model reports training on roughly 170 million observational cells together with more than 100 million perturbational cells spanning about 70 cellular contexts. The public Virtual Cell Challenge benchmark built alongside it comprises roughly 300,000 cells and 300 CRISPRi perturbations in a single cell type, H1 human embryonic stem cells.
The perturbational corpus has recently become comparable to the descriptive one in raw cell count, which is what the essay means by saying this has begun to change. It still covers only a small number of cellular contexts against the hundreds of primary human cell types the atlas effort has been enumerating: two in Orion, sixteen in Pisces, fifty lines in the largest chemical screen, roughly seventy contexts in the largest training set. The frontier is starting to move beyond immortalized lines: Pisces already includes iPSCs and differentiating iPSCs, and X-Cell reports zero-shot generalization to primary human CD4+ T cells. But even the largest efforts sample a few tens of contexts, before disease state, time course, and dose are considered at all. That is the sense in which the essay’s “narrow range of cell lines” holds.
The same corpus, in bytes, next to AlphaFold
In byte terms, today’s aggregate single-cell corpus is on the order of 10 to 50 terabytes (the CZ CELLxGENE Census and the Human Cell Atlas). Orion gives a concrete per-cell anchor: 8 million cells in 520 GB is about 65 KB per deeply sequenced cell. The AlphaFold Protein Structure Database, more than 214 million predicted structures spanning essentially all sequenced life (Varadi et al. 2024), occupies only 1 to 2 terabytes. That AlphaFold is the smaller of the two corpora is the point developed below. A 1-to-2-terabyte corpus was enough to largely solve protein folding, because folding is conserved and the data could be pooled across every species at once. The same byte-scale is nowhere close to sufficient for human cell biology, because the space is far larger and because it cannot be assembled by borrowing across species.
How large the space is
Perturbations. The human genome contains roughly 19,400 protein-coding genes (GENCODE 2025). Single-gene loss of function alone defines about 1.9 × 10⁴ perturbations; pairwise combinations define about 1.9 × 10⁸.
Contexts. Tabula Sapiens resolved 475 distinct cell types from 483,152 cells across 24 human tissues, and the Cell Ontology carries on the order of 3,000 cell-type terms. Several hundred is the conservative working figure for distinct human cell types, before disease state, time course, and dose are considered at all.
Multiplying only those two axes, single-gene perturbations across a few hundred cell types, gives on the order of 10⁷ distinct experimental conditions, with no replication and no combinatorial, temporal, or dose dimension included. Admit pairwise perturbations and it is 10¹⁰ to 10¹¹. Set that against what exists: on the order of 10⁴ distinct perturbations, assayed in fewer than a hundred contexts, almost all of them cell lines. The shortfall is roughly three orders of magnitude for the single-perturbation case, and seven or more if one insists on covering combinations.
In bytes, the same conclusion
Today’s cell atlas holds 10 to 50 terabytes drawn from about 10⁴ conditions. A causal dictionary that covered the roughly 10⁷ single-perturbation conditions at comparable depth would run to 10 to 100 petabytes, three to four orders of magnitude beyond today’s atlas, with combinatorial, temporal, and dose dimensions pushing it higher still. This is exactly the gap the chart depicts: from a cell atlas of 10–50 TB today, through a complete cell-type dictionary, to a causal dictionary of 10–100 PB.
The argument depends only on the size of the gap, several orders of magnitude between what exists today and what a causal model would need, which is robust to reasonable variation in the assumptions. It would survive every figure above being wrong by two orders of magnitude in either direction.
None of this addresses the other half of the equation: relating cellular mechanisms to human clinical outcomes. Most human disease is a systems-level dysfunction, a temporal interaction of multiple biologies across diverse cell types, and understanding it requires systems-level measurements that are far less scalable, often requiring living organisms. That is the subject of the next section. Article continues below…
Measuring disease-relevant biology
The essay draws a contrast between biological processes that are conserved across species, where cross-species data transfers and models trained on it generalize, and those that are intrinsically human, where it does not. The clearest poles of that spectrum are protein folding and neurodegeneration, and the contrast is what makes the AlphaFold comparison above more than an analogy.
Protein folding is near-universal.
A protein’s three-dimensional structure is determined far more by physics than by species, and is far more conserved than its sequence. Chothia and Lesk (1986) established that structure diverges much more slowly than sequence, so that proteins sharing as little as 20–25% sequence identity routinely adopt the same fold. This is why a single model trained on the Protein Data Bank, using experimental structures drawn from across the tree of life, predicts structures for any organism. The AlphaFold Protein Structure Database now holds more than 214 million predicted structures, a roughly 500-fold expansion from the 300,000 structures across 21 model-organism proteomes released in 2021, and covers almost the complete UniProt archive (Varadi et al. 2024; Jumper et al. 2021). The folding problem is conserved enough to be learned once, from all of life, and applied anywhere. This applies to folding, not to function: a protein’s interactions and physiological role need not transfer in the same way.
Neurodegeneration is not
Genome conservation is real but uneven. Around 85% of protein-coding genes are conserved between mice and humans, but conservation is not even across the genome, with higher divergence in immune (including microglial) and vascular genes (Boyanova et al. 2026), the biology that human genetics implicates in Alzheimer’s disease. The divergence is concrete at the level of the key disease proteins: rodent amyloid-β differs from human at three residues (R5G, Y10F, H13R), altering its processing and aggregation, which is why rodents do not normally develop amyloid pathology; and developmental differences in tau splicing prevent mice from recapitulating the human tau-isoform shifts central to disease (Boyanova et al. 2026).
The consequence is measurable. Across more than 100 genetic mouse models of Alzheimer’s disease, the two most widely used (5xFAD and APP knock-in) replicate only about 30% of the human protein alterations, rising to about 42% when tau and splicing pathology are added (Yarbro et al. 2025). The best available models capture under half of the human molecular disease, the quantitative form of the essay’s statement that rodents do not get Alzheimer’s and non-human primates do not recapitulate ALS.
Some processes, core metabolism for instance, are conserved across mammals and travel reasonably well. But the diseases where progress has been slowest, neurodegeneration above all, sit at the human-specific end, where cross-species data cannot substitute for measurement in human systems, and where those measurements are the most expensive, least available, and most ethically constrained to obtain. This is why the byte comparison with AlphaFold above is only part of the story. The same scale of data solved folding because folding pools across species, and cannot solve human neurodegeneration because it does not.
5. What AI can and cannot compress in the clinic
The essay’s final chart divides the clinical-development timeline into three parts and asks how much of it AI can realistically remove. The three parts are: the IND-enabling preclinical work required before a drug can enter humans; the operational overhead of running a trial (startup and site activation, patient recruitment, monitoring, data cleaning, and database lock); and the in-life biological observation window, the time patients must actually be dosed and followed for the clinical endpoint to mature.
Two features of that timeline are well documented. First, it is long and dominated by patient time. The canonical model of per-phase durations is Paul et al. (2010): 12 months of preclinical development, then 18, 30, and 30 months for Phases 1, 2, and 3, and 18 months from submission to launch. Measured durations from DiMasi et al. (2016) agree in magnitude, at a mean of 80.8 months from first-in-human to submission and 16.0 months of regulatory review. DiMasi’s individual phase means (33.1, 37.9, and 45.1 months) cannot be added together, because the phases overlap; their sum exceeds the measured first-in-human-to-submission interval by three years. Recruitment by itself has grown from an average of about 13 months in 2008–2011 to about 18 months in 2016–2019 (Brøgger-Mikkelsen et al. 2022).
The operational layer is substantial and the single largest source of unanticipated delay; slow enrollment is consistently the leading cause of trials running late (Fogel 2018). Industry and vendor projections frequently claim that AI will cut trial timelines by 30–50%. Those figures largely describe the operational layer, faster recruitment and automated data handling, and cost, rather than the total calendar time to a matured endpoint. They largely ignore the in-life observation window. To learn whether a drug changes the course of a disease, patients have to be treated and then followed until the outcome matures, an interval set by human biology, not by computation. This is the floor beneath any compression estimate: even if every operational inefficiency were eliminated, the biological waiting period remains. Netted against that floor, the share of the end-to-end timeline that AI can realistically remove is smaller. The chart puts it at roughly 11–17%, with the in-life observation window left essentially untouched.
That range is a derived estimate rather than a published benchmark, so the full arithmetic is set out here. The window divides as follows.
End-to-end window: 108 months. Candidate nomination to launch, built from Paul et al. (2010): 12 months of preclinical development, 18 + 30 + 30 months of clinical phases, and 18 months from submission to launch.
In-life observation: 54 months. Taken from Phases 2 and 3 together, 60 months in the Paul model, less the portion of that calendar time that is site activation and close-out rather than dosing and follow-up. Recruitment overlaps dosing, since patients enrolled early are already being followed while later ones are still being enrolled, so the in-life and operational buckets cannot simply be summed. We have assigned the overlap to in-life.
Regulatory review: 16 months. From DiMasi et al. (2016).
IND-enabling preclinical work: 12 months. From the Paul model.
Operational overhead: 26 months, taken as the residual (108 − 54 − 16 − 12), covering site activation, the enrollment time not already counted as in-life, monitoring, data cleaning, and database lock. For scale, Tufts CSDD benchmarking has put median time from protocol approval to full site activation at roughly 17 months, and last-patient-last-visit to database lock at about 37 days, the latter from an industry survey rather than the peer-reviewed literature.
Against that division, we assume AI can remove:
40–60% of the operational overhead, or 10.4–15.6 months. This is the aggressive end of what patient identification, site selection, and automated data handling can plausibly deliver, and it is deliberately close to the 30–50% vendor claims, since those claims describe this layer specifically.
10–25% of the IND-enabling work, or 1.2–3.0 months.
None of the in-life observation window, and none of regulatory review.
That totals 11.6 to 18.6 months out of 108, or 11% to 17%. If the in-life share is 40% of the window rather than 50%, the operational residual grows and the compressible fraction rises to roughly 15–23%. And if patient selection or better efficacy biomarkers genuinely shorten the follow-up interval, part of the in-life window becomes compressible after all. That is the route the essay points to as the only way through the biological floor.
References
Arc Institute. State: Arc Institute’s first virtual cell model. arcinstitute.org/news/virtual-cell-model-state (accessed 2026). Benchmark described in the Virtual Cell Challenge, Cell, 2025.
Arrowsmith J. Trial watch: Phase III and submission failures, 2007–2010; and Phase II failures, 2008–2010. Nature Reviews Drug Discovery. 2011;10(2):87; 10(5):328–329.
Boyanova S, Katsouri L, Krupic J, Hair K, Wang S-H, Wiseman FK. Considerations for the selection and phenotyping of mouse models for the study of Alzheimer’s disease. STAR Protocols. 2026;7(3):104633. doi:10.1016/j.xpro.2026.104633
Brøgger-Mikkelsen M, Zibert JR, Andersen AD, et al. Changes in key recruitment performance metrics from 2008–2019 in industry-sponsored phase III clinical trials. PLoS One. 2022;17(7):e0271819. doi:10.1371/journal.pone.0271819
Cell Ontology. The Cell Ontology in the age of single-cell omics. arXiv:2506.10037, 2025. A preprint, not peer-reviewed. arxiv.org/abs/2506.10037
Chothia C, Lesk AM. The relation between the divergence of sequence and structure in proteins. EMBO Journal. 1986;5(4):823–826. doi:10.1002/j.1460-2075.1986.tb04288.x
Cook D, Brown D, Alexander R, et al. Lessons learned from the fate of AstraZeneca’s drug pipeline: a five-dimensional framework. Nature Reviews Drug Discovery. 2014;13(6):419–431. doi:10.1038/nrd4309
CZ CELLxGENE Discover: a single-cell data platform for scalable exploration, analysis and modeling of aggregated data. Nucleic Acids Research. 2025;53(D1):D886. doi:10.1093/nar/gkae1142
DiMasi JA, Grabowski HG, Hansen RW. Innovation in the pharmaceutical industry: new estimates of R&D costs. Journal of Health Economics. 2016;47:20–33. doi:10.1016/j.jhealeco.2016.01.012
Every Cure. The Problem. everycure.org/the-problem (accessed 2026).
Fogel DB. Factors associated with clinical trials that fail and opportunities for improving the likelihood of success: a review. Contemporary Clinical Trials Communications. 2018;11:156–164. doi:10.1016/j.conctc.2018.08.001
GENCODE. Reference annotation for the human and mouse genomes, 2025 release. Nucleic Acids Research. 2025;53(D1):D966.
Huang AC, Hsieh T-HS, Zhu J, Michuda J, Teng A, Kim S, et al. X-Atlas/Orion: genome-wide Perturb-seq datasets via a scalable Fix-Cryopreserve platform for training dose-dependent biological foundation models. bioRxiv. 2025. doi:10.1101/2025.06.11.659105
Huang K, Chandak P, Wang Q, et al. A foundation model for clinician-centered drug repurposing. Nature Medicine. 2024;30(12):3601–3613. doi:10.1038/s41591-024-03233-x
Human Cell Atlas. Cellular atlases are unlocking the mysteries of the human body. Nature. 2024. doi:10.1038/d41586-024-03552-6
Hwang TJ, Carpenter D, Lauffenburger JC, et al. Failure of investigational drugs in late-stage clinical development and publication of trial results. JAMA Internal Medicine. 2016;176(12):1826–1833. doi:10.1001/jamainternmed.2016.6008
Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. doi:10.1038/s41586-021-03819-2
King EA, Davis JW, Degner JF. Are drug targets with genetic support twice as likely to be approved? PLoS Genetics. 2019;15(12):e1008489. doi:10.1371/journal.pgen.1008489
Kola I, Landis J. Can the pharmaceutical industry reduce attrition rates? Nature Reviews Drug Discovery. 2004;3(8):711–716. doi:10.1038/nrd1470
Mancuso M, Jacquet P, Brau R, Dhulesia A, Srinivasan A. Is biopharma doing enough to advance novel targets? L.E.K. Consulting Executive Insights. May 15, 2025; Figure 1b updated 29 July 2026 (targets with 50+ associated drugs changed from 38 to 37). lek.com
McDonald AM, Knight RC, Campbell MK, et al. What influences recruitment to randomised controlled trials? A review of trials funded by two UK funding agencies. Trials. 2006;7:9. doi:10.1186/1745-6215-7-9
Minikel EV, Painter JL, Dong CC, Nelson MR. Refining the impact of genetic evidence on clinical success. Nature. 2024;629:624–629. doi:10.1038/s41586-024-07316-0
Morgan P, Brown DG, Lennard S, et al. Impact of a five-dimensional framework on R&D productivity at AstraZeneca. Nature Reviews Drug Discovery. 2018;17(3):167–181. doi:10.1038/nrd.2017.244
Nelson MR, Tipney H, Painter JL, et al. The support of human genetic evidence for approved drug indications. Nature Genetics. 2015;47(8):856–860. doi:10.1038/ng.3314
Ochoa D, Karim M, Ghoussaini M, et al. Human genetic evidence supports two-thirds of the 2021 FDA-approved drugs. Nature Reviews Drug Discovery. 2022;21(8):551. doi:10.1038/d41573-022-00120-3
Pangalos MN, Rees S. AstraZeneca R&D: improving drug development productivity with better predictivity. Drug Discovery World. Summer 2018:18.
Paul SM, Mytelka DS, Dunwiddie CT, et al. How to improve R&D productivity: the pharmaceutical industry’s grand challenge. Nature Reviews Drug Discovery. 2010;9(3):203–214. doi:10.1038/nrd3078
Replogle JM, Saunders RA, Pogson AN, et al. Mapping information-rich genotype–phenotype landscapes with genome-scale Perturb-seq. Cell. 2022;185(14):2559–2575. doi:10.1016/j.cell.2022.05.013
Sun D, Gao W, Hu H, Zhou S. Why 90% of clinical drug development fails and how to improve it? Acta Pharmaceutica Sinica B. 2022;12(7):3049–3062. doi:10.1016/j.apsb.2022.02.002
Tabula Sapiens Consortium. The Tabula Sapiens: a multiple-organ, single-cell transcriptomic atlas of humans. Science. 2022;376(6594):eabl4896. doi:10.1126/science.abl4896
Tahoe-100M: a giga-scale single-cell perturbation resource. bioRxiv preprint, 2025. doi:10.1101/2025.02.20.639398
Tufts Center for the Study of Drug Development. Study start-up and site activation benchmarks; and Tufts CSDD / Veeva clinical data management industry survey. Industry benchmarking analyses, not peer-reviewed.
Varadi M, Bertoni D, Magana P, et al. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Research. 2024;52(D1):D368–D375. doi:10.1093/nar/gkad1011
Wang C, Karimzadeh M, Ravindra NG, Bounds LR, Alerasool N, Huang AC, et al. X-Cell: scaling causal perturbation prediction across diverse cellular contexts via diffusion language models. bioRxiv. 2026. doi:10.64898/2026.03.18.712807
Wong CH, Siah KW, Lo AW. Estimation of clinical trial success rates and related parameters. Biostatistics. 2019;20(2):273–286. doi:10.1093/biostatistics/kxx069
Yarbro JM, Han X, Dasgupta A, Yang K, Liu D, Shrestha HK, et al. Human and mouse proteomics reveals the shared pathways in Alzheimer’s disease and delayed protein turnover in the amyloidome. Nature Communications. 2025. doi:10.1038/s41467-025-56853-3









Excellent article with a very important message! Thank you for writing it, Daphne. One other point wrt operational and in-life biological work. Often, the patient recruitment and dosing seem to be conducted by contractual partners (research hospitals). where efficiency may not be in their best interests?