Post

Why AI-Assisted Drug Compounds Keep Failing Clinical Trials and What the Training Data Has to Do With It

Table of Contents
    Add a header to begin generating the table of contents

    AI accelerates early-stage drug discovery, with AI-discovered molecules clearing Phase I trials at 80–90%, far above the historic ~52% baseline (Int. J. Pharm. Sci. Drug Res., May 2026). But Phase II success stays near 40%, statistically indistinguishable from the historic 37% rate, since models trained on preclinical data miss translational biology.

    AI has genuinely improved early-stage drug discovery performance. It has not closed the translational gap, the space between predicting a molecule’s properties and predicting whether it works in humans, and that distinction matters more than the headlines around either “AI drug discovery” or “AI drug failures” tend to suggest.

    : Scientist examining a test tube while analyzing data on a digital monitor.

    Faster Discovery Has Not Reduced Clinical Risk

    Target identification, molecule generation, and lead optimization, the parts of the pipeline where AI models are trained on the richest, most structured data, are where AI-native biotech has made its most visible gains. Timelines that used to take years have compressed into months. But a faster start doesn’t guarantee a faster finish, and several AI-assisted compounds have entered clinical development with strong early profiles only to run into the question that has always defined this industry: does the mechanism actually produce a clinical benefit in humans.

    Where AI Drug Discovery Excels

    A May 2026 peer-reviewed review in the International Journal of Pharmaceutical Sciences and Drug Research looked at Phase I and Phase II trial outcomes across AI-native biopharma pipelines and found AI-discovered molecules clearing Phase I trials at an 80 to 90% rate, compared to roughly 52% historically. The review’s own interpretation is straightforward: this suggests AI is highly capable of designing or identifying molecules with drug-like properties.

    That’s a real, measurable performance win, and it maps directly onto the kind of data these models are trained on: binding affinity, ADMET properties, synthetic accessibility, and structural biology, all of which are represented extensively and consistently in public and proprietary datasets.

    AI-powered molecular drug discovery with digital analysis & pharmaceutical data.

    This isn’t a coincidence of what happened to be digitized first. Molecular property prediction is, structurally, a good fit for machine learning: bioactivity databases, crystal structures, and screening results give models large, labeled, consistent datasets, and the prediction target (does this molecule bind, is it stable, is it safe in early dosing) is measurable in a lab long before a patient is enrolled. Phase I success is, in a meaningful sense, a chemistry and pharmacokinetics problem, exactly the kind this generation of models was built to solve.

    Where Model Performance Breaks Down

    The same review found Phase II success rates for AI-discovered molecules at roughly 40%, statistically indistinguishable from the historic baseline of roughly 37%. That’s the core of the performance gap: a model can be highly effective at predicting whether a molecule has certain desirable, well-represented properties, and still have no real ability to predict whether the underlying biological mechanism will produce a meaningful clinical outcome in a heterogeneous patient population.

    Strong performance on measurable, well-represented preclinical signals does not automatically translate into strong performance on complex clinical outcomes. Those are two different prediction problems, and only one of them is well-covered by the data these models are trained on.

    The distinction is worth making concrete. Predicting a molecule’s binding affinity or ADMET profile is closer to an interpolation problem: the model operates within a well-sampled chemical space, drawing on thousands of structurally similar prior examples. Predicting whether a biological mechanism will produce a clinical benefit is closer to an extrapolation problem: the model has to reason from imperfect proxies, animal models, cell lines, organoids, to outcomes in genetically and clinically diverse human patients who rarely resemble the clean cohorts used in early discovery work. No amount of additional chemistry data changes the nature of that second problem, because it was never a chemistry problem to begin with.

    What Clinical Failures Reveal

    Individual cases illustrate the broader pattern rather than proving it. One topical AI-derived candidate for atopic dermatitis met its Phase IIa safety and tolerability endpoint but did not achieve its secondary efficacy endpoints in 2023. It illustrates the distinction between early clinical progression and demonstrated efficacy: the candidate advanced through earlier testing but failed to show the expected effect once measured against placebo.

    A separate AI-discovered candidate developed for cerebral cavernous malformation met its Phase II safety endpoint, but the exploratory efficacy trends seen at the higher dose did not hold up on longer-term follow-up, and the program was discontinued in 2025. Neither case establishes on its own that a training-data gap was the cause of the outcome. What both illustrate is the same broader pattern: candidates that performed as expected on the metrics their pipelines were built to predict, then met a different, less predictable test once measured against real patients.

    What ties these cases together isn’t the disease area, the modality, or even the specific company. It’s a drop-off between Phase I and Phase II that lines up closely with where the underlying training data itself thins out, from dense, structured, chemistry-native signal to sparse, context-dependent, clinical signal.

    The Training Data Behind the Translational Gap

    This is not simply a problem of needing more data. One contributing factor to the translational gap between early model performance and clinical outcomes is the mismatch between what these models are trained to predict and the biological and clinical questions that ultimately determine whether a candidate works in humans.

    The data AI models have relatively strong access to sits almost entirely on the preclinical side: molecule structures, binding affinity, ADMET properties, synthetic accessibility, structural biology, and high-throughput screening results. This is dense, structured, and well-suited to the kind of large-scale pattern learning these models are built for.

    Translational gap between preclinical data and clinical real-world data.

    The data that’s much thinner or effectively absent sits on the translational side: disease heterogeneity across real patient populations, mechanistic validation in humans rather than in preclinical models, the differences between animal or cell-based models and the patients who actually enroll in a trial, clinical context and treatment history, and expert reasoning about whether a given biological hypothesis is actually likely to translate. Adding more preclinical data alone is unlikely to close that gap, because the missing signal is fundamentally different in kind, not just in volume.

    Closing the Gap With Expert Feedback

    Closing this gap means building feedback loops around scientific reasoning and experimental outcomes, not just adding more structured chemistry data to the training set. In practice, that looks like:

    1. Expert review of raw experimental outputs, such as assay results, instrument logs, and reaction outcomes, so models learn from clean, well-interpreted signals rather than noisy lab data.
    2. Validation of AI-generated hypotheses, checking a proposed target or mechanism against known biology before it shapes a program’s direction.
    3. Review of experiment plans and next-best-experiment decisions, so a model’s proposed next step gets checked against real scientific judgment.
    4. Process supervision, assessing whether an AI agent’s reasoning, not just its output, was scientifically sound, then capturing that feedback as structured data fed back into subsequent model development.

    This kind of work sits at the highest end of what’s sometimes called task complexity in AI training, where credentialed domain expertise, not just annotation guidelines, is the qualifying skill.

    The value here is that it turns expert judgment into a reusable training signal instead of a one-time gate. When a domain expert flags that a target hypothesis doesn’t account for known disease heterogeneity, or that a preclinical model doesn’t translate the way a pipeline assumes, that judgment is worth capturing in structured, model-usable form, not just acting on once and discarding. Over enough cycles, that’s the kind of signal that can start to close the gap the Phase I and Phase II numbers point to.

    What This Means for ML and Data Strategy Teams

    A few implications follow directly:

    1.  Audit data by pipeline stage:

    A dataset can be enormous and still concentrated almost entirely on the part of the problem AI already solves well, leaving the preclinical-to-clinical spectrum thin at the end that matters most.

    2. Track performance by stage:

    A model that looks strong on one aggregate accuracy figure may simply be strong at Phase I-relevant prediction and untested on anything downstream.

    3. Expert review as data:

    The reasoning behind a judgment call, not just the resulting pass or fail label, is the signal missing from most pipelines.

    4. Study failures like successes:

    A program filed away as a business decision, without a structured account of which biological assumption didn’t hold, is a missed opportunity to feed that lesson back into the next generation of models.

    Conclusion

    Solving this particular problem requires more than faster models or larger training sets. AI drug discovery needs feedback loops that connect model outputs to expert-reviewed experimental and clinical evidence, particularly where biological complexity makes prediction difficult. This is the layer iMerit’s medical AI expert review and data annotation work supports, delivered in part through iMerit Scholars.

    Pharmaceutical and biomedical specialists review hypotheses, experiment plans, and model reasoning, not just final molecular outputs, so expert judgment becomes part of what the next generation of models learns from. The Phase I numbers show what AI has already learned to do well. Closing the gap to Phase II is a question of what these systems still need to learn, and from whom.

    Learn more about how iMerit supports Scientific AI & Autonomous Labs.