Artificial intelligence (AI)–designed drugs have shown early phase 1 success rates of 80% to 90% and phase 2 success rates of approximately 40%, exceeding conventional phase 1 rates of 40% to 65% while roughly matching conventional phase 2 rates, according to a presentation by Marina Chiara Garassino, MD, during the 27th Annual International Lung Cancer Congress.¹
“The implications [of AI in oncology] are enormous at multiple levels. We just need to have some fantasy and increase the possibility that we have to integrate [AI] into our lives,” Garassino, a professor of Medicine and director of the Thoracic Program at The University of Chicago in Illinois, said during the presentation.
AI in Lung Cancer: Highlights
- AI-designed drugs are showing higher early success rates than conventional drug development (80%-90% phase 1, ~40% phase 2), roughly doubling the probability a molecule ultimately reaches market (5%-10% to 9%-18%)
- Aggregate model performance can be misleading — one benchmarked model ranked best of four tested under a drug-blind split but worst of four under a cell-blind split, the scenario that most resembles selecting therapy for a new patient
- Explainable AI features improved physicians' ability to correctly predict true disease control rate (37%) and overall survival (36%) in a clinical usability study, with the largest gains among non-expert oncologists
During the presentation, Garassino reviewed the current landscape of AI applications in oncology across 3 domains: drug discovery and development, clinical decision-support agents, and diagnostic devices, along with the regulatory and implementation challenges facing each.
How is AI being used in drug discovery and development?
Early analyses suggest AI-discovered drugs have approximately twice the probability of reaching market compared with historical rates for conventionally-discovered drugs, an increase from 5% to 10% up to 9% to 18%, attributed in part to faster medicinal-chemist iteration time and quantitative systems pharmacology approaches.
Garassino arranged an AI drug-discovery toolkit into 5 functional classes, each followed by a distinct characteristic failure mode: target generation tools, which tend toward literature bias and rediscover well-studied genes rather than identifying novel targets relevant to rare subgroups; structure-prediction tools, which are limited to static ground-state conformations that are least informative in the setting of acquired resistance; engagement and affinity tools, described as the weakest class due to poor generalization and benchmark leakage, suitable only for triage purposes; molecule-generation tools, which primarily reduce cycle time rather than improve probability of clinical success; and property- and patient-level prediction tools, of which only the first 2 classes currently have a defensible evidentiary role.
Since these model classes fail independently, Garassino noted that concordance across classes may be more informative than confidence within any single model.
Garassino also highlighted TxGemma as an illustrative case of why aggregate performance metrics can be misleading: in an independent benchmark using the GDSC1 dataset, the same model ranked best of 4 models tested under a drug-blind split, but worst of 4 models (per-drug Spearman correlation = 0.13) under a cell-blind split—the split that most closely resembles selecting therapy for a new patient.²
“TXGemma is a multi-agent, and there are all these possibilities in just one multi-agent. So think that these new agents are thinking, reviewing the literature, finding the target through the databases, recreating the target, and doing the binding. It seems like science fiction,” Garassino said.
What progress has been made with autonomous clinical AI agents?
Citing a 2025 Nature Cancer study, Garassino described an autonomous oncology decision-making agent built on a large language model combined with multimodal tools, including narrow-AI pathology analysis, foundation-model radiology segmentation, and molecular-variant and literature-search tools; across 20 fictional multimodal patient cases, the agent reached approximately 87% accuracy.³
Separately, Prof Valmed is the first generative AI tool that has become CE-certified as a Class IIb medical device for supporting clinical decision-making in the European Union, functioning as a retrieval-augmented-generation–based “medical co-pilot” drawing on approximately 2.5 million curated documents, alongside the CME-accredited educational platform Valmed A(I)cademy.1
In the I3LUNG clinical usability study, explainable AI (XAI) features increased the probability that physicians correctly predicted true disease control rate by 37% and true overall survival (OS) by 36% across all participating physicians, with a larger benefit among non-experts than lung cancer experts for both disease control rate (53% vs 23%) and OS (61% vs 14%).
Garassino noted that implementation, interpretability, medical responsibility, and validation and regulation remain unresolved challenges for clinical AI agents, including evolving frameworks, such as the AI Act, Medical Device Regulation, In Vitro Diagnostic Regulation, and guidelines on general-purpose AI, and cited Italian regulatory warnings to physicians regarding improper use of non-specific AI platforms and chatbots for interpreting medical data.
What diagnostic devices are being developed for lung cancer screening?
Garassino also reviewed cough-analysis prescreening tools being studied in populations at high risk for lung cancer. In one dataset evaluating a multimodal model incorporating 60 features, the best-performing model showed a balanced accuracy of 0.694 (SD, 0.130) and F1 score of 0.600 (SD, 0.173) for distinguishing disease progression across treatment lines and radiographic findings; a support vector classifier model showed a sensitivity of 0.845 and specificity of 1.0 for progression-related outcomes, and a sensitivity of 0.975 and specificity of 0.190 for detecting pneumonia-related complications.
References
- Garassino MC. Potential of AI-assisted approaches in lung cancer. Presented at: 27th Annual International Lung Cancer Congress; July 24-25, 2026; Huntington Beach, CA.
- Sada Del Real K, Swamy VS, Arcagni J et al. Foundation models and deep learning for cancer drug response prediction: a framework for data, metrics, and validation. Brief Bioinform. 2026;27(3):bbag225. doi:10.1093/bib/bbag225
- Ferber D, El Nahhas OSM, Wölflein G, et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat Cancer. 2025;6(8):1337-1349. doi:10.1038/s43018-025-00991-6