Commentary|Articles|September 24, 2026

Positioning Artificial Intelligence to Help Explain Unexpected Cancer Trial Outcomes

Fact checked by: Caroline Seymour
Listen
0:00 / 0:00

Novel approaches to the use of AI in oncology may require discussion and ultimately modification of existing proprietary and regulatory barriers.

The role of artificial intelligence (AI) in multiple aspects of health care, including oncology, continues to rapidly expand.1 It is the rare month that goes by without a major announcement of a new product or upgrade of an existing platform that might be employed to improve operational efficiency or assist in the detection or management of an ever-expanding variety of human conditions.

In this commentary, another novel approach to the use of AI technology is proposed, one that surely requires considerable discussion and ultimately modification of existing proprietary and regulatory barriers if it is to proceed from a basic concept to a meaningful use case.

When randomized trials are developed to study an existing standard-of-care (SOC) clinical option in cancer care vs an investigative approach, certain assumptions are required in the design to ensure comparable populations within the study arms. Further, suggested potential results based on existing data are considered to select appropriate sample sizes to permit a meaningful statistical outcome. Finally, from an ethical perspective it is important that equipoise is maintained and there are no known or suspected reasons to believe that one approach will be superior (or conversely inferior) to other arms in the study.

But what happens when the objective findings of a cancer study seriously challenge those assumptions, when an experimental regimen is actually revealed to be substantially inferior to the SOC “control arm”, or one study arm results in major improvement in overall survival while exhibiting a far smaller effect on progression-free survival and where there is no obvious explanation for the surprising observation?

Two previously published phase 3 randomized trials from the domain of gynecologic oncology represent examples of the types of studies highlighted above.2,3 For several reasons it is unlikely that answers to the relevant questions being posed will ever be forthcoming, including the substantial time that has passed from the reporting of results and the almost certain absence of either paper or electronic medical records. However, despite these facts the trial result details provide a provocative window into what might very well have been possible if necessary legal, regulatory, and ethical review permitted such an effort.

The first report was that of a phase 3 randomized study comparing a novel antineoplastic agent, canfosfamide, a glutathione analogue prodrug that demonstrated activity in in vitro ovarian cancer models.2 A phase 2 trial of the agent (n = 34) in platinum-resistant ovarian cancer (PROC) reported in 2005 revealed a 15% objective response rate with 50% of individuals achieving what was defined as “disease stabilization”, a respectable outcome at that time within this highly chemotherapy-resistant population.4 The study did not show an excessive drug toxicity profile and the median survival of the entire population was 423 days; again, not an unreasonable outcome in this setting.

However, in the subsequently conducted phase 3 trial in 461 research subjects which compared single-agent canfosfamide with either single-agent pegylated liposomal doxorubicin (PLD) or topotecan (whichever agent was not employed as therapy for the patient in the second-line setting) as third-line treatment of PROC, the investigative regimen was associated with a statistically significant inferior outcome, both in progression-free survival ([PFS] median 2.3 months vs 4.3 months, P < .001) and overall survival ([OS] median 8.5 months vs 13.5 months, P < .001).2

In addition to the still unanswered question as to why this study was not stopped early when inferior survival trends might initially have been observed, it remains unknown why this agent with legitimate laboratory-based biological as well as documented clinical activity in the setting of PROC would be associated with inferior survival when compared with two drugs with seemingly equivalent (modest) utility in the identical circumstance.

Were the trial results simply a rare outcome due to unrecognized imbalances within the trial population, despite best efforts to adequately control for this important research issue through the process of randomization? Or was something else going on here, perhaps a toxicity, drug interaction, or influence of a relevant comorbidity that was underappreciated in preclinical evaluation or prior clinical studies?

And now the point of including this study in the current commentary: Could an AI-based examination of the medical records following, or possibly even prior to patient-study entry discover an overlooked detail that could not only explain the surprising results, but also be useful in future drug development efforts? A reasonable question to ponder.

The second study to be highlighted was a phase 3 randomized trial of single-agent PLD vs topotecan as second-line therapy of ovarian cancer.3 While the study was conducted in both platinum-resistant and platinum-sensitive recurrent ovarian cancer (defined as disease recurrence 6 months following the completion of primary chemotherapy), the focus here will be on the potentially persistently platinum-sensitive patient population. In this group treatment with PLD was associated with a modest 5.6-week (28.9 weeks vs 23.3 weeks) improvement in PFS but a striking 36.9-week (108 weeks vs 71.1 weeks) benefit in OS. Considering the observation that it is far more common to observe a trial-associated improvement in PFS with diminishing or no effect on OS due to therapies received after patients are removed from a study, the survival outcome in this experience is even more surprising.

Could this outcome reflect the fact patients receiving the less marrow toxic PLD were far more likely to successfully be retreated with carboplatin? Remember, this was the potentially “platinum-sensitive” patient subgroup compared with the more marrow toxic topotecan, or are there other reasons for the observation buried within the medical record? And, again, if an AI-based exploration was able to help establish a meaningful explanation, might this be of value in future clinical investigative efforts as well as potentially informing SOC clinical practice in the management of recurrent ovarian cancer? 

References

  1. Li J, Zhang L, Yu Z, et al. The impact of AI on modern oncology from early detection to personalized cancer treatment. npj Precis. Onc. Published online January 24, 2026. doi:10.1038/s41698-026-01276-6
  2. Vergote I, Finkler N, del Campo J, et al. Phase 3 randomized study of canfosfamide (Telcyta, TLK286) versus pegylated liposomal doxorubicin or topotecan as third-line therapy in patients with platinum-refractory or -resistant ovarian cancer. Eur J Cancer. 2009;45(13):2324-2332. doi:10.1016/j.ejca.2009.05.016
  3. Gordon AN, Fleagle JT, Guthrie D, et al. Recurrent epithelial ovarian carcinoma: a randomized phase III study of pegylated liposomal doxorubicin versus topotecan. J Clin Oncol. 2001;19(14):3312-3322. doi:10.1200/JCO.2001.19.14.3312
  4. Kavanagh JJ, Gershenson DM, Choi H, et al. Multi-institutional phase 2 study of TLK286 (TELCYTA, a glutathione S-transferase P1-1 activated glutathione analog prodrug) in patients with platinum and paclitaxel refractory or resistant ovarian cancer. Int J Gynecol Cancer. 2005;15(4):593-600. doi:10.1111/j.1525-1438.2005.00114.x

Related to this article