AI in Medicine & Oncology Digest

July 30, 2026 — Re-cutting a Negative Trial, and the Limits of Retrospective AI
Curated by Dr. Allan Pereira — Moffitt Cancer Center

Top 4 Updates

#1
Source: The Lancet Digital Health  |  Authors: Bertsimas & Margonis (MIT / Memorial Sloan Kettering / Charité), senior author Singer (MSK); reanalysis of the multicentre EORTC STRASS randomised trial  |  Published: July 27, 2026
Score 13/20PEER-REVIEWEDinternal-retrospectiveONC: primary
Lancet Digital Health base 8 + internal retrospective derivation on randomised-trial data with no external cohort (+1) + addresses a real treatment-selection gap with rigour (+2) + oncology primary (+2) = 13
STRASS, the only completed randomised trial of preoperative radiotherapy in retroperitoneal sarcoma, randomised 266 patients to radiotherapy plus surgery versus surgery alone and was negative overall; its subgroup analysis and the STREXIT extension have since been read as supporting radiotherapy for liposarcoma broadly. Here a random survival forest was trained on all 266 randomised patients to predict 5-year abdominal recurrence-free survival under each treatment option from pretreatment variables, and an optimal policy tree then partitioned patients by predicted benefit. The tree produced seven subgroups; three (152 of 266 patients) were predicted to benefit, and two reached statistical significance — well-differentiated liposarcoma in patients aged 60 or younger, and dedifferentiated liposarcoma treated with curative-intent surgery. In those two subgroups combined, 3-year abdominal recurrence-free survival was 79% (95% CI 70–89) with radiotherapy versus 58% (48–72) without (HR 0.40, 95% CI 0.22–0.71, p=0.0016), with cumulative abdominal recurrence falling from 33% to 17%. By contrast, the STRASS- and STREXIT-defined subgroups showed no significant improvement when re-evaluated within STRASS.
Why it matters: A negative trial is not the same as an ineffective treatment, and this is a disciplined attempt to answer the harder question of who benefits — using randomised data, which removes the confounding that sinks most retrospective AI treatment-selection work. If it holds, it argues for narrower rather than broader use of preoperative radiotherapy in retroperitoneal sarcoma.
Limitations: The policy tree was both derived and evaluated in the same 266 randomised patients, with no external validation cohort, so the subgroups are hypothesis-generating rather than practice-defining; the authors are explicit that focused confirmatory trials are needed before any of this becomes standard of care.
Post angle: AI re-cuts a negative sarcoma RCT and finds the minority who actually benefit from preop RT — but it was derived in-sample, so it is a trial hypothesis, not a treatment rule.
#2
Source: Science Advances  |  Authors: Chen & Zheng, senior author Lin (South China University of Technology; Sun Yat-sen University Cancer Center)  |  Published: July 29, 2026
Score 10/20PEER-REVIEWEDinternal-retrospectiveONC: primary
Science Advances base 6 + internal retrospective validation in a single case–control cohort (+1) + novel first-in-class low-cost detection platform with a plausible clinical path (+1) + oncology primary (+2) = 10
The artificial perception system (APS) pairs a structurally defined DNA–carbon nanotube sensor array with machine-learning classifiers: the array emits multichannel fluorescence 'fingerprints' from serum, which the models decode into a disease class. Across 253 serum samples spanning liver, lung and ovarian cancer plus non-cancer controls, APS achieved a mean sensitivity of 89% and specificity of 96%; early-stage lung cancer was detected at 92% sensitivity and 95% specificity, at an estimated cost of roughly $4 per test. SHAP and Mantel analyses were used to probe which sensor channels drove classification, which the authors argue supports biological plausibility rather than pure pattern-matching. What this demonstrates is analytical feasibility and cost in a modest case–control cohort — not screening performance in an unselected population.
Why it matters: Cost is the binding constraint on multicancer early detection in most of the world, and a ~$4 sensor-plus-model assay is a genuinely different economic proposition from sequencing-based liquid biopsy. Whether the signal survives contact with a real screening population is the entire question.
Limitations: 253 samples in what appears to be a single-source case–control design, with no independent external cohort and no unselected screening population; case–control sensitivity and specificity systematically flatter real-world performance, where low cancer prevalence would drive positive predictive value far below what these numbers suggest.
Post angle: A ~$4 nanotube-plus-ML serum test hits 89%/96% across three cancers — impressive proof of concept, but it is case–control, so do not read it as screening performance.
#3
Source: npj Digital Medicine  |  Authors: Zhu (KCL / Oxford), senior author Oliver (Oxford); multi-site PSYCHS interview cohort  |  Published: July 23, 2026
Score 10/20PEER-REVIEWEDinternal-retrospective
npj Digital Medicine base 7 + internal retrospective validation against researcher ratings, multi-site but no independent cohort (+1) + addresses a recognised safety and generalisability gap with unusual rigour (+2) = 10
Eleven open-weight LLMs were evaluated on 678 partial PSYCHS interview transcripts from 373 participants (77.7% at clinical high risk for psychosis, CHR-P), asked both to infer CHR-P status and to estimate severity and frequency across 15 symptom domains against researcher-rated scores. The largest models performed best — Llama-3.3-70B reached accuracy 0.80 with sensitivity 0.93 but specificity of only 0.58 — while symptom-score agreement with human raters was good (ICC 0.74 and 0.75). Generated summaries were largely faithful to the source transcripts, with clinically relevant confabulation in 3% of cases, and errors skewed toward over-pathologising non-clinical experiences. Performance was broadly similar across demographic groups but varied by site, and smaller models were competitive at substantially lower compute cost.
Why it matters: This is the rarer kind of LLM-in-medicine paper: it reports the numbers that actually determine whether a model is deployable — specificity, confabulation rate, demographic parity and site-to-site variation — instead of a single headline accuracy. The 0.93/0.58 sensitivity–specificity split is the honest story, and it is one many benchmark papers would have buried.
Limitations: Retrospective transcript analysis benchmarked against researcher ratings, not against clinician decisions or patient outcomes, and with no independent replication cohort; a 77.7% CHR-P base rate is far above any real clinic, and 0.58 specificity would generate a large false-positive burden if used for triage.
Post angle: Open-weight LLMs rate psychosis-risk interviews at 93% sensitivity but 58% specificity, with 3% confabulation — a model of how to report LLM performance honestly.
#4
Source: npj Digital Medicine  |  Authors: Mayer (Heidelberg), senior authors Mahal & Ditzen (Heidelberg / Zurich)  |  Published: July 29, 2026
Score 9/20PEER-REVIEWED
npj Digital Medicine base 7 + no model validation, this is a human-subjects experiment (0) + addresses a recognised gap in how patients experience AI-mediated communication, with a physiological endpoint (+2) = 9
163 healthy participants took part in standardised simulated consultations in which bad news was delivered either by a human physician (in person or by video call) or by an 'AI physician' presented as a chatbot or an avatar. Crucially, the AI conditions used a Wizard-of-Oz design: participants were told they were interacting with an AI while a human actually controlled the interaction, so the study measures the effect of believing the clinician is AI, not the performance of any real system. Human-delivered consultations produced higher stress than the AI-styled formats on both subjective ratings and salivary cortisol, and greater perceived credibility was associated with stronger stress responses. Memory retrieval was lowest in the AI-chatbot condition.
Why it matters: Lower patient stress is often treated as an unambiguous win for AI-mediated communication, and this suggests it may partly reflect the encounter being taken less seriously — participants who found the situation more credible were more stressed, and those in the chatbot arm remembered less. That trade-off matters most in exactly the conversations oncology has every day.
Limitations: Healthy volunteers in a simulated scenario, not patients receiving real diagnoses, and no actual AI system was tested; a single experiment of n=163 with modality effects that may not survive in a real clinical relationship.
Post angle: Patients were less stressed when they thought a physician was AI — but they also remembered less, and credibility tracked with stress. A caution for anyone automating serious conversations.

Also Worth Watching

  1. Caristo Diagnostics (company announcement of a regulatory action) — the company announced on July 29 that FDA granted De Novo authorisation (stated as DEN250042) for CaRi-Heart, which derives a fat attenuation index from pericoronary fat on routine coronary CT angiography and outputs a 10-year cardiovascular mortality risk; the underlying biomarker has prior peer-reviewed validation and ORFAN registry outcomes data. Flagged rather than led: as of July 30 the FDA De Novo database returns no record for DEN250042 (database last updated July 27), so this is verified only against the company's own release plus trade-press coverage, not a primary FDA source.
  2. Science (peer-reviewed) — Hicformer, a deep-learning framework applied to GAGE-seq single-cell data from post-mortem Alzheimer's and control brain tissue, shows that 3D genome architecture is necessary to predict cell-type-specific, disease-relevant expression changes; included as a Tier-1 item, but this is mechanistic genomics using AI as a tool, with no diagnostic or clinical claim.
  3. NAR Cancer (peer-reviewed) — a supervised classifier plus attention-guided autoencoder classifies cancer stem-like cell states and deconvolves bulk RNA across more than 25,000 tumour profiles (TCGA, PRECOG, relapse and checkpoint-inhibitor cohorts), with CSC abundance correlating with worse disease-free survival and reduced immunotherapy efficacy; a discovery tool on public retrospective data, released as a Python package, with no prospective or clinical testing.
  4. JAMA (Viewpoint / expert commentary) — a Vanderbilt informatics group's argument for rethinking clinical decision support in the generative-AI era; opinion, not data, but a Tier-1 venue signal on where CDS implementation thinking is heading.
  5. JAMA (Viewpoint / expert commentary) — a Viewpoint on the privacy, regulatory and clinical implications of patients disclosing health information to consumer AI chatbots, co-authored by an oncology faculty member; opinion, no empirical data, but the topic pairs directly with the consultation-perception study above.
  6. DEN Open (peer-reviewed) — in a 63-patient single-centre retrospective pilot, switching from white-light to linked-color imaging cut median false-positive CAD detections from 5 to 2 per case (p<0.001) while both modalities maintained 100% lesion-detection sensitivity; small, non-randomised and run in a fixed white-light-then-LCI sequence, and the authors themselves call for prospective counterbalanced confirmation — but a reminder that image acquisition, not just the model, drives CADe performance.
← Back to all digests
Curated by Dr. Allan Pereira — Moffitt Cancer Center · @DrAllanPereira