Science1 publisher3 min readPublished
Radiomics for brain tumor recurrence: 0.91 AUC, one dataset, and a scaling knob
A Scientific Reports workflow reports a held-out AUC of 0.91 against 0.81 for PCA and 0.75 for mRMR on the UMMC benchmark. The authors call the findings preliminary and externally unvalidated.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The paper "A new radiomics-based approach for predicting brain tumor recurrence" by Olatunde, Oyetunde, Khasawneh et al. was published in Scientific Reports (2026).
- The study combined MRI-derived radiomic features with clinical variables to predict brain tumor recurrence after radiotherapy.
- The evaluated workflow, designed for limited-sample settings, identifies the most predictive feature family, incorporates a categorical feature based on primary tumor location, and applies rank-correlation-based hierarchical clustering within the selected feature family.
- A categorical scaling parameter was introduced to reduce the dominance of one-hot encoded clinical variables over scaled numerical radiomic features.
- On the University of Mississippi Medical Center (UMMC) benchmark dataset, the proposed workflow achieved a held-out test ROC AUC of 0.91, compared with 0.81 for PCA and 0.75 for mRMR baselines, and outperformed previously reported methods on the same dataset.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
A group publishing in Scientific Reports has put forward a radiomics workflow for predicting brain tumor recurrence after radiotherapy whose central argument is about the feature space rather than the classifier [1][2]. On the University of Mississippi Medical Center benchmark dataset, the pipeline reports a held-out test AUC of 0.91 against 0.81 for a principal component analysis baseline and 0.75 for mRMR feature selection [5].
The stated diagnosis is familiar to anyone who has tried to fit imaging features to a clinical endpoint: the high-dimensional radiomic feature space raises overfitting risk and cuts generalizability, and it does so worst on small samples [6]. The two standard escapes each cost something. PCA compresses the space but makes the resulting features hard to interpret, and direct feature selection is distorted by multicollinearity among radiomic variables, which are often near-duplicates of one another [7].
The proposed answer is structural discipline instead of a new model. The workflow picks the single most predictive feature family, adds a categorical variable for primary tumor location, and then applies rank-correlation-based hierarchical clustering inside the selected family to collapse redundant features [3]. Restricting the clustering to one family is the load-bearing choice: it keeps features interpretable in the way PCA components are not, while attacking the collinearity that trips up ranking-based selection [3][7].
The second contribution is smaller and more practical. The authors introduce a categorical scaling parameter specifically to stop one-hot encoded clinical variables from dominating the scaled numerical radiomic features [4]. That is an honest admission about preprocessing that most papers skip. A binary indicator column and a standardized texture feature do not carry the same effective weight in a distance-based or regularized model, and with a handful of clinical dummies against dozens of radiomic columns, the encoding choice can quietly decide which signal the model sees.
Now the discount rate. The margins are 0.10 AUC over PCA and 0.16 over mRMR, on one benchmark dataset, at a single held-out evaluation [5][10]. The published abstract does not report how many patients the UMMC dataset contains, how large the held-out split was, confidence intervals on any of the three AUC figures, or whether the categorical scaling parameter was tuned strictly inside cross-validation [11]. In a small-sample setting, which is exactly the setting the paper is designed for, a 0.10 AUC gap can turn on a few patients changing sides, and a newly introduced scaling hyperparameter is precisely the kind of knob that can absorb information from the test split if the protocol is loose. The authors do not oversell: they state the findings are preliminary and require external validation before clinical adoption [8]. The work reports no specific grant funding and no competing interests [9].
What to watch: whether the full text, which is open access under a Creative Commons NonCommercial NoDerivatives license, reports the cohort size and the tuning protocol for the scaling parameter [12][11]. Then whether the pipeline survives a second institution's scans, where scanner and protocol differences usually destroy radiomic feature stability. Until an external cohort exists, the transferable result here is not 0.91; it is the two design habits, family-restricted clustering and explicit categorical scaling, which cost nothing to test against your own baseline [3][4].