Instead of building yet another giant AI to read cancer slides, researchers just taught five existing ones to work as a team, and the combination spotted tumour subtypes, gene mutations, and which patients would respond to immunotherapy better than any of them alone. Ask five experienced pathologists to look at the same tumour slide and you’ll get five slightly different sets of observations. One notices the stroma. One fixates on the immune cells crowding the tumour edge. One flags a nucleus that looks wrong. Put them in a room together and you usually end up with a better answer than any of them would have given alone. That, in essence, is the idea behind ELF, and it has now been tested on more than a thousand cancer patients.
Twenty models, no clear winner
Over the past few years, AI labs have poured enormous resources into pathology foundation models: huge neural networks trained on hundreds of thousands of digitised tissue slides, designed to be adapted later for specific clinical jobs. At least twenty now exist. Here’s the awkward part. When independent researchers benchmarked them, no single model won. One would top the leaderboard for breast cancer subtyping and then flop at predicting a gene mutation in the colon. Another would reverse the pattern. Performance swung around not just between tasks but between different datasets for the same task. For a hospital trying to choose one, that’s a genuine headache. And the usual fix, training a bigger model from scratch, costs a fortune in computing and electricity. So a team led by Ruijiang Li at Stanford University School of Medicine, working with colleagues at Memorial Sloan Kettering, tried the cheaper, stranger option: stop competing and start combining.
How ELF actually works
ELF stands for Ensemble Learning of Foundation models. It takes five of the leading existing models, namely GigaPath, CONCH, Virchow2, H-Optimus-0 and UNI, and runs every slide through all of them. Each produces its own numerical impression of the tissue. ELF then learns to fuse those five impressions into one compact summary of the entire slide. The training used 53,699 whole slide images from 20 different organs, drawn from eleven public datasets. Crucially, ELF wasn’t told “this is cancer of type X” in fine detail. It mostly learned by being shown the same tissue region as described by different models and being pushed to recognise that they were, in fact, the same thing.
The researchers also checked whether the five models were simply agreeing with each other, in which case the whole exercise would be pointless. They weren’t. When the team compared which parts of a slide each model paid attention to, the correlations were strikingly low. The models really were looking at different things.
What it did in testing
The team then threw 125 separate clinical tasks at it, spanning nearly 18,000 patients. On six tumour subtyping challenges, ELF beat the previous best model, TITAN, by an average of 2.2%. On the hardest test, sorting breast tissue into seven categories, the margin was 16.3%, though it’s worth saying that all the models struggled badly there. ELF’s balanced accuracy was just 0.457, meaning it was wrong more often than right. Then came the party trick of computational pathology: predicting molecular results from an ordinary stained slide, no DNA sequencing required. Across 84 combinations of biomarker and cancer type, ELF led the field. For microsatellite instability in colorectal cancer, a finding that directly changes treatment, it reached an average AUC of 0.886 across four patient groups, including three from institutions it had never seen.
The numbers that matter most, though, concern immunotherapy. Checkpoint inhibitors transformed cancer care, but most patients don’t get lasting benefit, and the standard test, PD-L1, is famously unreliable. Across 13 cohorts and eight cancer types, ELF hit a mean AUC of 0.724, against 0.612 to 0.664 for the competition. In lung cancer patients at Memorial Sloan Kettering, combining ELF with PD-L1 and mutation burden lifted performance to 0.722. Among patients the combined model flagged as likely responders, the actual durable response rate rose from 53% to 59%, while the group wrongly written off as non-responders was cut in half. When the team looked at where the model was staring, it was concentrating on immune rich and fibrotic regions, exactly the tissue features oncologists already associate with immunotherapy success. Reassuringly unmysterious.
The bit that surprised me
ELF has 0.15 million parameters. TITAN has 48.5 million. Prov-GigaPath has 86.3 million. The ensemble is roughly 500 times smaller than one of the models it outperforms, because all the heavy lifting happens in the five borrowed models. Nobody had to train a new giant. The cost shows up elsewhere. You now have to run five models per slide. Sequentially that’s about 14 minutes, and running them in parallel on separate GPUs drops it to roughly three and a half. Given that standard tissue staining takes hours to days, that’s not the bottleneck anyone should worry about.
Before anyone gets carried away
The authors are unusually blunt about the limits. Most of the results come from retrospective, often single institution datasets. A stained slide alone probably isn’t enough to make treatment decisions, and genomics and clinical records would need to be folded in. And no prospective trial has tested whether any of this actually helps a real patient in a real clinic. What ELF does establish is a principle worth remembering as AI keeps arriving in medicine. Sometimes the answer isn’t a bigger brain. It’s a second opinion, and a third, and a fourth, and a fifth. The code and model weights are publicly available, which means other labs can start poking holes in it immediately. That’s how this is supposed to work.
Sources:
Luo, X., Wang, X., Eweje, F., Zhang, X., Gomez Marti, J. L., Cascarino, S., Yang, S., Li, Y., Quinton, R., Xiang, J., Ji, Y., Li, Z., Chen, Y., Bergstrom, C., Kim, T., Olguin, F. M., Yuan, K., Abikenari, M., Heider, A., … Li, R. (2026). Ensemble learning of pathology foundation models for precision oncology. Cancer Cell, 44, 1–16. https://doi.org/10.1016/j.ccell.2026.08.008
Lipkova, J., & Kather, J. N. (2024). The age of foundation models. Nature Reviews Clinical Oncology, 21, 769–770. https://doi.org/10.1038/s41571-024-00941-8
Neidlinger, P., El Nahhas, O. S. M., Muti, H. S., Lenz, T., Hoffmeister, M., Brenner, H., van Treeck, M., Langer, R., Dislich, B., Behrens, H. M., et al. (2026). Benchmarking foundation models as feature extractors for weakly supervised computational pathology. Nature Biomedical Engineering, 10, 1113–1123. https://doi.org/10.1038/s41551-025-01516-3























