In a prospective cohort of 128 people, a machine-learning model using saliva bacteria distinguished people with early-onset colorectal cancer from healthy controls with an AUC of 0.780. This is an early association study, and the model needs independent testing before it could be considered for screening.
Research published in BMC microbiology ·
| Published | |
|---|---|
| Journal | BMC microbiology |
| Study design | Prospective cohort study with case-control comparison; saliva microbiome profiling and machine-learning model development. |
| Who took part | 65 people with early-onset colorectal cancer and 63 control individuals. |
| What was measured | How accurately saliva microbiome profiles distinguished early-onset colorectal cancer cases from controls. |
Why this is interesting
Colorectal cancer diagnosed at younger ages can be difficult to detect early, and a saliva-based test would be far less invasive than colonoscopy if it proved accurate enough. This paper asks whether patterns in oral bacteria could contribute to such a test.
What was already known Colorectal cancer screening can detect precancerous polyps and cancers before symptoms develop, but the best-established screening approaches do not currently use saliva. Early-onset colorectal cancer, usually defined as disease diagnosed before age 50, has been increasing in recent years. Researchers have linked changes in the microbiome, the community of microorganisms living in and on the body, to colorectal cancer, but it has remained uncertain whether bacteria found in saliva carry a sufficiently reliable signal for detection.
What this study adds This study provides an initial signal that saliva bacterial profiles differed between 65 people with early-onset colorectal cancer and 63 healthy controls. Its best model had moderate discriminatory performance within this small cohort. It does not establish that a saliva test can screen the general population, identify cancer before symptoms, or replace established diagnostic testing.

- Study group 65 cases, 63 controls
- Sample tested Saliva bacterial profiles
- Best model AUC 0.780
- Case recall 0.929 in this cohort
What the researchers tested
This was a prospective cohort study, as graded by the editor. The investigators collected saliva from 65 people with early-onset colorectal cancer, or EOCRC, and 63 control individuals without the disease. They then sequenced a section of bacterial genetic material, the V3-V4 region of the 16S ribosomal RNA gene, to identify and compare bacteria in the saliva samples.
The team used these microbial profiles to train several machine-learning models. Machine learning is a set of statistical methods that looks for combinations of features, here bacterial patterns, that best separate one group from another. The measured outcome was diagnostic discrimination: how well a model could distinguish the EOCRC group from the control group in this cohort.
This design can identify an association between a salivary microbiome pattern and having EOCRC. It cannot show that oral bacteria caused the cancer, that cancer caused the bacterial changes, or that changing oral bacteria would alter anyone’s cancer risk. It also cannot yet show that the test works in people attending routine screening, people with bowel symptoms, or people with other illnesses that may affect the oral microbiome.
The bacterial differences and model performance
The overall number and distribution of bacterial types within each individual sample, called alpha diversity, did not differ between the cancer and control groups. The groups did differ in beta diversity, a measure of how different the microbial communities are from one another across samples. That finding says the group-level composition differed, not that every person with EOCRC had a distinctive saliva profile.
At the genus level, the EOCRC group had higher relative abundance of Prevotella, Actinomyces, and Corynebacterium. Several genera were less abundant, including Fusobacterium, Haemophilus, Peptococcus, Eikenella, and the Eubacterium_yurii group. Relative abundance means the proportion of sequencing reads assigned to a bacterial group. It does not directly measure the total number of bacteria in the mouth.
Among the tested algorithms, a neural-network model performed best. It achieved an area under the receiver-operating-characteristic curve, or AUC, of 0.780. AUC describes how well a test separates two groups across possible positive-test thresholds: 0.5 is no better than chance, while 1.0 is perfect separation. An AUC of 0.780 is a potentially useful early signal, but it is well short of proving a test is ready for clinical screening.
The model’s recall was 0.929. Recall is the proportion of known EOCRC cases that the model classified as positive in this dataset. The paper does not report the corresponding specificity, which is the proportion of controls correctly classified as negative, nor does it give confidence intervals around the AUC or recall. Those missing details limit assessment of how many false-positive results such an approach might generate.
Why an internal model can look better than a future test
The central caution is the size and composition of the dataset. The model was developed from 128 participants, divided into known cancer cases and healthy controls. Algorithms can learn features that happen to separate a particular small dataset, including differences related to recruitment, diet, oral health, medication use, sample handling, or geography. This is often called overfitting.
An AUC of 0.780 in a small case-control cohort may fall substantially when independent researchers test the model in a new group of people. That external validation should include people representative of the intended use. A screening population includes many people without symptoms and with a much lower prevalence of cancer than a case-control study. A symptomatic population includes people with benign bowel conditions, inflammatory disease, infections, polyps, and cancers at different sites and stages. These are harder, and more clinically relevant, tests of specificity and usefulness.
The paper also compares people already known to have EOCRC with healthy controls. Screening aims to find disease before it is known, ideally at an earlier and treatable stage. The text does not report whether the model identified early-stage cancers especially well, whether it distinguished cancer from advanced precancerous polyps, or whether bacterial patterns remained predictive after accounting for factors that can shape saliva microbiota.
I would view this as biomarker discovery rather than a screening-test result. The biological observation is worth pursuing, particularly because saliva collection is easy and non-invasive. The current data do not support using salivary microbiome profiles to decide who does or does not need standard colorectal cancer assessment.
What would need to come next
A credible next study would pre-specify the bacterial signature and model before testing it in a large, independent set of participants. It should recruit people across the age range relevant to early-onset disease and include both screening participants and people undergoing diagnostic work-up for symptoms. Investigators would need to report sensitivity, specificity, false-positive and false-negative results, confidence intervals, and performance by cancer stage.
It would also be useful to know whether a saliva signature adds information beyond established clinical factors and whether it can identify advanced polyps as well as cancers. For a test intended for screening, researchers must show that its performance remains acceptable where cancer is uncommon. A model that detects most known cases but sends many unaffected people to invasive testing may not be useful in practice.
This study moves the question from speculation to a testable signal. It does not yet answer whether the signal is stable enough, specific enough, or accurate enough to change screening care. Those distinctions matter when a test is proposed for people who may feel well and are being asked to trust a result.
The numbers
- 65 EOCRC patients and 63 control individualsParticipantsThe cohort used to profile saliva bacteria and develop the models.
- AUC 0.780Best model discriminationThe neural-network model’s ability to separate cases from controls in this cohort.
- 0.929Best model recallThe proportion of known EOCRC cases identified by that model in this dataset.
What to take from this
- Saliva bacterial communities differed between the EOCRC and control groups, but this is an association, not evidence that the bacteria cause cancer.
- The best model reached an AUC of 0.780 within 128 participants, an encouraging early result that requires external validation.
- The study did not establish how the model would perform in real-world screening, among symptomatic patients, or against benign bowel conditions.
- No saliva microbiome test is established by this paper as a replacement for current colorectal cancer screening or diagnostic assessment.
What this study cannot tell us
The cohort was small, with 65 cancer cases and 63 controls, and the model appears to have been developed and assessed within that same cohort. The abstract reports no external validation, confidence intervals, specificity, cancer-stage analysis, or comparison with people who have symptoms or non-cancer colorectal conditions. Case-control comparisons between known cancer patients and healthy controls often make diagnostic separation appear stronger than it will be in a real screening setting. The trial registration was retrospective, meaning it was registered after the study had begun.
Worth asking your oncology team
These are questions this study raises, not recommendations. Your team knows your case; this article does not.
- Does my age, family history, symptoms, or prior testing affect which established colorectal cancer screening or diagnostic approach is appropriate for me?
- Are there any validated blood-, stool-, or tissue-based biomarkers relevant to my particular cancer care, as distinct from this early saliva-based research?
- If saliva microbiome testing is offered commercially, what evidence shows that it has been validated in people like me and that its results would change clinical decisions?
The source
Zhen J, Dong M, Li Y, Cao B, Liao F, Lin D, Zhang J, Liu C, Zheng X, Dong W.. Human salivary microbiome as a potential non-invasive biomarker for early-onset colorectal cancer screening: a prospective study.. BMC microbiology. 2026
This article summarises published research for general information. It is not medical advice, and it is not a substitute for a conversation with your own oncology team, who know your case. Do not start, stop, or change any treatment or supplement on the basis of what you read here.
