06
2026
The π-HuB Secretariat is glad to share a seminal methodological review jointly led by Prof. Xin Zhou and Prof. Jing Yang, focusing on statistical power and sample-size estimation for human microbiome research. Published in Med on July 17, 2026, the review is titled Power and sample-size estimation in human microbiome research. Given the close interplay between microbiome, human proteome, host physiology and disease mechanisms, this comprehensive summary addresses long-standing statistical bottlenecks in microbiome cohort studies and delivers actionable guidance for large-scale omics research.
Human microbiome research is widely applied to explore complex diseases including diabetes, inflammatory bowel disease and cancers, mostly relying on high-throughput metagenomic sequencing to compare microbial communities. However, microbiome data feature compositionality, sparsity and severe zero inflation, which greatly complicate statistical modeling and raise sample-size requirements.
Substantial inter-individual and intra-individual microbial variations further exacerbate the issue. Many existing studies suffer from inadequate statistical power and poor reproducibility. Classic cases include inconsistent conclusions about the Firmicutes/Bacteroidetes ratio as an obesity biomarker, and contradictory findings on microbial markers for diabetes. While power analysis is well-established in other omics fields, targeted, user-friendly guidelines for microbiome research have long been lacking, especially for clinical researchers and early-career investigators. This review fills this critical gap.
This work systematically sorts out four core components of power analysis: significance level (α), statistical power (1-β), effect size and sample size. A notable feature of microbiome research is that the definition of effect size depends heavily on analytical models.
The review thoroughly analyzes mainstream alpha diversity and beta diversity metrics, along with their applicable scenarios and corresponding sample-size requirements:
Alpha diversity metrics: Chao1 is sensitive to rare taxa but has high variability, requiring deeper sequencing and larger sample sizes. The Shannon-Wiener index balances richness and evenness, and 15–20 samples per group are sufficient for most general studies. The Simpson dominance index focuses on dominant taxa with low variance, and merely 10–15 samples per group can achieve adequate power. Fisher’s alpha fits microbial communities following log-series distribution and is less affected by sequencing depth.
Beta diversity metrics: Jaccard distance only reflects species presence or absence, being vulnerable to data sparsity and thus demanding larger cohorts. Bray-Curtis dissimilarity is based on species abundance and sensitive to shifts of dominant microbes. Weighted UniFrac integrates abundance and phylogenetic information with good stability and low sample demands, while unweighted UniFrac targets rare taxa and phylogenetic differences and requires more samples.
As the common analytical tool for beta diversity comparison, PERMANOVA does not follow conventional distribution assumptions. Simulation-based methods are recommended for its power calculation and sample-size estimation.
For microbial biomarker mining, the review compares the strengths and limitations of various statistical approaches. Parametric methods including t-test, ANOVA, ANCOM, Dirichlet-multinomial model (DMM) and Negative Binomial (NB) model, as well as zero-inflated models are elaborated. Non-parametric tests like Mann-Whitney U test and Kruskal-Wallis test are recommended for data that violate normality assumptions.
Multiple hypothesis testing is a major factor restricting statistical power. The review suggests reasonable feature filtering and separate discovery-validation cohorts as effective solutions. Beyond basic association analysis, the paper also discusses correlation analysis, mediation analysis and Mendelian randomization (MR) for causal exploration. Correlation studies often face small effect sizes and need large samples; mediation analysis requires 2 to 3 times more samples than direct association research due to weaker indirect effects; MR is limited by the low heritability of microbiome traits and also relies on large cohorts to obtain reliable results.
Longitudinal designs use each participant as their own control, effectively reducing inter-individual interference and improving the capacity to capture subtle microbial changes over time. This review summarizes multiple strategies to boost power in longitudinal research: adopting matched design and proper sampling frequency to cut down variability; applying mixed-effects models, generalized estimating equations (GEE) and other specialized tools to adapt to repeated measurement data; increasing sequencing depth and unifying experimental protocols to reduce technical noise and batch effects. These strategies provide solid methodological support for dynamic omics tracking.
This methodological review delivers valuable insights and practical references. Rigorous cohort design, rational sample size planning and standardized statistical analysis are fundamental to both microbiome research and large-scale human proteome studies carried out under π-HuB. The complete set of power analysis methods, diversified statistical model selection strategies and simulation-based sample size calculation approaches summarized in this work can be fully referenced for π-HuB’s cohort construction, multi-omics data processing and result interpretation.
In addition, the research paradigms covering cross-sectional comparison, longitudinal dynamic monitoring and multi-layer causal inference are highly compatible with π-HuB’s research layout. The solutions proposed for addressing data compositionality, sparsity, zero inflation and multiple testing issues can help standardize analytical workflows, improve data reliability and reproducibility, and support high-quality output of follow-up research. This work also sets a valuable reference for expanding the combined research of proteomics and microbiology within the framework of π-HuB.
Reference
Zhou Q, Lu Y, Wang L, et al. Power and sample-size estimation in human microbiome research. Med. Published online June 16, 2026. doi:10.1016/j.medj.2026.101174
We use cookies to personalize and enhance your browsing experience on our website. By clicking "Accept All", you agree to use cookies. You can read our Cookie Policy for more information.