Search
Browse
Statistics
Feeds

metadeconfoundR: covariate analysis of high-dimensional cross-sectional omics data

[thumbnail of Accepted Manuscript]
Preview
PDF (Accepted Manuscript) - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
872kB
[thumbnail of Supplementary Data]
Preview
PDF (Supplementary Data) - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
3MB

Item Type:Article
Title:metadeconfoundR: covariate analysis of high-dimensional cross-sectional omics data
Creators: Birkner, Till ORCID logoORCID: https://orcid.org/0000-0003-2656-2821, Chen, Chia-Yu ORCID logoORCID: https://orcid.org/0000-0003-1765-7132, Essex, Morgan ORCID logoORCID: https://orcid.org/0000-0001-8758-7497, Dahm, Kilian ORCID logoORCID: https://orcid.org/0000-0002-6819-2622, Löber, Ulrike ORCID logoORCID: https://orcid.org/0000-0001-7468-9531, Ulas, Thomas ORCID logoORCID: https://orcid.org/0000-0002-9785-4197, Jarquín-Díaz, Víctor Hugo ORCID logoORCID: https://orcid.org/0000-0003-3758-1091 and Forslund-Startceva, Sofia Kirke ORCID logoORCID: https://orcid.org/0000-0003-4285-6993
Abstract:MOTIVATION: Identifying disease biomarkers from large molecular datasets is complicated by correlated and confounded signals like comorbidities and treatment regimens, batch effects, and cohort biases. These effects bias statistical inference and clinical conclusions. Robust methodologies are fundamental for reliable biomarker discovery. RESULTS: metadeconfoundR is an R package for conservative biomarker discovery in (multi-)omics case-control datasets. It has a scalable two-step confounder-aware statistical framework for retaining only associations with independent support. It identifies covariate-naive univariate associations between omics features and metadata, then re-evaluates these associations using parallel post-hoc nested linear model testing to account for potential confounders. Confounded associations are flagged if they fully reduce to at least one other variable. metadeconfoundR supports parallel computation for large-scale datasets, offers visualization and tools for interpreting results and secondary analyses. We benchmark metadeconfoundR against state-of-the-art methods for identifying biomarkers using simulated ground truth derived from microbiome data, and demonstrate its ability to disentangle confounding effects while preserving statistical power, offering particular advantage when multiple covariates are present. metadeconfoundR functions for any -omics data type with continuous or categorical metadata/covariates. AVAILABILITY: metadeconfoundR is available on CRAN (https://cran.r-project.org/web/packages/metadeconfoundR/) and GitHub (https://github.com/TillBirkner/metadeconfoundR).
Keywords:Microbiome, Metagenomics, Multi-Omics, Biomarkers, Software, Confounding, R-Package, Model-Selection, Case-Control Studies
Source:Bioinformatics Advances
ISSN:2635-0041
Publisher:Oxford University Press
Page Range:vbag242
Date:19 August 2026
Official Publication:https://doi.org/10.1093/bioadv/vbag242
Related to:

Repository Staff Only: item control page

Downloads

Downloads per month over past year

Open Access
MDC Library