Search
Browse
Statistics
Feeds

quantmsdiann: a scalable SDRF-driven DIA-NN workflow for reanalysis of single-cell, spatial, and bulk proteomics datasets

[thumbnail of Preprint]
Preview
PDF (Preprint) - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
403kB
[thumbnail of Supplementary Files] MS Word (Supplementary Files)
1MB
Item Type:Preprint
Title:quantmsdiann: a scalable SDRF-driven DIA-NN workflow for reanalysis of single-cell, spatial, and bulk proteomics datasets
Creators: Yue, Qi-Xuan ORCID logoORCID: https://orcid.org/0009-0004-8120-9339, Shen, Yufei, Dai, Chengxin ORCID logoORCID: https://orcid.org/0000-0001-6943-5211, Larrea-Sebal, Asier ORCID logoORCID: https://orcid.org/0000-0001-9107-4299, Webel, Henry ORCID logoORCID: https://orcid.org/0000-0001-8833-7617, Nimo, Jose ORCID logoORCID: https://orcid.org/0000-0002-1565-7799, Coscia, Fabian ORCID logoORCID: https://orcid.org/0000-0002-2244-5081, Kok, Orhun, Slavov, Nikolai ORCID logoORCID: https://orcid.org/0000-0003-2035-1820, Vizcaíno, Juan Antonio ORCID logoORCID: https://orcid.org/0000-0002-3905-4335, Sachsenberg, Timo ORCID logoORCID: https://orcid.org/0000-0002-2833-6070, Bai, Mingze ORCID logoORCID: https://orcid.org/0000-0002-9782-2056, Demichev, Vadim ORCID logoORCID: https://orcid.org/0000-0002-2424-9412 and Perez-Riverol, Yasset ORCID logoORCID: https://orcid.org/0000-0001-6579-6941
Abstract:Public proteomics archives now hold thousands of data-independent acquisition (DIA) datasets, but reusing them is difficult: each was processed with a different software configuration, and most lack standardized metadata. Here, we present quantmsdiann, an open-source Nextflow/nf-core workflow that runs DIA-NN in parallel across cloud and high-performance computing (HPC) infrastructure, guided by the experimental design declared in SDRF format. The workflow provides pinned container profiles and builds recipes for each supported DIA-NN version under BioContainers, resolving dependencies automatically, and reads all major vendor formats and the HUPO-PSI mzML standard. It exports harmonized quantification tables as MSstats input, in the Quantitative Proteomics eXchange (QPX) format, and a pmultiqc quality-control report. The parallel design reanalyzes a 2,300-run single-cell dataset in 2.2 hours on 300 HPC nodes. We performed multiple experiments and benchmarks of quantmsdiann on single-cell datasets; ProteoBench DIA-NN single-machine submissions or public datasets in ProteomeXchange. The benchmark against ProteoBench single-machine DIA-NN modules demonstrated no differences between DIA-NN single-machine runs and parallelization in quantmsdiann; while upgrading DIA-NN from 1.8.1 to a current release increased protein-group identifications by up to 17% in single-cell datasets. Remarkably, reanalysis of public DIA datasets with quantmsdiann and the latest version of DIA-NN always exceeds the originally deposited counts, recovering up to 59% more protein groups, with the largest gains on deposits processed with older or non-DIA-NN engines and smaller gains where a recent DIA-NN release was already used. quantmsdiann is a step toward scalable, reproducible reanalysis of the growing DIA data archive and a foundation for large-scale DIA meta-analysis and atlas building. It is available at https://github.com/bigbio/quantmsdiann.
Keywords:DIA-NN, Single-Cell Proteomics, Spatial Proteomics, Phosphoproteomics, plexDIA, Dataindependent Acquisition, Nextflow, nf-Core, SDRF, Reproducibility, QPX, pmultiqc
Source:Research Square
Publisher:Research Square
Article Number:rs-10319687/v1
Date:14 July 2026
Official Publication:https://doi.org/10.21203/rs.3.rs-10319687/v1
Related to:

Repository Staff Only: item control page

Downloads

Downloads per month over past year

Open Access
MDC Library