openHSU logo
Log In(current)
  1. Home
  2. Helmut-Schmidt-University / University of the Federal Armed Forces Hamburg
  3. Publications
  4. 3 - Publication references (without full text)
  5. HarmonizR: blocking and singular feature data adjustment improve runtime efficiency and data preservation

HarmonizR: blocking and singular feature data adjustment improve runtime efficiency and data preservation

Publication date
2025-02-11
Document type
Forschungsartikel
Author
Schlumbohm, Simon  
Neumann, Julia E.
Neumann, Philipp  
Organisational unit
High Performance Computing  
DOI
10.1186/s12859-025-06073-9
URI
https://openhsu.ub.hsu-hh.de/handle/10.24405/22992
Publisher
BioMed Central
Series or journal
BMC Bioinformatics
ISSN
1471-2105
Periodical volume
26
Periodical issue
1
Article ID
47
Peer-reviewed
✅
Part of the university bibliography
✅
Funding(s)
Publikationsfonds der HSU/UniBw H  
Additional Information
Language
English
Keyword
Batch effects
Proteomics
Computational efficiency
Big data
Dataset integration
Abstract
Background:
Data adjustment is an essential tool for increasing statistical power during analysis, for example in case of complex multi-experiment data from (single-cell) RNA, proteomics and other omics data. Despite its benefits, data integration introduces internal biases—so-called batch effects. Due to the inherent presence of missing values by such methods and their additional introduction by means of data integration, renowned algorithms such as ComBat and limma are unable to perform batch effect adjustment. Recently, the HarmonizR framework was presented for these cases, which is a tool for missing value tolerant data adjustment.
Results:
In this contribution, we provide significant improvements to the HarmonizR approach. A novel blocking strategy is introduced to severely reduce runtime, while still supporting parallel architectures. Additionally, a “unique removal” strategy has been integrated into HarmonizR to maintain even more features for adjustment in datasets, showing a feature rescue of up to 103.9% for our tested datasets. In this work, we show (1) severely improved runtime for both small and large, real datasets and (2) the ability retain more features from the integrated dataset during adjustment, showing a feature rescue of up to 103.9% for our tested datasets.
Conclusion:
The proposed improvements tackle the previous shortcomings of the published HarmonizR version. Since HarmonizR was mainly developed for dataset integration on rare tumor entities, it did not include runtime improvements beyond parallelization, which has been addressed in this update. An additionally welcome update regarding improved feature rescue furthermore enhances the algorithms ability to quickly and robustly perform batch effect reduction.
Description
This article is licensed under a Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/).
Version
Published version
Access right on openHSU
Metadata only access

  • Privacy policy
  • Send Feedback
  • Imprint