Omics data in relative values are almost subcompositionally coherent
Marina Martínez-Álvaro, Michael Greenacre, A. Blasco
Introduction Omics data are compositional and often expressed as relative abundances after total sum scaling normalization. An important statistical issue with compositional data is the lack of subcompositional coherence, meaning that relative abundances change when data are re-normalized after removing or adding features. While this problem is well documented for small compositions, it has not been investigated in large Omics datasets, which typically contain hundreds or thousands of features and where subcompositions are ubiquitous. Subcompositions arise, for example, when using different reference datasets, sequencing depths or when filtering low-abundant features from the database. In such cases, the most abundant features are preferentially retained, whereas variation between original or full compositions and subcompositions is mainly driven by less abundant features. The standard solution to this problem is the use of logratio transformations, but these complicate interpretations and require handling zeros, which are frequent in Omics data and whose imputation introduces spurious variability. Methods Here, we evaluated subcompositional coherence in five representative Omics datasets: fecal 16S metagenomics, rumen metagenomics (taxonomic and functional levels), liver transcriptomics, and plasma metabolomics, considering both unsupervised and supervised learning contexts. We generated 100 random subcompositions comprising one-third of the original features under an abundance-weighted subcomposition scheme and compared their statistical outputs with those from the full composition. Results and discussion Raw Omics data showed near-perfect coherence: relative abundances, pairwise correlations and sample distances all exhibited very high (scaled) concordances (≥0.98–0.99). Outputs from commonly used supervised models (linear regression, PLS, random forest, and linear mixed models with a Gaussian kernel) were also highly subcompositionally coherent. We conclude that large Omics datasets expressed as relative abundances are almost subcompositionally coherent when considering a weighted subcomposition scheme, thereby challenging one of the criticisms of using relative data in the Omics field over logratio transformations.