Skip to main content
The same Analysis Module with the same set of parameters can produce widely different results depending on the pre-processing settings. Pre-processing parameters can affect diversity values, distance matrices, clustering patterns, feature importance, and statistical comparisons.

What to look for when setting pre-processing parameters

Several properties of the Data Table determine how strongly pre-processing affects your results. How much these properties matter depends mainly on the type of sequencing data (long-read or short-read 16S, or WGS), the profile type (taxonomic or functional), and the complexity and variability of the sampled environment (e.g., soil samples tend to be more complex, whereas vaginal samples tend to be less complex):
  • Sparsity: a feature is sparse across the Data Table when it is not detected in most samples of the cohort. Sparse features are typically taxa or genes with extremely low prevalence in the Data Table.
  • Sequencing depth: how many reads each sample has, and how much this differs between samples in the cohort. Uneven sequencing depth introduces technical bias into comparisons and can be mistaken for a real biological difference in the variable of interest.

Sparsity and Sequencing depth variability depending on the Cosmos-Hub Profiling Workflow

Pre-processing impact on Analysis Module outcomes

References

  • Zhou R et al. Data pre-processing for analyzing microbiome data. Brief Bioinform / NIH review. 2023;24(6)
  • Busato S et al. Compositionality, sparsity, spurious heterogeneity, and other pitfalls in microbiome machine learning. Patterns. 2022;3(12)