Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics
Dec 27, 2022·

·
0 min read
A. subramanian
*†
,M. alperovich
*
Yiming Yang
Dr. Bo Li
Abstract
Quality control (QC) of cells, a critical first step in single-cell RNA sequencing data analysis, has largely relied on arbitrarily fixed data-agnostic thresholds applied to QC metrics such as gene complexity and fraction of reads mapping to mitochondrial genes. The few existing data-driven approaches perform QC at the level of samples or studies without accounting for biological variation. We first demonstrate that QC metrics vary with both tissue and cell types across technologies, study conditions, and species. We then propose data-driven QC (ddqc), an unsupervised adaptive QC framework to perform flexible and data-driven QC at the level of cell types while retaining critical biological insights and improved power for downstream analysis. ddqc applies an adaptive threshold based on the median absolute deviation on four QC metrics (gene and UMI complexity, fraction of reads mapping to mitochondrial and ribosomal genes). ddqc retains over a third more cells when compared to conventional data-agnostic QC filters. Finally, we show that ddqc recovers biologically meaningful trends in gradation of gene complexity among cell types that can help answer questions of biological interest such as which cell types express the least and most number of transcripts overall, and ribosomal transcripts specifically. ddqc retains cell types such as metabolically active parenchymal cells and specialized cells such as neutrophils which are often lost by conventional QC. Taken together, our work proposes a revised paradigm to quality filtering best practices—iterative QC, providing a data-driven QC framework compatible with observed biological diversity.
Type
Publication
Genome Biology
Authors
Authors

Authors
Bioinformatics Software Engineer
Yiming Yang is a bioinformatics software engineer in Li Lab.

Authors
Principal Scientist II
Dr. Bo Li is a Principal Scientist II at AI for Biology and Translation (AIBT), Genentech, Inc. His research focuses on three major topics: Science, Technology and Computational Methods. For Science, his team works on lung cancer and especially small cell lung cancer. For Technology, his team evaluates and adopts cutting-edge high-throughput data generation technologies, such as Cellanome and SBX sequencing. For Computational Methods, his team develops novel computational and deep learning tools for enabling insight discovery from high-throughput multi-modal data.
Before joining in Genentech, he was an Assistant Professor of Medicine at Harvard Medical School and the director of Bioinformatics and Computational Biology at Center for Immunology and Inflammatory Diseases, Massachusetts General Hospital.
He received his Ph.D. in computer science from UW-Madison and completed two postdoctoral trainings with Dr. Lior Pachter at UC Berkeley and Dr. Aviv Regev at Broad Institute.
He is best known for developing RSEM, an impactful RNA-seq transcript quantification software. RSEM is cited 22,602 times (Google Scholar) and adopted by several big consortia such as TCGA, ENCODE, GTEx and TOPMed.