We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:


Current browse context:


Change to browse by:

References & Citations


(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Quantitative Biology > Quantitative Methods

Title: Multiscale methods for signal selection in single-cell data

Abstract: Analysis of single-cell transcriptomics often relies on clustering cells and then performing differential gene expression (DGE) to identify genes that vary between these clusters. These discrete analyses successfully determine cell types and markers; however, continuous variation within and between cell types may not be detected. We propose three topologically motivated mathematical methods for unsupervised feature selection that consider discrete and continuous transcriptional patterns on an equal footing across multiple scales simultaneously. Eigenscores ($\text{eig}_i$) rank signals or genes based on their correspondence to low-frequency intrinsic patterning in the data using the spectral decomposition of the Laplacian graph. The multiscale Laplacian score (MLS) is an unsupervised method for locating relevant scales in data and selecting the genes that are coherently expressed at these respective scales. The persistent Rayleigh quotient (PRQ) takes data equipped with a filtration, allowing the separation of genes with different roles in a bifurcation process (e.g., pseudo-time). We demonstrate the utility of these techniques by applying them to published single-cell transcriptomics data sets. The methods validate previously identified genes and detect additional biologically meaningful genes with coherent expression patterns. By studying the interaction between gene signals and the geometry of the underlying space, the three methods give multidimensional rankings of the genes and visualisation of relationships between them.
Comments: 32 pages, 15 figures, 1 table. Revised and published in Entropy, special issue Applications of Topological Data Analysis in the Life Sciences
Subjects: Quantitative Methods (q-bio.QM); Social and Information Networks (cs.SI); Algebraic Topology (math.AT); Spectral Theory (math.SP); Machine Learning (stat.ML)
Journal reference: Entropy 2022, 24(8), 1116
DOI: 10.3390/e24081116
Cite as: arXiv:2206.07760 [q-bio.QM]
  (or arXiv:2206.07760v2 [q-bio.QM] for this version)

Submission history

From: Otto Sumray [view email]
[v1] Wed, 15 Jun 2022 18:42:26 GMT (7745kb,D)
[v2] Thu, 6 Oct 2022 11:29:22 GMT (8029kb,D)

Link back to: arXiv, form interface, contact.