A seminar by Associate Professor Heejung Shim from University of Melbourne
Title: Modelling spatial transcriptomics: from flexible cell-type deconvolution to multi-scale spatial factor analysis
Abstract: Spatial transcriptomics enables the study of gene expression within its spatial context, but introduces key statistical challenges, including mixed cellular composition and complex spatial structure. In this talk, I present two complementary modelling approaches, with a particular focus on the statistical ideas underlying flexible use of prior information and multi-scale modelling of spatial structure.
First, I introduce FlexiDeconv, a cell-type deconvolution method based on a Bayesian topic model. A key feature of this method is its flexible use of reference information, allowing the model to balance prior information from scRNA-seq with signals from observed spatial data, and to adapt when the reference is incomplete or mismatched, a common challenge in practice. I will also briefly discuss ongoing work extending this flexible deconvolution framework to isoform-level spatial transcriptomics, where the increased dimensionality and uncertainty associated with isoform-level measurements introduce additional modelling challenges.
I then focus in more detail on WaveFactor, a wavelet-based Bayesian sparse factor model for identifying spatial gene expression patterns. Wavelet and related multi-scale approaches have been used in signal processing and have also been successfully adapted to high-throughput sequencing data to model spatial structure across genomic locations (Shim and Stephens, 2015; Shim et al., 2024). The key idea is to represent a spatially structured signal at multiple resolutions, rather than choosing a single spatial scale. Such a representation can capture both fine- and broad-scale patterns while providing a sparse representation of spatial structure. I will introduce the intuition behind wavelet-based multi-scale modelling and discuss how these ideas are incorporated into WaveFactor. WaveFactor can additionally incorporate gene-set information to guide factor inference, while allowing for uncertainty and potential errors in these annotations.
Together, these methods illustrate how flexible modelling of prior information and multi-scale modelling of spatial structure can improve our ability to extract biologically meaningful signals from spatial transcriptomics data.
Co-authors: This presentation is based on joint work with Yichen Jiang and Yulin Wu.
For further information, please contact RSFAS Seminars.
All information collected by the University is governed by the ANU Privacy Policy.

