Welcome!
We are a computational biology group in the School of Life Sciences and Technology at Tongji University. Our goal is to understand the diverse mechanisms by which noncoding sequences regulate gene expression at the transcriptome level. We specialize in developing statistical and integrative methods for the analysis of multi-omics high-throughput data, and we apply them to study the role of repetitive sequences — particularly transposable elements (TEs) — in the control of gene expression. In parallel, we investigate transcriptional and post-transcriptional regulation, and the mechanisms governing the biogenesis and biological functions of noncoding RNAs. Our work advances the understanding of noncoding sequences and their relevance to human health and disease.
Research Directions
Developing AI-Driven Computational Methods
High-throughput sequencing has made data generation routine, but extracting biological meaning from it has become the bottleneck. We develop the computational and statistical methods that address it — algorithms that convert raw reads into interpretable biological insight. Our work covers three classes of problem: learning regulatory grammar directly from DNA and RNA sequence with deep neural networks; inferring the rates and coordination of RNA processing steps from nascent transcription; and disentangling biological variation from technical noise in single-cell and long-read measurements. Methodologically we combine machine learning with probabilistic modeling and classical statistics, and we favor models that stay interpretable, because a model is useful to us only when it points to a mechanism we can test. These methods serve the lab's other projects and are released as open software for the wider community.
Decipher TE-Mediated Gene Expression
A central focus of our lab is to understand how transposable elements (TEs) shape gene expression. Once regarded largely as genomic parasites, TEs are now recognized as a major source of regulatory innovation, contributing to both genome evolution and the origins of disease. Many TEs carry functional transcriptional regulatory elements, which adds considerable complexity to their roles in the cell. Our goal is twofold: to uncover the mechanisms by which TEs give rise to new transcripts, and to determine how those transcripts contribute to cellular function. Beyond acting as alternative promoters, TEs can serve as cis-regulatory elements that recruit transcription factors and modulate gene activity. We therefore investigate how TEs operate in specific cellular contexts, taking temporal and spatial factors into account and relating them to epigenetic modifications. Addressing these questions requires a multidisciplinary approach that combines experimental and computational methods, and it is through this integration that we aim to clarify how TEs shape gene regulation, genome function and evolution.
Unraveling Kinetics of RNA Processing
Our lab investigates the kinetics of RNA processing, a process that profoundly shapes the diversity of gene expression. Phenotypic variation between individuals often arises not from differences in the genetic code itself, but from subtle changes in how that code is regulated. Recent work has resolved individual molecular steps that control mRNA abundance, revealing that transcriptome diversity emerges from the distinct ways in which RNA molecules are processed. Conventional expression studies typically measure steady-state RNA levels, yet living systems are inherently dynamic. This raises a central question in functional genomics: how do numerous dynamic molecular processes combine to determine the cellular RNA landscape? We therefore set out to measure the rate and efficiency of pivotal milestones in the RNA life cycle, including transcriptional elongation, splicing, and cleavage and polyadenylation. To do so, we combine experimental techniques for capturing nascent RNA with high-throughput sequencing, then apply computational and mathematical modeling to read the rhythm of each processing step along a gene. These genome-wide measurements reveal principles that govern efficient RNA processing within cells.
Decoding Noncoding RNA Complexity
A substantial proportion of human transcripts lack the potential to encode proteins, and are therefore classified as noncoding RNAs. These molecules show expression patterns specific to particular diseases, and their involvement in transcriptional and post-transcriptional regulation makes them plausible drivers of disease. Our objective is to characterize the biogenesis and molecular functions of noncoding RNAs within the broader landscape of genetics and disease, with the ultimate aim of advancing precision medicine. To this end, we combine mathematical and statistical modeling with cutting-edge experimental approaches, which allows us to dissect the properties of long noncoding RNAs with distinctive structures, such as circular RNAs (circRNAs), as well as small noncoding RNAs including snRNAs, tsRNAs and rsRNAs. Through this work we seek to clarify the subtle regulatory roles of noncoding RNAs and their impact on human health and disease.
Available positions
We are looking for motivated researchers to join the Xiao-Ou Zhang Lab. Our work spans noncoding sequences, gene regulation and precision medicine for human disease. Using computational and data-driven approaches, we develop algorithms, statistical tools and bioinformatics pipelines that make sense of diverse sequencing data, from second- and third-generation sequencing to single-cell assays.
If you would like to be part of our research environment, send your CV and a short statement of research interests to zhangxiaoou@tongji.edu.cn. We welcome applications from postdocs, graduate students, research assistants and undergraduate interns.