The textbook version of the genome is tidy: DNA is DNA, RNA is RNA, and the two do not mix. The real version is messier. Roughly one million ribonucleotides sit embedded in every human genome, and we spent years working out where they are and what they do.
That work is now published in Cell. I led the bioinformatics analyses; Georgia Tech wrote it up here.
How RNA ends up in DNA
Ribonucleotides get incorporated because a cell's pool of RNA building blocks vastly outnumbers its pool of DNA building blocks. Polymerases discriminate well, but not perfectly, and across a whole genome even a tiny error rate produces a large absolute number.
The clean-up crew is RNase H2, an enzyme whose job is to recognise a ribonucleotide sitting in a DNA context and start its removal. So the standard story has been: mistake, then repair.
Where they land is not random
Mapping them at base resolution across human cell types, we found the placement is anything but random. Embedded ribonucleotides cluster at:
- transcription start sites
- CpG islands
- R-loops and G-quadruplexes
- telomeres
And their density scales with how actively a gene is transcribed. A mark that tracks transcriptional state and differs by cell type is doing something more interesting than waiting to be repaired.
The twist: they change DNA's shape
When RNase H2 nicks an embedded ribonucleotide, DNA supercoiling shifts. Remove RNase H2 and negative supercoiling builds up at promoters, with Top1 stepping in to compensate.
Embedded ribonucleotides are not only damage awaiting repair. They are epigenetic modulators of transcription-associated DNA topology.
That is the part I did not expect going in. The chemical composition of the strand feeds directly into its physical organisation, and physical organisation feeds back into transcription.
Why it matters beyond the mechanism
Mutations in RNase H2 cause Aicardi-Goutières syndrome, a rare genetic disorder driven by runaway immune activation. The same enzymes are tied to genome instability and tumorigenesis.
A base-resolution map of where these ribonucleotides sit, plus evidence that they shape DNA topology, opens a Pandora's box of questions about how and where ribonucleotide processing feeds genome instability and inflammation. I find that a good problem to have.
Not a solo effort
None of this would exist without my co-authors: Francesca Storici, my advisor, whose command of the biology shaped every question we asked; Natasa Jonoska, for the statistical framework behind our claims of non-randomness; and Yeunsoo Lee, for the experiments that carried the supercoiling story from hypothesis to evidence.