Incorporation of Phylogenetic and Comparative Genomic Information in Deep Neural Network for RNA Secondary Structure Prediction
Abstract
In RNA secondary structure prediction, deep neural network (DNN) models
have become mainstream and achieved State-Of-The-Art performance.
Biological prior information such as hard constraints on RNA structure,
including Watson-Crick base pairing and minimum loop length, have been
included in these neural network architectures to further promote prediction
accuracy. However, comparative genomic information from species evolution
has never been incorporated into such models. In conserving RNA structure
across evolutionary time, the proportion of compensatory double substitutions
& compatible single substitutions in base-pairing region is expected to be
higher than that in unpaired regions (e.g. loops or bulges in secondary
structure). We hypothesize that incorporating this evolutionary signal can
better differentiate base-pairs and loops, to improve the accuracy of secondary
structure prediction. In this thesis, we propose a new approach of extracting
and encoding these characteristic substitution patterns in a DNN architecture
to improve RNA secondary structure prediction accuracy. I show that the
proposed approach substantially improves prediction performance on both
synthetic and real-world RNA sequences.