Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Novel techniques for measuring the effect of neighbouring bases on mutation and their applications

dc.contributor.authorZhu, Yicheng
dc.date.accessioned2020-02-17T12:29:34Z
dc.date.available2020-02-17T12:29:34Z
dc.date.issued2020
dc.description.abstractUnderstanding factors influencing mutations can improve detection of novel mutations, the diagnostic signatures of disease-causing mutagens, and facilitate the development of more accurate models of genetic divergence. Hypermutability of CpG demonstrates the existence of mutation motifs, sequences of flanking bases that influence point mutation processes. These motifs can also be indicative of specific underlying mutation mechanisms. I developed novel log-linear models for identifying mutation motifs that allow further comparisons of these mutation motifs, and of the complete mutation spectra between samples. Mutation motifs are visualised using a sequence logo type method. In this thesis, I applied the methods to examine each of the possible 12 point mutations in about 13.6 million human germline mutations (inferred from single-nucleotide polymorphisms recorded in the Ensembl database) and about 181,000 melanoma mutations from the COSMIC database. My method recovered the well-known CpG effect, which a conventional motif detection method failed to do. I established that all point mutations have significant and distinct mutation motifs. While the major effects of flanking bases lie within 2 bp of the mutated position, I refute previous reports that the effect magnitude decays monotonically with distance. Comparison between autosomes and X-chromosomes supported a reduced contribution from methylation-induced C to T mutation on the X-chromosome, consistent with a previous prediction. In addition, analyses of malignant melanoma confirmed reported characteristic features of this cancer, such as strand asymmetry of mutation processes. Further, I found that neighbouring influences in malignant melanoma differ significantly from those affecting germline mutations. Interestingly, for C to T mutation, the CpG effect was no longer evident, and was largely substituted by different neighbouring mechanisms. Moreover, the observed neighbouring influence is able to reflect the chemical influences of mutagenic processes after exposure to ultraviolet light. Based on this observation, I hypothesised that information regarding the mechanistic origin of point mutations is present in surrounding DNA sequences, and sequence neighbourhood can be used to identify the mechanistic origin of particular mutations. Machine learning classifiers were developed to assess the above hypothesis and discriminate between N-ethyl-N-nitrosourea (ENU)-induced and spontaneous point mutations in the mouse germline. ENU is a synthetic chemical employed in mutagenesis studies, introducing novel point mutations to genomes. My classification results reveal that a combination of k-mer size and representation of second-order interactions among nucleotides was able to improve classification performance compared to the naive classifier approach. In conclusion, this work demonstrates that neighbouring bases have a profound effect on the occurrence of mutations. The statistical methods reported in this research can be used to examine the role of flanking sequence on mutation processes from polymorphism data, which further enable identification of differences in the operation of mechanisms of mutation between genomic regions, cell types or species. In addition, the machine learning classification results have important implications for modelling context-dependent effects on sequence evolution.
dc.identifier.otherb71497390
dc.identifier.urihttp://hdl.handle.net/1885/201728
dc.language.isoen_AU
dc.titleNovel techniques for measuring the effect of neighbouring bases on mutation and their applications
dc.typeThesis (PhD)
local.contributor.supervisorHuttley, Gavin
local.identifier.doi10.25911/5f58af958e859
local.identifier.proquestYes
local.mintdoimint
local.thesisANUonly.author62900a07-43b7-4ff6-836e-a93e4ecbf274
local.thesisANUonly.key1a72f5eb-18a7-902c-f605-7fd7f3431836
local.thesisANUonly.title000000011626_TC_1

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Y_Zhu_PhD_Thesis_of_ANU.pdf
Size:
4.32 MB
Format:
Adobe Portable Document Format
Description:
Thesis Material