Alignment-free sequence comparison for biologically realistic sequences of moderate length
Loading...
Date
Authors
Burden, Conrad J
Jing, Junmei
Wilson, Susan R
Journal Title
Journal ISSN
Volume Title
Publisher
Walter de Gruyter
Abstract
The D2 statistic, defined as the number of matches of words of some pre-specified length k, is a computationally fast alignment-free measure of biological sequence similarity. However there
is some debate about its suitability for this purpose as the variability in D2 may be dominated by the terms that reflect the noise in each of the single sequences only. We examine the extent of the problem and the effectiveness of overcoming it by using two mean-centred variants of this statistic,
D2* and D2c. We conclude that all three statistics are potentially useful measures of sequence similarity, for which reasonably accurate p-values can be estimated under a null hypothesis of sequences composed of identically and independently distributed letters. We show that D2 and D2c, and to a somewhat lesser extent D2*, perform well in tests to classify moderate length query
sequences as putative cis-regulatory modules.
Description
Keywords
Citation
Collections
Source
Statistical Applications in Genetics and Molecular Biology 11.1 (2012):1-28
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
Downloads
File
Description