Content deleted Content added
Help needed: Complexity measure |
m Disambiguating links to Complexity measure (link changed to Computer linguistics) using DisamAssist. |
||
Line 2:
When a [[nucleotide]] sequence is written as text using a four-letter alphabet, the repetitiveness of the text, that is, the repetition of its [[N-gram]]s (words), can be calculated and serves as a measure of sequence complexity. Thus, the more complex a [[DNA sequence]], the richer its [[oligonucleotide]] vocabulary, whereas repetitious sequences have relatively lower complexities. Subsequent work improved the original algorithm described in [[Edward Trifonov|Trifonov]] (1990),<ref name=Trifonov1990 /> without changing the essence of the linguistic complexity approach.<ref name=Gabrielian1999>{{Cite journal | last1 = Gabrielian | first1 = A. | title = Sequence complexity and DNA curvature | doi = 10.1016/S0097-8485(99)00007-8 | journal = Computers & Chemistry | volume = 23 | issue = 3–4 | pages = 263–201 | year = 1999 | pmid = | pmc = }}</ref><ref name=Orlov2004>{{Cite journal | last1 = Orlov | first1 = Y. L. | last2 = Potapov | first2 = V. N. | doi = 10.1093/nar/gkh466 | title = Complexity: An internet resource for analysis of DNA sequence complexity | journal = Nucleic Acids Research | volume = 32 | issue = Web Server issue | pages = W628–W633 | year = 2004 | pmid = 15215465| pmc =441604 }}</ref><ref name=Janson2004>{{Cite journal | last1 = Janson | first1 = S. | last2 = Lonardi | first2 = S. | last3 = Szpankowski | first3 = W. | author3-link = Wojciech Szpankowski | title = On average sequence complexity | doi = 10.1016/j.tcs.2004.06.023 | journal = Theoretical Computer Science | volume = 326 | pages = 213–227 | year = 2004 | pmid = | pmc = }}</ref>
The meaning of LC may be better understood by regarding the presentation of a sequence as a [[Tree (data structure)|tree]] of all subsequences of the given sequence. The most complex sequences have maximally balanced trees, while the measure of imbalance or tree asymmetry serves as a [[Computer linguistics|complexity measure]]
{{nb5}} <math>C = U_1 U_2...U_i....U_w </math>
|