PubMed Health⌕ Search

Biomedical subjects

Heikki Hyyrö

Publications and source records attributed to Heikki Hyyrö.

3 recordsLinked to original sources

Genome-wide selection of unique and valid oligonucleotides.

Functional genomics methods are used to investigate the huge amount of information contained in genomes. Numerous experimental methods rely on the use of oligo- or polynucleotides. Nucleotide strand hybridization forms the underlying principle for these methods. For all these techniques, the probes should be unique for analyzed genes. In addition to being unique for the studied genes, the probes should fulfill a large number of criteria to be usable and valid. The criteria include for example, avoidance of self-annealing, suitable melting temperature and nucleotide composition. We developed a method for searching unique and valid oligonucleotides or probes for genes so that there is not even a similar (approximate) occurrence in any other location of the whole genome. By using probe size 25, we analyzed 17 complete genomes representing a wide range of both prokaryotic and eukaryotic organisms. More than 92% of all the genes in the investigated genomes contained valid oligonucleotides. Extensive statistical tests were performed to characterize the properties of unique and valid oligonucleotides. Unique and valid oligonucleotides were relatively evenly distributed in genes except for the beginning and end, which were somewhat overrepresented. The flanking regions in eukaryotes were clearly underrepresented among suitable oligonucleotides. In addition to distributions within genes, the effects on codon and amino acid usage were also studied.

Amino Acids↗

On exact string matching of unique oligonucleotides.

Unique, gene-specific oligonucleotides are used for many genetic investigations such as polymerase chain reaction, gene cloning, microarray technology and antisense DNA studies. It is a computationally demanding task to extract these oligonucleotides from DNA databases. We studied the problem from the point of view of the string matching problem. We implemented and tested several exact string matching algorithms and modified the implementations to be as effective as possible. Ten different implementations were tested on yeast genomic sequence data. The run times for the best algorithms were significantly improved compared to conventional approaches, while in principle, i.e. in respect of theoretical time complexity, these algorithms do not actually differ essentially from each other.

Algorithms↗

An O(N2) algorithm for discovering optimal Boolean pattern pairs.

We consider the problem of finding the optimal combination of string patterns, which characterizes a given set of strings that have a numeric attribute value assigned to each string. Pattern combinations are scored based on the correlation between their occurrences in the strings and the numeric attribute values. The aim is to find the combination of patterns which is best with respect to an appropriate scoring function. We present an O(N2) time algorithm for finding the optimal pair of substring patterns combined with Boolean functions, where N is the total length of the sequences. The algorithm looks for all possible Boolean combinations of the patterns, e.g., patterns of the form p and not q, which indicates that the pattern pair is considered to occur in a given string s, if p occurs in s, AND q does NOT occur in s. An efficient implementation using suffix arrays is presented, and we further show that the algorithm can be adapted to find the best k-pattern Boolean combination in O(Nk) time. The algorithm is applied to mRNA sequence data sets of moderate size combined with their turnover rates for the purpose of finding regulatory elements that cooperate, complement, or compete with each other in enhancing and/or silencing mRNA decay.

3' Untranslated Regions↗