PubMed Health⌕ Search

Biomedical subjects

A Milosavljević

Publications and source records attributed to A Milosavljević.

14 recordsLinked to original sources

DNA sequence recognition by hybridization to short oligomers: experimental verification of the method on the E. coli genome.

A newly developed method for sequence recognition by hybridization to short oligomers is verified for the first time in genome-scale experiments. The experiments involved hybridization of 15,328 randomly selected 2-kb genomic clones of Escherichia coli with 997 short oligomer probes to detect complementary oligomers within the clones. Lists of oligomers detected within individual clones were compiled into a database. The database was then searched using known E. coli sequences as queries. The goal was to recognize the clones that are identical or similar to the query sequences. A total of 76 putative recognitions were tested in two separate but complementary recognition experiments. The results indicate high specificity of recognition. Current and prospective applications of this novel method are discussed.

Base Sequence↗

Genome-scale DNA sequence recognition by hybridization to short oligomers.

Recently developed hybridization technology (Drmanac et al. 1994) enables economical large-scale detection of short oligomers within DNA fragments. The newly developed recognition method (Milosavljević 1995b) enables comparison of lists of oligomers detected within DNA fragments against known DNA sequences. We here describe an experiment involving a set of 4,513 distinct genomic E.coli clones of average length 2kb, each hybridized with 636 randomly selected short oligomer probes. High hybridization signal with a particular probe was used as an indication of the presence of a complementary oligomer in the particular clone. For each clone, a list of oligomers with highest hybridization signals was compiled. The database consisting of 4,513 oligomer lists was then searched using known E.coli sequences as queries in an attempt to identify the clones that match the query sequence. Out of a total of 11 clones that were recognized at highest significance level by our method, 8 were single-pass sequenced from both ends. The single-pass sequenced ends were then compared against the query sequences. The sequence comparisons confirmed 7 out of the total of 8 examined recognitions. This experiment represents the first successful example of genome-scale sequence recognition based on hybridization data.

Base Sequence↗

Clone clustering by hybridization.

DNA sequencing by hybridization (SBH) Format 1 technique is based on experiments in which thousands of short oligomers are consecutively hybridized with dense arrays of clones. In this paper we present the description of a method for obtaining hybridization signatures for individual clones that guarantees reproducibility despite a wide range of variations in experimental circumstances, a sensitive method for signature comparison at prespecified significance levels, and a clustering algorithm that correctly identifies clusters of significantly similar signatures. The methods and the algorithm have been verified experimentally on a control set of 422 signatures that originate from 9 distinct clones of known sequence. Experiments indicate that only 30 to 50 oligomer probes suffice for correct clustering. This information about the identity of clones can be used to guide both genomic and cDNA sequencing by SBH or by standard gel-based methods.

Algorithms↗

DNA sequence recognition by hybridization to short oligomers.

A format 1 technology for performing massive hybridization experiments has been developed as part of the sequencing by hybridization (SBH) project. Arrays of tens of thousands of clones are interrogated with short oligomer probes to determine sets of oligomers that are present in individual clones. SBH requires highly discriminative hybridizations with a large number of probes. One of the main uses of a reconstructed DNA sequence is in a similarity search against databases of known DNA. We argue that sequence reconstruction, even partial, should not be performed for this particular purpose; we provide an information-theoretic proof that the oligomer lists obtained from hybridization experiments should be used directly for similarity searches. We propose a similarity search method that takes full advantage of the subword structure of positively identified oligomers within a clone. The method tolerates error in hybridization experiments, requires fewer probes than necessary for sequencing, and is computationally efficient. To enable direct sequence recognition, we apply the recently developed method of sequence comparison that is based on minimal length encoding and algorithimic mutual information. The method has been tested on both real and simulated data and has led to a correct identification of clones based on hybridizations with 109 short oligomer probes. The method is applicable to hybridization data that comes from both format 1 and format 2 (sequencing chip) hybridization experiments. The sequence recognition method can provide targeting information for large-scale DNA sequencing by gel-based methods or by hybridization.

Algorithms↗

Sequence comparisons via algorithmic mutual information.

One of the main problems in DNA and protein sequence comparisons is to decide whether observed similarity of two sequences should be explained by their relatedness or by mere presence of some shared internal structure, e.g., shared internal tandem repeats. The standard methods that are based on statistics or classical information theory can be used to discover either internal structure or mutual sequence similarity, but cannot take into account both. Consequently, currently used methods for sequence comparison employ "masking" techniques that simply eliminate sequences that exhibit internal repetitive structure prior to sequence comparisons. The "masking" approach precludes discovery of homologous sequences of moderate or low complexity, which abound at both DNA and protein levels. As a solution to this problem, we propose a general method that is based on algorithmic information theory and minimal length encoding. We show that algorithmic mutual information factors out the sequence similarity that is due to shared internal structure and thus enables discovery of truly related sequences. We extend that recently developed algorithmic significance method (Milosavljević & Jurka 1993) to show that significance depends exponentially on algorithmic mutual information.

Algorithms↗

Discovering simple DNA sequences by the algorithmic significance method.

A new method, 'algorithmic significance', is proposed as a tool for discovery of patterns in DNA sequences. The main idea is that patterns can be discovered by finding ways to encode the observed data concisely. In this sense, the method can be viewed as a formal version of the Occam's Razor principle. In this paper the method is applied to discover significantly simple DNA sequences. We define DNA sequences to be simple if they contain repeated occurrences of certain 'words' and thus can be encoded in a small number of bits. Such definition includes minisatellites and microsatellites. A standard dynamic programming algorithm for data compression is applied to compute the minimal encoding lengths of sequences in linear time. An electronic mail server for identification of simple sequences based on the proposed method has been installed at the Internet address pythia/anl.gov.

Algorithms↗

Discovering sequence similarity by the algorithmic significance method.

The minimal-length encoding approach is applied to define concept of sequence similarity. A sequence is defined to be similar to another sequence or to a set of keywords if it can be encoded in a small number of bits by taking advantage of common subwords. Minimal-length encoding of a sequence is computed in linear time, using a data compression algorithm that is based on a dynamic programming strategy and the directed acyclic word graph data structure. No assumptions about common word ("k-tuple") length are made in advance, and common words of any length are considered. The newly proposed algorithmic significance method provides an exact upper bound on the probability that sequence similarity has occurred by chance, thus eliminating the need for any arbitrary choice of similarity thresholds. Preliminary experiments indicate that a small number of keywords can positively identify a DNA sequence, which is extremely relevant in the context of partial sequencing by hybridization.

Algorithms↗

[Hemolytic and aplastic crisis as a cause of anemia in congenital hemolytic diseases (spherocytosis)].

The author shows 45 patients suffering from congenital spherocytosis, tested by Cr-51. The author in 8 cases has found decreased T1/2 without anemia and splenomegaly, and in 13 cases with reduced T1/2, with anemia and without splenomegaly, and in 24 cases remarkably reduced T1/2, with anemia and congestive splenomegaly (proved increased blood plasma volume and increased index spleen: heart first day of testing). The author emphasizes the importance of "aplastic crisis" which is expressed in infection or other disturbances, as the cause of anemia, as well as the importance of "hemolytic crisis" which is expressed in increased splenic activity, also as the cause of anemia in congenital spherocytosis.

Adolescent↗