PubMed Health⌕ Search

Biomedical subjects

C Korostensky

Publications and source records attributed to C Korostensky.

7 recordsLinked to original sources

Darwin v. 2.0: an interpreted computer language for the biosciences.

MOTIVATION: We announce the availability of the second release of Darwin v. 2.0, an interpreted computer language especially tailored to researchers in the biosciences. The system is a general tool applicable to a wide range of problems. RESULTS: This second release improves Darwin version 1.6 in several ways: it now contains (1) a larger set of libraries touching most of the classical problems from computational biology (pairwise alignment, all versus all alignments, tree construction, multiple sequence alignment), (2) an expanded set of general purpose algorithms (search algorithms for discrete problems, matrix decomposition routines, complex/long integer arithmetic operations), (3) an improved language with a cleaner syntax, (4) better on-line help, and (5) a number of fixes to user-reported bugs. AVAILABILITY: Darwin is made available for most operating systems free of char ge from the Computational Biochemistry Research Group (CBRG), reachable at http://chrg.inf.ethz.ch. CONTACT: darwin@inf.ethz.ch

Algorithms↗

Using traveling salesman problem algorithms for evolutionary tree construction.

MOTIVATION: The construction of evolutionary trees is one of the major problems in computational biology, mainly due to its complexity. RESULTS: We present a new tree construction method that constructs a tree with minimum score for a given set of sequences, where the score is the amount of evolution measured in PAM distances. To do this, the problem of tree construction is reduced to the Traveling Salesman Problem (TSP). The input for the TSP algorithm are the pairwise distances of the sequences and the output is a circular tour through the optimal, unknown tree plus the minimum score of the tree. The circular order and the score can be used to construct the topology of the optimal tree. Our method can be used for any scoring function that correlates to the amount of changes along the branches of an evolutionary tree, for instance it could also be used for parsimony scores, but it cannot be used for least squares fit of distances. A TSP solution reduces the space of all possible trees to 2n. Using this order, we can guarantee that we reconstruct a correct evolutionary tree if the absolute value of the error for each distance measurement is smaller than f2.gif" BORDER="0">, where f3.gif" BORDER="0">is the length of the shortest edge in the tree. For data sets with large errors, a dynamic programming approach is used to reconstruct the tree. Finally simulations and experiments with real data are shown.

Algorithms↗

A genetic system based on split-ubiquitin for the analysis of interactions between membrane proteins in vivo.

A detection system for interactions between membrane proteins in vivo is described. The system is based on split-ubiquitin [Johnsson, N. & Varshavsky, A. (1994) Proc. Natl. Acad. Sci. USA 91, 10340-10344]. Interaction between two membrane proteins is detected by proteolytic cleavage of a protein fusion. The cleavage releases a transcription factor, which activates reporter genes in the nucleus. As a result, interaction between membrane proteins can be analyzed by the means of a colorimetric assay. We use membrane proteins of the endoplasmic reticulum as a model system. Wbp1p and Ost1p are both subunits of the oligosaccharyl transferase membrane protein complex. The Alg5 protein also localizes to the membrane of the endoplasmic reticulum, but does not interact with the oligosaccharyltransferase. Specific interactions are detected between Wbp1p and Ost1p, but not between Wbp1p and Alg5p. The new system might be useful as a genetic and biochemical tool for the analysis of interactions between membrane proteins in vivo.

Cloning, Molecular↗

An algorithm for the identification of proteins using peptides with ragged N- or C-termini generated by sequential endo- and exopeptidase digestions.

We have developed an algorithm (MassDynSearch) for identifying proteins using a combination of peptide masses with small associated sequences (tags). Unlike the approach developed by Matthias Mann, 'Tag searching', in which the sequence tags are generated by gas phase fragmentation of peptides in a mass spectrometer, 'Rag Tag' searching uses peptide tags which are generated enzymatically or chemically. The protein is digested either chemically or with an endopeptidase and the resultant mixture is then subjected to partial exopeptidase degradation. The mixture is analyzed by matrix assisted laser desorption and ionization time of flight mass spectrometry and a list of intact peptide masses is generated, each associated with a set of degradation product masses which serve as unique tags. These 'tagged masses' are used as the input to an algorithm we have written, MassDynSearch, which searches protein and DNA databases for proteins which contain similar tagged motifs. The method is simple, rapid and can be fully automated. The main advantage of this approach is that the specificity of the initial digestion is unimportant since multiple peptides with tags are used to search the database. This is especially useful for proteins like membrane, cytoskeletal, and other proteins where specific endopeptidases are less efficient and lower specificity proteases such as chymotrypsin, pepsin, and elastase must be used.

Algorithms↗

A predicted consensus structure for the N-terminal fragment of the heat shock protein HSP90 family.

A secondary structure has been predicted for the heat shock protein HSP90 family from an aligned set of homologous protein sequences by using a transparent method in both manual and automated implementation that extracts conformational information from patterns of variation and conservation within the family. No statistically significant sequence similarity relates this family to any protein with known crystal structure. However, the secondary structure prediction, together with the assignment of active site positions and possible biochemical properties, suggest that the fold is similar to that seen in N-terminal domain of DNA gyrase B (the ATPase fragment).

Algorithms↗

Probing protein function using a combination of gene knockout and proteome analysis by mass spectrometry.

Recently the determination of the genome sequences of three procaryotes (Haemophilus influenzae, Methanococcus jannaschii and Mycoplasma genitalium) as well as the first eucaryotic genome (Saccharomyces cerevisiae) were completed. Between 40-60% of the genes were found to code for proteins to which no function could be assigned. We describe an approach which combines proteome analysis (mapping of expressed proteins isolated by two-dimensional polyacrylamide gel electrophoresis to the genome) with genetic manipulations to study the complex pattern of protein regulation occurring in Escherichia coli in response to sulfate starvation. We have previously described the upregulation of eight spots on two-dimensional (2-D) gels in response to sulfate starvation and the assignment of six of these to entries in the E. coli genome sequence (Quadroni et al., Eur. J. Biochem. 1996, 239, 773-781). Here we describe the identification of the remaining two proteins which are encoded in a sulfate-controlled operon in the 21.5' region of the E. coli genome. Upregulated protein spots were cut from multiple 2-D gels collected and run on a modified funnel gel to concentrate the proteins and remove the sodium dodecyl sulfate before digestion. The peptide masses obtained from the digests were used to search the SwissProt database or a six-frame translation of the EMBL DNA database using a peptide mass fingerprinting algorithm. A digest can be reanalyzed after deuterium exchange to obtain a second, orthogonal data set to increase the confidence level of protein identification. The digests of the remaining unidentified proteins were used for peptide fragment generation using either post-source decay in a matrix-assisted laser desorption ionization (MALDI) time-of-flight mass spectrometer or collision-induced dissociation (CID) coupled mass spectrometry (MS/MS) with triple stage quadrupole or ion trap mass spectrometers. The spectra were used as peptide fragment fingerprints to search the SwissProt and EMBL databases.

Amino Acid Sequence↗

Evaluation measures of multiple sequence alignments.

Multiple sequence alignments (MSAs) are frequently used in the study of families of protein sequences or DNA/RNA sequences. They are a fundamental tool for the understanding of the structure, functionality and, ultimately, the evolution of proteins. A new algorithm, the Circular Sum (CS) method, is presented for formally evaluating the quality of an MSA. It is based on the use of a solution to the Traveling Salesman Problem, which identifies a circular tour through an evolutionary tree connecting the sequences in a protein family. With this approach, the calculation of an evolutionary tree and the errors that it would introduce can be avoided altogether. The algorithm gives an upper bound, the best score that can possibly be achieved by any MSA for a given set of protein sequences. Alternatively, if presented with a specific MSA, the algorithm provides a formal score for the MSA, which serves as an absolute measure of the quality of the MSA. The CS measure yields a direct connection between an MSA and the associated evolutionary tree. The measure can be used as a tool for evaluating different methods for producing MSAs. A brief example of the last application is provided. Because it weights all evolutionary events on a tree identically, but does not require the reconstruction of a tree, the CS algorithm has advantages over the frequently used sum-of-pairs measures for scoring MSAs, which weight some evolutionary events more strongly than others. Compared to other weighted sum-of-pairs measures, it has the advantage that no evolutionary tree must be constructed, because we can find a circular tour without knowing the tree.

Algorithms↗