PubMed Health⌕ Search

PubMed · 10380193

A probabilistic approach to consensus multiple alignment.

Abstract

We consider the problem of obtaining the maximum a posteriori probability (MAP) estimate of a consensus ancestral sequence for a set of DNA sequences. Our maximization method, called ASA (dnA Sequence Alignment), can be applied to the refinement of noisy regions of a DNA assembly, to the alignment of genomic functional sites, or to the alignment of any set of DNA sequences related by a star-like phylogeny. Along with the optimal consensus, ASA finds suboptimal solutions together with their relative probabilities. The probabilistic approach makes it possible to establish the limits to which an ancestor can in principle be recovered from diverged sequences. In simulations on rather short synthetic sequences (of length up to 80) with different coverage and error rates ranging from 5% to 30%, ASA restored the consensus from noisy observations essentially as best as is theoretically possible for the given error rates. We also illustrate the performance of ASA on the alignment of E.Coli promoters and the Alu-Sb subfamily of human repeat sequences. Since our model is a special case of a profile HMM, we give a comparison between these two approaches, as well as with other DNA alignment methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

B Lazareva-Ulitsky, D Haussler. 1999. A probabilistic approach to consensus multiple alignment.. https://doi.org/10.1142/9789814447300_0015

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

miR-503-3p promotes epithelial-mesenchymal transition in breast cancer by directly targeting SMAD2 and E-cadherin.

Although progress in clinical and basic research has significantly increased our understanding of breast cancer, little is known about the molecular mechanism underlying breast cancer metastasis. Identification of effective therapeutic targets to prevent breast cancer metastasis is urgently needed. The function of miR-503-3p has been investigated in other cancers, but its role in breast cancer remains undefined. Here, we found that miR-503-3p was overexpressed in breast cancer tissue and plasma compared with adjacent normal breast tissue and with plasma from healthy individuals. Moreover, we identified miR-503-3p to be an oncogene of breast cancer cell proliferation, migration and invasion. Upregulation of miR-503-3p in breast cancer cells inhibited expression of epithelial-mesenchymal transition (EMT)-related protein SMAD2 and the epithelial marker protein E-cadherin by directly binding to their mRNA 3' untranslated region, whereas increased expression of mesenchymal marker proteins, including vimentin and N-cadherin. Taken together, our findings support a critical role for miR-503-3p in induction of breast cancer EMT and suggest that plasma miR-503-3p may be a useful diagnostic biomarker for breast cancer.

Base Sequence↗

Identification and characterization of Prp45p and Prp46p, essential pre-mRNA splicing factors.

Through exhaustive two-hybrid screens using a budding yeast genomic library, and starting with the splicing factor and DEAH-box RNA helicase Prp22p as bait, we identified yeast Prp45p and Prp46p. We show that as well as interacting in two-hybrid screens, Prp45p and Prp46p interact with each other in vitro. We demonstrate that Prp45p and Prp46p are spliceosome associated throughout the splicing process and both are essential for pre-mRNA splicing. Under nonsplicing conditions they also associate in coprecipitation assays with low levels of the U2, U5, and U6 snRNAs that may indicate their presence in endogenous activated spliceosomes or in a postsplicing snRNP complex.

Base Sequence↗

Choline-binding domain as a novel affinity tag for purification of fusion proteins produced in Pichia pastoris.

The choline-binding domain (ChoBD) of the carboxy-terminal region of the Streptococcus pneumoniae amidase LYTA (C-LYTA) presents a strong affinity for tertiary amines. We report a method for single-step purification of proteins expressed in the methylotrophic yeast Pichia pastoris based on the fusion of C-LYTA to the protein of interest. We show that C-LYTA can be efficiently expressed and secreted in this host. Tagged proteins fused to this binding domain can be purified on inexpensive DEAE matrices. It therefore provides a useful system for the purification of recombinant proteins with high specificity suitable for industrial purposes.

Base Sequence↗