PubMed Health⌕ Search

PubMed · 15878123

Identifying optimal incomplete phylogenetic data sets from sequence databases.

Abstract

We introduce a new method for identifying optimal incomplete data sets from large sequence databases based on the graph theoretic concept of alpha-quasi-bicliques. The quasi-biclique method searches large sequence databases to identify useful phylogenetic data sets with a specified amount of missing data while maintaining the necessary amount of overlap among genes and taxa. The utility of the quasi-biclique method is demonstrated on large simulated sequence databases and on a data set of green plant sequences from GenBank. The quasi-biclique method greatly increases the taxon and gene sampling in the data sets while adding only a limited amount of missing data. Furthermore, under the conditions of the simulation, data sets with a limited amount of missing data often produce topologies nearly as accurate as those built from complete data sets. The quasi-biclique method will be an effective tool for exploiting sequence databases for phylogenetic information and also may help identify critical sequences needed to build large phylogenetic data sets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Changhui Yan, J Gordon Burleigh, Oliver Eulenstein. 2005-03-21. Identifying optimal incomplete phylogenetic data sets from sequence databases.. https://doi.org/10.1016/j.ympev.2005.02.008

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

miR-503-3p promotes epithelial-mesenchymal transition in breast cancer by directly targeting SMAD2 and E-cadherin.

Although progress in clinical and basic research has significantly increased our understanding of breast cancer, little is known about the molecular mechanism underlying breast cancer metastasis. Identification of effective therapeutic targets to prevent breast cancer metastasis is urgently needed. The function of miR-503-3p has been investigated in other cancers, but its role in breast cancer remains undefined. Here, we found that miR-503-3p was overexpressed in breast cancer tissue and plasma compared with adjacent normal breast tissue and with plasma from healthy individuals. Moreover, we identified miR-503-3p to be an oncogene of breast cancer cell proliferation, migration and invasion. Upregulation of miR-503-3p in breast cancer cells inhibited expression of epithelial-mesenchymal transition (EMT)-related protein SMAD2 and the epithelial marker protein E-cadherin by directly binding to their mRNA 3' untranslated region, whereas increased expression of mesenchymal marker proteins, including vimentin and N-cadherin. Taken together, our findings support a critical role for miR-503-3p in induction of breast cancer EMT and suggest that plasma miR-503-3p may be a useful diagnostic biomarker for breast cancer.

Base Sequence↗

Characterization of the AlkS/P(alkB)-expression system as an efficient tool for the production of recombinant proteins in Escherichia coli fed-batch fermentations.

The availability of suitable, well-characterized, and robust expression systems remains an essential requirement for successful metabolic engineering and recombinant protein production. We investigated the suitability of the Pseudomonas putida GPo1-derived AlkS/P(alkB) expression system in strictly aqueous cultures. By applying the apolar inducer dicyclopropylketone (DCPK) to express green fluorescent protein (GFP) from this system in Escherichia coli and analyzing the resulting cultures on single-cell level by flow cytometry, we found that this expression system gives rise to a homogeneous population of cells, even though the overall system is expected to have a positive feed-back element in the expression of the regulatory gene alkS. Overexpressing E. coli's serine hydroxymethyltransferase gene glyA, we showed that the system was already fully turned on at inducer concentrations as low as 0.005% (v/v). This allows efficient mass production of recombinant enzymes even though DCPK concentrations decreased from 0.05% to 0.01% over the course of a fully aerated cultivation in aqueous medium. Therefore, we elaborated the optimum induction procedure for production of the biocatalytically promising serine hydroxymethyltransferase and found volumetric and specific productivity to increase with specific growth rate in glucose-limited fed-batch cultures. Acetate excretion as a result of recombinant protein production could be avoided in an optimized fermentation protocol by switching earlier to a linear feed. This protocol resulted in a production of a final cell dry weight (CDW) concentration of 52 g/L, producing recombinant GlyA with a maximum specific activity of 6.3 U/mg total protein.

Base Sequence↗

Genome evolution and functional divergence in Yersinia.

The steadily increasing number of prokaryotic genomes has accelerated the study of genome evolution; in particular, the availability of sets of genomes from closely related bacteria has made exploration of questions surrounding the evolution of pathogenesis tractable. Here we present the results of a detailed comparison of the genomes of Yersinia pseudotuberculosis IP32593 and three strains of Yersinia pestis (CO92, KIM10, and 91001). There appear to be between 241 and 275 multigene families in these organisms. There are 2,568 genes that are identical in the three Y. pestis strains, but differ from the Y. pseudotuberculosis strain. The changes found in some of these families, such as the kinases, proteases, and transporters, are illustrative of how the evolutionary jump from the free-living enteropathogen Y. pseudotuberculosis to the obligate host-borne blood pathogen Y. pestis was achieved. We discuss the composition of some of the most important families and discuss the observed divergence between Y. pseudotuberculosis and Y. pestis homologs.

Base Sequence↗