PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.

Arabidopsis thaliana is the most widely-studied plant today. The concerted efforts of over 11 000 researchers and 4000 organizations around the world are generating a rich diversity and quantity of information and materials. This information is made available through a comprehensive on-line resource called the Arabidopsis Information Resource (TAIR) (http://arabidopsis.org), which is accessible via commonly used web browsers and can be searched and downloaded in a number of ways. In the last two years, efforts have been focused on increasing data content and diversity, functionally annotating genes and gene products with controlled vocabularies, and improving data retrieval, analysis and visualization tools. New information include sequence polymorphisms including alleles, germplasms and phenotypes, Gene Ontology annotations, gene families, protein information, metabolic pathways, gene expression data from microarray experiments and seed and DNA stocks. New data visualization and analysis tools include SeqViewer, which interactively displays the genome from the whole chromosome down to 10 kb of nucleotide sequence and AraCyc, a metabolic pathway database and map tool that allows overlaying expression data onto the pathway diagrams. Finally, we have recently incorporated seed and DNA stock information from the Arabidopsis Biological Resource Center (ABRC) and implemented a shopping-cart style on-line ordering system.

Arabidopsis↗

Expression, purification and characterization of secreted recombinant human insulin-like growth factor-I (IGF-I) and the potent variant des(1-3) IGF-I in Chinese hamster ovary cells.

Recombinant human insulin-like growth factor-I (hIGF-I) and a biologically potent variant lacking the N-terminal tripeptide (des(1-3)IGF-I) were produced from transfected Chinese hamster ovary cells. The constructs encoding the signal peptide, sequence of the mature peptide and a C-terminal extension peptide were expressed under the control of a Rous sarcoma virus promoter. Successfully transfected clones secreting correctly processed recombinant hIGF-I or des(1-3)IGF-I were selected by their secretion of IGF-I-like activity into the culture medium. The recombinant peptides were purified to homogeneity as assessed by high-performance liquid chromatography and N-terminal sequence analysis. The purified recombinant peptides exhibited biological potencies equivalent to authentic IGF-I and des(1-3)IGF-I respectively.

Amino Acid Sequence↗

Gene discovery in neuropharmacological and behavioral studies using Affymetrix microarray data.

We describe methods and software tools for doing data analysis based on Affymetrix microarray data, emphasizing often neglected issues. In our experience with neuroscience studies, experimental design and quality assessment are vital. We also describe in detail the pre-processing methods we have found useful for Affymetrix data. Finally, we summarize the statistical literature and describe some pitfalls in the post-processing analysis.

Animals↗

Integrative analysis of the cancer transcriptome.

DNA microarrays have been widely applied to the study of human cancer, delineating myriad molecular subtypes of cancer, many of which are associated with distinct biological underpinnings, disease progression and treatment response. These primary analyses have begun to decipher the molecular heterogeneity of cancer, but integrative analyses that evaluate cancer transcriptome data in the context of other data sources are often capable of extracting deeper biological insight from the data. Here we discuss several such integrative computational and analytical approaches, including meta-analysis, functional enrichment analysis, interactome analysis, transcriptional network analysis and integrative model system analysis.

Animals↗

Targeted analysis and discovery of posttranslational modifications in proteins from methanogenic archaea by top-down MS.

For more complete characterization of DNA-predicted proteins (including their posttranslational modifications) a "top-down" approach using high-resolution tandem MS is forwarded here by its application to methanogens in both hypothesis-driven and discovery modes, with the latter dependent on new automation benchmarks for intact proteins. With proteins isolated from ribosomes and whole-cell lysates of Methanococcus jannaschii (approximately 1,800 genes) using a 2D protein fractionation method, 72 gene products were identified and characterized with 100% sequence coverage via automated fragmentation of intact protein ions in a custom quadrupole/Fourier transform hybrid mass spectrometer. Three incorrect start sites and two modifications were found, with one of each determined for MJ0556, a 20-kDa protein with an unknown methylation at approximately 50% occupancy in stationary phase cells. The separation approach combined with the quadrupole/Fourier transform hybrid mass spectrometer allowed targeted and efficient comparison of histones from M. jannaschii, Methanosarcina acetivorans (largest Archaeal genome, 5.8 Mb), and yeast. This finding revealed a striking difference in the posttranslational regulation of DNA packaging in Eukarya vs. the Archaea. This study illustrates a significant evolutionary step for the MS tools available for characterization of WT proteins from complex proteomes without proteolysis.

Amino Acid Sequence↗

CSS-Palm: palmitoylation site prediction with a clustering and scoring strategy (CSS).

UNLABELLED: Palmitoylation is an important post-translational lipid modification of proteins. Unlike prenylation and myristoylation, palmitoylation is a reversible covalent modification, allowing for dynamic regulation of multiple complex cellular systems. However, in vivo or in vitro identification of palmitoylation sites is usually time-consuming and labor-intensive. So in silico predictions could help to narrow down the possible palmitoylation sites, which can be used to guide further experimental design. Previous studies suggested that there is no unique canonical motif for palmitoylation sites, so we hypothesize that the bona fide pattern might be compromised by heterogeneity of multiple structural determinants with different features. Based on this hypothesis, we partition the known palmitoylation sites into three clusters and score the similarity between the query peptide and the training ones based on BLOSUM62 matrix. We have implemented a computer program for palmitoylation site prediction, Clustering and Scoring Strategy for Palmitoylation Sites Prediction (CSS-Palm) system, and found that the program's prediction performance is encouraging with highly positive Jack-Knife validation results (sensitivity 82.16% and specificity 83.17% for cut-off score 2.6). Our analyses indicate that CSS-Palm could provide a powerful and effective tool to studies of palmitoylation sites. AVAILABILITY: CSS-Palm is implemented in PHP/PERL+MySQL and can be freely accessed at http://bioinformatics.lcd-ustc.org/css_palm/ CONTACT: yaoxb@ustc.edu.cn; xuyn@bmb.uga.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bionformatics online.

Algorithms↗

Host range relationships and the evolution of canine parvovirus.

Canine parvovirus (CPV) is an example of an unusual class of emerging virus-those that gain an altered host range through genetic variation and subsequently become widespread pathogens of their new and previously resistant host species. CPV was first detected in 1978 as the cause of new diseases in dogs throughout the world, when it rapidly spread throughout domestic populations, as well as becoming widespread in wild dogs. CPV was soon shown to be a variant of the long recognized feline panleukopenia virus (FPV), from which it differed in less than 1% at the nucleotide sequence level. Genetic analysis showed that virtually all of the biological differences between CPV and FPV, including the canine host range, were determined by three or four sequence differences in the viral capsid protein gene. Analysis of the atomic structures of the CPV and FPV capsids showed that the differences controlling host range were located within two different structural regions and were exposed on the capsid surface. The CPV which first emerged in 1978 appeared to be derived from a single ancestral sequence, which has allowed the ready analysis of the subsequent evolution of the virus in nature. Sequence analysis has also revealed that CPV strains have undergone a series of evolutionary selections in nature which have resulted in the global distribution of new virus variants. This was first seen in the global replacement between 1979 and 1981 of the original (1978) strain of the virus by a genetically and antigenically variant strain, and the subsequent widespread selection of other variants which have also become globally distributed. The genetic and antigenic variation in the virus strains was also correlated with changes in the host range of the virus, in particular in the ability to replicate in cats, and in canine host range differences seen in tissue culture cells.

Amino Acid Sequence↗

Gene expression analysis on biochemical networks using the Potts spin model.

MOTIVATION: Microarray technology allows us to profile the expression of a large subset or all genes of a cell. Biochemical research over the last three decades has elucidated an increasingly complete image of the metabolic architecture. For less complex organisms, such as Escherichia coli, the biochemical network has been described in much detail. Here, we investigate the clustering of such networks by applying gene expression data that define edge lengths in the network. RESULTS: The Potts spin model is used as a nearest neighbour based clustering algorithm to discover fragmentation of the network in mutants or in biological samples when treated with drugs. As an example, we tested our method with gene expression data from E.coli treated with tryptophan excess, starvation and trpyptophan repressor mutants. We observed fragmentation of the tryptophan biosynthesis pathway, which corresponds well to the commonly known regulatory response of the cells.

Algorithms↗

"Plasmo2D": an ancillary proteomic tool to aid identification of proteins from Plasmodium falciparum.

Bioinformatics tools to aid gene and protein sequence analysis have become an integral part of biology in the post-genomic era. Release of the Plasmodium falciparum genome sequence has allowed biologists to define the gene and the predicted protein content as well as their sequences in the parasite. Using pI and molecular weight as characteristics unique to each protein, we have developed a bioinformatics tool to aid identification of proteins from Plasmodium falciparum. The tool makes use of a Virtual 2-DE generated by plotting all of the proteins from the Plasmodium database on a pI versus molecular weight scale. Proteins are identified by comparing the position of migration of desired protein spots from an experimental 2-DE and that on a virtual 2-DE. The procedure has been automated in the form of user-friendly software called "Plasmo2D". The tool can be downloaded from http://144.16.89.25/Plasmo2D.zip.

Animals↗

A bioinformatics pipeline for high-throughput microbial multilocus sequence typing (MLST) analyses.

Multilocus sequence typing (MLST) analysis for semi-routine applications is hindered by the downstream, manually intensive steps of processing the raw sequence data files. This report describes the development of an MLST pipeline that automates DNA sequence editing and analysis in order to significantly reduce the time required for processing data. Validation using a pneumococcal dataset revealed complete agreement between the results generated by manual and automated workflows. The MLST pipeline was developed for both double-strand and single-strand sequencing.

Bacteria↗

Iterative stepwise discriminant analysis: a meta-algorithm for detecting quantitative sequence motifs.

An algorithm is presented for detecting a quantitative pattern in peptide fragments that bind class II major histocompatibility complex (MHC) molecules. It is referred to as a meta-algorithm because it requires successive applications of Stepwise Discriminate Analysis (SDA). On every iteration the best subsequence candidates are selected from sequences known to bind class II MHC molecules. When SDA compares probable binding subsequences with subsequences known not to bind class II MHC molecules, a quantitative model emerges that is capable of classifying subsequences as binding or non-binding. In an iterative manner, the resultant model is utilized as a criterion for selecting probable binding subsequence candidates. The procedure is repeated until models converge. In the illustrated examples, the final models correctly classify over 95% of the peptides in a database of peptides whose binding affinity for HLA-DR1 is known. The final model can then be used to predict the binding affinity of peptides that have not yet been laboratory tested.

Algorithms↗

Structured motifs search.

In this paper, we describe an algorithm for the localization of structured models, i.e. sequences of (simple) motifs and distance constraints. It basically combines standard pattern matching procedures with a constraint satisfaction solver, and it has the ability, not present in similar tools, to search for partial matches. A significant feature of our approach, especially in terms of efficiency for the application context, is that the (potentially) exponentially many solutions to the considered problem are represented in compact form as a graph. Moreover, the time and space necessary to build the graph are linear in the number of occurrences of the component patterns.

Algorithms↗

Cellular and molecular biology of Candida albicans estrogen response.

Candida albicans is the most common etiological agent of vaginal candidiasis. Elevated host estrogen levels and the incidence of vaginal candidiasis are positively associated. Elevated estrogen levels may affect host and/or fungal cells. This study investigates the effect of 17-beta-estradiol, 17-alpha-estradiol, ethynyl estradiol, and estriol on several C. albicans strains at concentrations ranging from 10(-5) to 10(-10) M. The addition of 17-beta-estradiol or ethynyl estradiol to C. albicans cells caused an increase in the number of cells forming germ tubes and an increase in germ tube length in a dose- and strain-dependent manner. The addition of 17-alpha-estradiol or estriol did not have a significant effect on germ tube formation by the cultured cells. Exposure to exogenous estrogens did not significantly change the biomass of any C. albicans culture tested. The transcriptional profile of estrogen-treated C. albicans cells showed increased expression of CDR1 and CDR2 across several strain-estrogen concentration-time point combinations, suggesting that these genes are the most responsive to estrogen exposure. Analysis of strain DSY654, which lacks the CDR1 and CDR2 coding sequences, showed a significantly decreased number of germ tube-forming cells in the presence of 17-beta-estradiol. PDR16 was the most highly up-regulated gene in strain DSY654 under these growth conditions. The cell biology and gene expression data from this study led to a model that proposes how components of the phospholipid and sterol metabolic pathways may interact to affect C. albicans germ tube formation and length.

Biomass↗

Exophiala crusticola anam. nov. (affinity Herpotrichiellaceae), a novel black yeast from biological soil crusts in the Western United States.

A novel black yeast-like fungus, Exophiala crusticola, is described based on two closely related isolates from biological soil crust (BSC) samples collected on the Colorado Plateau (Utah) and in the Great Basin desert (Oregon), USA. Their morphology places them in the anamorphic genus Exophiala, having affinities to the family Herpotrichiellaceae (Ascomycota). Phylogenetic analysis of their D1/D2 large subunit nuclear ribosomal RNA (LSU nrRNA) gene sequences suggests that they represent a distinct species. The closest known putative relative to Exophiala crusticola is Capronia coronata Samuels, isolated from decorticated wood in Westland County, New Zealand. The holotype for Exophiala crusticola anam. nov. is UAMH 10686 and the type strain is CP141bT (=ATCC MYA-3639T=CBS 119970T=DSM 16793T). Dark-pigmented fungi appear to constitute an important heterotrophic component of soil crusts and Exophiala crusticola represents the first description of a dematiaceous fungus isolated from BSCs.

DNA, Fungal↗

Apoptosis Gene Information System--AGIS.

Genes implicated in apoptosis have great relevance to biology, medicine and oncology. Here, we describe a unique resource, Apoptosis Gene Information System (AGIS) that provides data for over 2400 genes involved directly or indirectly, in apoptotic pathways of more than 350 different organisms. The organization of this information system is based on the principle of one-gene, one record. AGIS will be updated on a six monthly basis as new information becomes available. AGIS can be accessed at: http://www.cellfate.org/AGIS/.

Amino Acid Sequence↗

Long branch attraction, taxon sampling, and the earliest angiosperms: Amborella or monocots?

BACKGROUND: Numerous studies, using in aggregate some 28 genes, have achieved a consensus in recognizing three groups of plants, including Amborella, as comprising the basal-most grade of all other angiosperms. A major exception is the recent study by Goremykin et al. (2003; Mol. Biol. Evol. 20:1499-1505), whose analyses of 61 genes from 13 sequenced chloroplast genomes of land plants nearly always found 100% support for monocots as the deepest angiosperms relative to Amborella, Calycanthus, and eudicots. We hypothesized that this conflict reflects a misrooting of angiosperms resulting from inadequate taxon sampling, inappropriate phylogenetic methodology, and rapid evolution in the grass lineage used to represent monocots. RESULTS: We used two main approaches to test this hypothesis. First, we sequenced a large number of chloroplast genes from the monocot Acorus and added these plus previously sequenced Acorus genes to the Goremykin et al. (2003) dataset in order to explore the effects of altered monocot sampling under the same analytical conditions used in their study. With Acorus alone representing monocots, strongly supported Amborella-sister trees were obtained in all maximum likelihood and parsimony analyses, and in some distance-based analyses. Trees with both Acorus and grasses gave either a well-supported Amborella-sister topology or else a highly unlikely topology with 100% support for grasses-sister and paraphyly of monocots (i.e., Acorus sister to "dicots" rather than to grasses). Second, we reanalyzed the Goremykin et al. (2003) dataset focusing on methods designed to account for rate heterogeneity. These analyses supported an Amborella-sister hypothesis, with bootstrap support values often conflicting strongly with cognate analyses performed without allowing for rate heterogeneity. In addition, we carried out a limited set of analyses that included the chloroplast genome of Nymphaea, whose position as a basal angiosperm was also, and very recently, challenged. CONCLUSIONS: These analyses show that Amborella (or Amborella plus Nymphaea), but not monocots, is the sister group of all other angiosperms among this limited set of taxa and that the grasses-sister topology is a long-branch-attraction artifact leading to incorrect rooting of angiosperms. These results highlight the danger of having lots of characters but too few and, especially, molecularly divergent taxa, a situation long recognized as potentially producing strongly misleading molecular trees. They also emphasize the importance in phylogenetic analysis of using appropriate evolutionary models.

Artifacts↗

Mitochondrial DNA and retroviral RNA analyses of archival oral polio vaccine (OPV CHAT) materials: evidence of macaque nuclear sequences confirms substrate identity.

Inoculation of live experimental oral poliovirus vaccines (OPV CHAT) during the 1950s in central Africa has been proposed to account for the introduction of HIV into human populations. For this to have occurred, it would have been necessary for chimpanzee rather than macaque kidney epithelial cells to have been included in the preparation of early OPV materials. Theoretically, this could have led to contamination with a progenitor of HIV-1 derived from a related simian immunodeficiency virus of chimpanzees (SIVCPZ). In this article we present further detailed analyses of two samples of OPV, CHAT 10A-11 and CHAT 6039/Yugo, which were used in early human trials of poliovirus vaccination. Recovery of poliovirus by culture techniques confirmed the biological viability of the vaccines and sequence analysis of poliovirus RNA specifically identified the presence of the CHAT strain. Independent nested sets of oligonucleotide primers specific for HIV-1/SIVCPZ and HIV-2/SIVMAC/SIVSM phylogenetic lineages, respectively, indicated no evidence of HIV/SIV RNA in either vaccine preparation, at a sensitivity of 100 RNA equivalents/ml. Analysis of cellular substrate by the amplification of two distinct regions of mitochondrial DNA (D-loop control region and 12S ribosomal sequences) revealed no evidence of chimpanzee cellular sequences. However, this approach positively identified rhesus and cynomolgus macaque DNA for the CHAT 10A-11 and CHAT 6039/Yugo vaccine preparations, respectively. Analysis of multiple clones of mtDNA 12S rDNA indicated a relatively high number of nuclear mitochondrial DNA sequences (numts) in the CHAT 10A-11 material, but confirmed the macaque origin of cellular substrate used in vaccine preparation. These data reinforce earlier findings on this topic providing no evidence to support the contention that poliovirus vaccination was responsible for the introduction of HIV into humans and sparking the AIDS pandemic.

Animals↗

Postgenomic bioinformatic analysis of yeast artificial chromosome sequence.

The free availability of multiple genomic sequences represents one of the greatest advances in biology of the new millennium, and promises to revolutionize our ability to determine and treat the causes of human disease. This chapter highlights a number of basic, freely available, and user-friendly bioinformatic techniques that can be used to predict the functional genetic contents of specific yeast artificial chromosome (YAC) clones. The content of this chapter is written for the level of graduate students, who may be relatively inexperienced with the use of computers for analyzing DNA sequences. The basic instructions that allow the identification of the genomic sequence of interest and to download this sequence onto a personal computer from an online database are presented. Simple instructions are also given on how to perform basic sequence manipulations, how to use online tools to design polymerase chain reaction primers, and how to map restriction sites. Also described are more complicated programs that rapidly and efficiently perform genome alignments that, in addition to predicting the location of protein coding sequences, allow the prediction of functional genomic sequences, such as cis regulatory elements and scaffold/matrix attachment sites. The availability of genomic sequences and the rapidly expanding numbers of predictive programs that allow the predictive analysis of these sequences promises to greatly facilitate the use of YAC clones in the search for the causes of disease.

Chromosomes, Artificial, Yeast↗