PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Molecular Sequence Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Indelign: a probabilistic framework for annotation of insertions and deletions in a multiple alignment.

MOTIVATION: A quantitative study of molecular evolutionary events such as substitutions, insertions and deletions from closely related genomes requires (1) an accurate multiple sequence alignment program and (2) a method to annotate the insertions and deletions that explain the 'gaps' in the alignment. Although the former requirement has been extensively addressed, the latter problem has received little attention, especially in a comprehensive probabilistic framework. RESULTS: Here, we present Indelign, a program that uses a probabilistic evolutionary model to compute the most likely scenario of insertions and deletions consistent with an input multiple alignment. It is also capable of modifying the given alignment so as to obtain a better agreement with the evolutionary model. We find close to optimal performance and substantial improvement over alternative methods, in tests of Indelign on synthetic data. We use Indelign to analyze regulatory sequences in Drosophila, and find an excess of insertions over deletions, which is different from what has been reported for neutral sequences. AVAILABILITY: The Indelign program may be downloaded from the website http://veda.cs.uiuc.edu/indelign/ SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.

Algorithms↗

Generation and analysis of large-scale expressed sequence tags (ESTs) from a full-length enriched cDNA library of porcine backfat tissue.

BACKGROUND: Genome research in farm animals will expand our basic knowledge of the genetic control of complex traits, and the results will be applied in the livestock industry to improve meat quality and productivity, as well as to reduce the incidence of disease. A combination of quantitative trait locus mapping and microarray analysis is a useful approach to reduce the overall effort needed to identify genes associated with quantitative traits of interest. RESULTS: We constructed a full-length enriched cDNA library from porcine backfat tissue. The estimated average size of the cDNA inserts was 1.7 kb, and the cDNA fullness ratio was 70%. In total, we deposited 16,110 high-quality sequences in the dbEST division of GenBank (accession numbers: DT319652-DT335761). For all the expressed sequence tags (ESTs), approximately 10.9 Mb of porcine sequence were generated with an average length of 674 bp per EST (range: 200-952 bp). Clustering and assembly of these ESTs resulted in a total of 5,008 unique sequences with 1,776 contigs (35.46%) and 3,232 singleton (65.54%) ESTs. From a total of 5,008 unique sequences, 3,154 (62.98%) were similar to other sequences, and 1,854 (37.02%) were identified as having no hit or low identity (<95%) and 60% coverage in The Institute for Genomic Research (TIGR) gene index of Sus scrofa. Gene ontology (GO) annotation of unique sequences showed that approximately 31.7, 32.3, and 30.8% were assigned molecular function, biological process, and cellular component GO terms, respectively. A total of 1,854 putative novel transcripts resulted after comparison and filtering with the TIGR SsGI; these included a large percentage of singletons (80.64%) and a small proportion of contigs (13.36%). CONCLUSION: The sequence data generated in this study will provide valuable information for studying expression profiles using EST-based microarrays and assist in the condensation of current pig TCs into clusters representing longer stretches of cDNA sequences. The isolation of genes expressed in backfat tissue is the first step toward a better understanding of backfat tissue on a genomic basis.

Adipose Tissue↗

SCORPION, a molecular database of scorpion toxins.

Increasing interest in the studies of toxins and the requirements for better structural and functional annotations have created a need for improved data management in the field of toxins. The molecular database, SCORPION, contains more than 200 entries of fully referenced scorpion toxin data including primary sequences, three-dimensional structures, structural and functional annotations of scorpion toxins along with relevant literature references. SCORPION has a set of search tools that allow users to extract data and perform specific queries. These entries have been compiled from public databases and literature, cleaned of errors and enriched with additional structural and functional information. The grouping of scorpion toxins provides a basis for extending and clarifying the existing structural and functional classifications. The bioinformatics modules in SCORPION facilitate analyses aimed at classification of scorpion toxins and identification of sequence patterns associated with specific structural or functional properties of scorpion toxins. The SCORPION database is accessible via the Internet at sdmc.krdl.org.sg:8080/scorpion.

Animals↗

Association of genes to genetically inherited diseases using data mining.

Although approximately one-quarter of the roughly 4,000 genetically inherited diseases currently recorded in respective databases (LocusLink, OMIM) are already linked to a region of the human genome, about 450 have no known associated gene. Finding disease-related genes requires laborious examination of hundreds of possible candidate genes (sometimes, these are not even annotated; see, for example, refs 3,4). The public availability of the human genome draft sequence has fostered new strategies to map molecular functional features of gene products to complex phenotypic descriptions, such as those of genetically inherited diseases. Owing to recent progress in the systematic annotation of genes using controlled vocabularies, we have developed a scoring system for the possible functional relationships of human genes to 455 genetically inherited diseases that have been mapped to chromosomal regions without assignment of a particular gene. In a benchmark of the system with 100 known disease-associated genes, the disease-associated gene was among the 8 best-scoring genes with a 25% chance, and among the best 30 genes with a 50% chance, showing that there is a relationship between the score of a gene and its likelihood of being associated with a particular disease. The scoring also indicates that for some diseases, the chance of identifying the underlying gene is higher.

Chromosome Mapping↗

Cloning and characterization of Enterobacter sakazakii pigment genes and in situ spectroscopic analysis of the pigment.

Enterobacter sakazakii is considered an opportunistic foodborne pathogen that is characterized by formation of yellow-pigmented colonies. Because of the lack of basic knowledge about Enterobacter sakazakii genetics, the BAC approach and the heterologous expression of the pigment in Escherichia coli were used to elucidate the molecular structure of the genes responsible for pigment production in Enterobacter sakazakii strain ES5. Sequencing and annotation of a 33.025 bp fragment revealed seven ORFs that could be assigned to the carotenoid biosynthesis pathway. The gene cluster had the organization crtE-idi-XYIBZ, with the crtE-idi-XYIB genes putatively transcribed as an operon and the crtZ gene transcribed in the opposite orientation. The carotenogenic nature of the pigment of Enterobacter sakazakii wt was ascertained by in situ analysis using visible microspectroscopy and resonance Raman microspectroscopy.

Carotenoids↗

MMDB: Entrez's 3D-structure database.

Three-dimensional structures are now known within many protein families and it is quite likely, in searching a sequence database, that one will encounter a homolog with known structure. The goal of Entrez's 3D-structure database is to make this information, and the functional annotation it can provide, easily accessible to molecular biologists. To this end Entrez's search engine provides three powerful features. (i) Sequence and structure neighbors; one may select all sequences similar to one of interest, for example, and link to any known 3D structures. (ii) Links between databases; one may search by term matching in MEDLINE, for example, and link to 3D structures reported in these articles. (iii) Sequence and structure visualization; identifying a homolog with known structure, one may view molecular-graphic and alignment displays, to infer approximate 3D structure. In this article we focus on two features of Entrez's Molecular Modeling Database (MMDB) not described previously: links from individual biopolymer chains within 3D structures to a systematic taxonomy of organisms represented in molecular databases, and links from individual chains (and compact 3D domains within them) to structure neighbors, other chains (and 3D domains) with similar 3D structure. MMDB may be accessed at http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Structure.

Animals↗

Origin and molecular evolution of receptor tyrosine kinases with immunoglobulin-like domains.

Receptor tyrosine kinases (RTKs) are involved in the control of fundamental cellular processes in metazoans. In vertebrates, RTK could be grouped in distinct classes based on the nature of their cognate ligand and modular composition of their extracellular domain. RTK with immunoglobulin-like domains (IG-like RTK) encompass several RTK classes and have been found in early metazoans, including sponges. Evolution of IG-like RTK is characterized by extended molecular and functional diversification, which prompted us to study their evolutionary history. For that purpose, a nonredundant data set including annotated protein sequences of IG-like RTK (n = 85) was built, representing 19 species ranging from sponges to humans. Phylogenetic trees were generated from alignment of conserved regions using maximum likelihood approach. Molecular phylogeny strongly suggests that IG-like RTK diversification occurred according to a complex scenario. In particular, we propose that specific cis duplications of a common ancestor to both platelet-derived growth factor receptor (class III) and vascular endothelial growth factor receptor (class V) families preceded two trans duplications. In contrast, other IG-like RTK genes, like Musk and PTK7, apparently did not evolve by duplications, whereas fibroblast growth factor receptors (class IV) evolved through two rounds of trans duplications. The proposed model of IG-like RTK evolution is supported by high bootstrap values and by the clustering of genes encoding class III and class V RTKs at specific chromosomal locations in mouse and human genomes.

Animals↗

Functional profiling of the Saccharomyces cerevisiae genome.

Determining the effect of gene deletion is a fundamental approach to understanding gene function. Conventional genetic screens exhibit biases, and genes contributing to a phenotype are often missed. We systematically constructed a nearly complete collection of gene-deletion mutants (96% of annotated open reading frames, or ORFs) of the yeast Saccharomyces cerevisiae. DNA sequences dubbed 'molecular bar codes' uniquely identify each strain, enabling their growth to be analysed in parallel and the fitness contribution of each gene to be quantitatively assessed by hybridization to high-density oligonucleotide arrays. We show that previously known and new genes are necessary for optimal growth under six well-studied conditions: high salt, sorbitol, galactose, pH 8, minimal medium and nystatin treatment. Less than 7% of genes that exhibit a significant increase in messenger RNA expression are also required for optimal growth in four of the tested conditions. Our results validate the yeast gene-deletion collection as a valuable resource for functional genomics.

Cell Size↗

Genome-wide bioinformatic and molecular analysis of introns in Saccharomyces cerevisiae.

Introns have typically been discovered in an ad hoc fashion: introns are found as a gene is characterized for other reasons. As complete eukaryotic genome sequences become available, better methods for predicting RNA processing signals in raw sequence will be necessary in order to discover genes and predict their expression. Here we present a catalog of 228 yeast introns, arrived at through a combination of bioinformatic and molecular analysis. Introns annotated in the Saccharomyces Genome Database (SGD) were evaluated, questionable introns were removed after failing a test for splicing in vivo, and known introns absent from the SGD annotation were added. A novel branchpoint sequence, AAUUAAC, was identified within an annotated intron that lacks a six-of-seven match to the highly conserved branchpoint consensus UACUAAC. Analysis of the database corroborates many conclusions about pre-mRNA substrate requirements for splicing derived from experimental studies, but indicates that splicing in yeast may not be as rigidly determined by splice-site conservation as had previously been thought. Using this database and a molecular technique that directly displays the lariat intron products of spliced transcripts (intron display), we suggest that the current set of 228 introns is still not complete, and that additional intron-containing genes remain to be discovered in yeast. The database can be accessed at http://www.cse.ucsc.edu/research/compbi o/yeast_introns.html.

Computational Biology↗

From PREDs and open reading frames to cDNA isolation: revisiting the human chromosome 21 transcription map.

A supernumerary copy of human chromosome 21 (HC21) causes Down syndrome. To understand the molecular pathogenesis of Down syndrome, it is necessary to identify all HC21 genes. The first annotation of the sequence of 21q confirmed 127 genes, and predicted an additional 98 previously unknown "anonymous" genes (predictions (PREDs) and open reading frames (C21orfs)), which were foreseen by exon prediction programs and/or spliced expressed sequence tags. These putative gene models still need to be confirmed as bona fide transcripts. Here we report the characterization and expression pattern of the putative transcripts C21orf7, C21orf11, C21orf15, C21orf18, C21orf19, C21orf22, C21orf42, C21orf50, C21orf51, C21orf57, and C21orf58, the GC-rich sequence DNA-binding factor candidate GCFC (also known as C21orf66), PRED12, PRED31, PRED34, PRED44, PRED54, and PRED56. Our analysis showed that most of the C21orfs originally defined by matching spliced expressed sequence tags were correctly predicted, whereas many of the PREDs, defined solely by computer prediction, do not correspond to genuine genes. Four of the six PREDs were incorrectly predicted: PRED44 and C21orf11 are portions of the same transcript, PRED31 is a pseudogene, and PRED54 and PRED56 were wrongly predicted. In contrast, PRED12 (now called C21orf68) and PRED34 (C21orf63) are now confirmed transcripts. We identified three new genes, C21orf67, C21orf69, and C21orf70, not previously predicted by any programs. This revision of the HC21 transcriptome has consequences for the entire genome regarding the quality of previous annotations and the total number of transcripts. It also provides new candidates for genes involved in Down syndrome and other genetic disorders that map to HC21.

Animals↗

Molecular identification of the first insect ecdysis triggering hormone receptors.

The Drosophila Genome Project website (www.flybase.org) contains an annotated gene sequence (CG5911), coding for a G protein-coupled receptor. We cloned the cDNA corresponding to this sequence and found that the gene has not been correctly predicted. The corrected gene CG5911 has five introns and six exons (1-6). Alternative splicing yields two cDNAs called A (containing exons 1-5) and B (containing exons 1-4, 6). We expressed these splicing variants in Chinese hamster ovary cells and found that the corrected CG5911-A and -B cDNAs coded for two different G protein-coupled receptors that could be activated by low concentrations of Drosophila ecdysis triggering hormones-1 and -2. Ecdysis (cuticle shedding) is an important behaviour, allowing growth and metamorphosis in insects and other arthropods. Our paper is the first report on the molecular identification of ecdysis triggering hormone receptors from insects.

Alternative Splicing↗

Genome Report: De novo genome assembly of the greater Bermuda land snail, Poecilozonites bermudensis (Mollusca: Gastropoda), confirms ancestral genome duplication.

Poecilozonites bermudensis, the greater Bermuda land snail, is a critically endangered species and one of only two extant members in its genus. These snails are one of Bermuda's few endemic animal clades and their rich fossil record was the basis for the punctuated equilibria model of speciation. Once thought extinct, recent conservation efforts have focused on the recovery of the species, yet no genomic information or other molecular sequences have been available to inform these initiatives. We present a high-quality, annotated genome for P. bermudensis generated using PacBio long read and Omni-C short read sequencing. The resulting assembly is approximately 1.36 Gb with a scaffold N50 of 44.t Mb and 31 chromosome-length scaffolds. Nearly 43 percent of the genome was identified as repeat content. This assembly will serve as a resource for the conservation and study of P. bermudensis, and its only close extant and also critically endangered relative, P. circumfirmatus. Additionally, this genome adds to the growing body of data needed for a more complete understanding of gastropod evolution and for evolutionary processes in general.

Annotation↗

GOPET: a tool for automated predictions of Gene Ontology terms.

BACKGROUND: Vast progress in sequencing projects has called for annotation on a large scale. A Number of methods have been developed to address this challenging task. These methods, however, either apply to specific subsets, or their predictions are not formalised, or they do not provide precise confidence values for their predictions. DESCRIPTION: We recently established a learning system for automated annotation, trained with a broad variety of different organisms to predict the standardised annotation terms from Gene Ontology (GO). Now, this method has been made available to the public via our web-service GOPET (Gene Ontology term Prediction and Evaluation Tool). It supplies annotation for sequences of any organism. For each predicted term an appropriate confidence value is provided. The basic method had been developed for predicting molecular function GO-terms. It is now expanded to predict biological process terms. This web service is available via http://genius.embnet.dkfz-heidelberg.de/menu/biounit/open-husar CONCLUSION: Our web service gives experimental researchers as well as the bioinformatics community a valuable sequence annotation device. Additionally, GOPET also provides less significant annotation data which may serve as an extended discovery platform for the user.

Artificial Intelligence↗

T1DBase, a community web-based resource for type 1 diabetes research.

T1DBase (http://T1DBase.org) is a public website and database that supports the type 1 diabetes (T1D) research community. The site is currently focused on the molecular genetics and biology of T1D susceptibility and pathogenesis. It includes the following datasets: annotated genome sequence for human, rat and mouse; information on genetically identified T1D susceptibility regions in human, rat and mouse, and genetic linkage and association studies pertaining to T1D; descriptions of NOD mouse congenic strains; the Beta Cell Gene Expression Bank, which reports expression levels of genes in beta cells under various conditions, and annotations of gene function in beta cells; data on gene expression in a variety of tissues and organs; and biological pathways from KEGG and BioCarta. Tools on the site include the GBrowse genome browser, site-wide context dependent search, Connect-the-Dots for connecting gene and other identifiers from multiple data sources, Cytoscape for visualizing and analyzing biological networks, and the GESTALT workbench for genome annotation. All data are open access and all software is open source.

Animals↗

Phylogenetic analysis of general bacterial porins: a phylogenomic case study.

Bacterial porin proteins allow for the selective movement of hydrophilic solutes through the outer membrane of Gram-negative bacteria. The purpose of this study was to clarify the evolutionary relationships among the Type 1 general bacterial porins (GBPs), a porin protein subfamily that includes outer membrane proteins ompC and ompF among others. Specifically, we investigated the potential utility of phylogenetic analysis for refining poorly annotated or mis-annotated protein sequences in databases, and for characterizing new functionally distinct groups of porin proteins. Preliminary phylogenetic analysis of sequences obtained from GenBank indicated that many of these sequences were incompletely or even incorrectly annotated. Using a well-curated set of porins classified via comparative genomics, we applied recently developed bayesian phylogenetic methods for protein sequence analysis to determine the relationships among the Type 1 GBPs. Our analysis found that the major GBP classes (ompC, phoE, nmpC and ompN) formed strongly supported monophyletic groups, with the exception of ompF, which split into two distinct clades. The relationships of the GBP groups to one another had less statistical support, except for the relationships of ompC and ompN sequences, which were strongly supported as sister groups. A phylogenetic analysis comparing the relationships of the GenBank GBP sequences to the correctly annotated set of GBPs identified a large number of previously unclassified and mis-annotated GBPs. Given these promising results, we developed a tree-parsing algorithm for automated phylogenetic annotation and tested it with GenBank sequences. Our algorithm was able to automatically classify 30 unidentified and 15 mis-annotated GBPs out of 78 sequences. Altogether, our results support the potential for phylogenomics to increase the accuracy of sequence annotations.

Algorithms↗

Molecular identification of the first insect proctolin receptor.

The website of the Drosophila Genome Project (www.flybase.org) contains the sequence of an annotated gene CG6986, which is predicted to code for a G protein-coupled receptor. We cloned the cDNA of this gene and expressed it in Chinese hamster ovary cells. Screening of a neuropeptide library revealed that the expressed receptor was specific for the neuropeptide proctolin (EC(50), 6x10(-10)M). Proctolin (RYLPT) was the first invertebrate neuropeptide to be fully sequenced (already in 1975) and occurs with identical structure in both crustaceans and insects, where it has myo- and neurostimulatory actions. Northern blots showed that the Drosophila proctolin receptor was only weakly expressed in embryos, larvae, pupae, and in the thoraces and abdomina of adult flies, but strongly in the heads of adult animals. The Drosophila receptor reported here is the first invertebrate proctolin receptor to be identified.

Amino Acid Sequence↗

Annotated expressed sequence tags for studies of the regulation of reproductive modes in aphids.

The damaging effect of aphids to crops is largely determined by the spectacular rate of increase of populational expansion due to their parthenogenetic generations. Despite this, the molecular processes triggering the transition between the parthenogenetic and sexual phases between their annual life cycle have received little attention. Here, we describe a collection of genes from the cereal aphid Rhopalosiphum padi expressed during the switch from parthenogenetic to sexual reproduction. After cDNA cloning and sequencing, 726 expressed sequence tags (EST) were annotated. The R. padi EST collection contained a substantial number (139) of bacterial endosymbiont sequences. The majority of R. padi cDNAs encoded either unknown proteins (56%) or housekeeping polypeptides (38%). The large proportion of sequences without similarities in the databases is related to both their small size and their high GC content, corresponding probably to the presence of 5'-unstranslated regions. Fifteen genes involved in developmental and differentiation events were identified by similarity to known genes. Some of these may be useful candidates for markers of the early steps of sexual differentiation.

Amino Acid Sequence↗