PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Molecular Sequence Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Annotated expressed sequence tags for studies of the regulation of reproductive modes in aphids.

The damaging effect of aphids to crops is largely determined by the spectacular rate of increase of populational expansion due to their parthenogenetic generations. Despite this, the molecular processes triggering the transition between the parthenogenetic and sexual phases between their annual life cycle have received little attention. Here, we describe a collection of genes from the cereal aphid Rhopalosiphum padi expressed during the switch from parthenogenetic to sexual reproduction. After cDNA cloning and sequencing, 726 expressed sequence tags (EST) were annotated. The R. padi EST collection contained a substantial number (139) of bacterial endosymbiont sequences. The majority of R. padi cDNAs encoded either unknown proteins (56%) or housekeeping polypeptides (38%). The large proportion of sequences without similarities in the databases is related to both their small size and their high GC content, corresponding probably to the presence of 5'-unstranslated regions. Fifteen genes involved in developmental and differentiation events were identified by similarity to known genes. Some of these may be useful candidates for markers of the early steps of sexual differentiation.

Amino Acid Sequence↗

Improving interoperability between microbial information and sequence databases.

BACKGROUND: Biological resources are essential tools for biomedical research. Their availability is promoted through on-line catalogues. Common Access to Biological Resources and Information (CABRI) is a service for distribution of biological resources and related data collected by 28 European culture collections. Linking this information to bioinformatics databanks can make the collections' holdings more visible after a search in molecular biology databanks and vice-versa. Identification of links to sequence databases can be useful, but annotation and indexing problems, together with compilation errors, immediately arise. In this paper, we present our efforts for the identification of cross-references between CABRI catalogues and the EMBL Data Library and related results. RESULTS: An SRS site with both EMBL and CABRI catalogues has been set up. Ad-hoc changes in indexing scripts allowed to achieve homogeneous index keys and SRS link features have been used to identify links between databases. After manual checking and comparison with an alternative procedure, about 67,500 valid cross-references were identified, added to the EMBL Data Library and are now distributed with it. HTML links can be established from EMBL to CABRI network service. Procedures can be executed whenever needed. CONCLUSION: Links between EMBL and CABRI catalogues constitute an improved access to micro-organisms of certified quality and can produce positive effects on biomedical research. Further links between CABRI catalogues and other bioinformatics databases can now easily be defined by using these cross-references. Linking genetic information onto natural resources information may stand model for the integration of other databases containing empirical data on these materials.

Base Sequence↗

Tools for integrated sequence-structure analysis with UCSF Chimera.

BACKGROUND: Comparing related structures and viewing the structures in the context of sequence alignments are important tasks in protein structure-function research. While many programs exist for individual aspects of such work, there is a need for interactive visualization tools that: (a) provide a deep integration of sequence and structure, far beyond mapping where a sequence region falls in the structure and vice versa; (b) facilitate changing data of one type based on the other (for example, using only sequence-conserved residues to match structures, or adjusting a sequence alignment based on spatial fit); (c) can be used with a researcher's own data, including arbitrary sequence alignments and annotations, closely or distantly related sets of proteins, etc.; and (d) interoperate with each other and with a full complement of molecular graphics features. We describe enhancements to UCSF Chimera to achieve these goals. RESULTS: The molecular graphics program UCSF Chimera includes a suite of tools for interactive analyses of sequences and structures. Structures automatically associate with sequences in imported alignments, allowing many kinds of crosstalk. A novel method is provided to superimpose structures in the absence of a pre-existing sequence alignment. The method uses both sequence and secondary structure, and can match even structures with very low sequence identity. Another tool constructs structure-based sequence alignments from superpositions of two or more proteins. Chimera is designed to be extensible, and mechanisms for incorporating user-specific data without Chimera code development are also provided. CONCLUSION: The tools described here apply to many problems involving comparison and analysis of protein structures and their sequences. Chimera includes complete documentation and is intended for use by a wide range of scientists, not just those in the computational disciplines. UCSF Chimera is free for non-commercial use and is available for Microsoft Windows, Apple Mac OS X, Linux, and other platforms from http://www.cgl.ucsf.edu/chimera.

Computer Graphics↗

Genetic differentiation of the Aspergillus section Flavi complex using AFLP fingerprints.

Twenty-four isolates of Aspergillus sojae, A. parasiticus, A. oryzae and A. flavus, including a number that have the capacity to produce aflatoxin, have been compared using amplified fragment length polymorphisms (AFLPs). Based on analysis of 12 different primer combinations, 500 potentially polymorphic fragments have been identified. Analysis of the AFLP data consistently and clearly separates the A. sojae/A. parasiticus isolates from the A. oryzae/A. flavus isolates. Furthermore. there are markers that can be used to distinguish the A. sojae isolates from those of A. parasiticus, which form the basis for species-specific markers. However, whilst there were many polymorphisms between isolates within the A. oryzae/A. flavus subgroup, no markers could be identified that distinguish between the two species. Sequencing of the ribosomal DNA ITS (internal transcribed spacers) from selected isolates also separated the A. sojae/A. parasiticus subgroup from the A. oryzae/A. flavus subgroup, but was unable to distinguish between the A. sojae and A. parasiticus isolates. Some ITS variation was found between isolates within the A. oryzae/A. flavus subgroup, but did not correlate with the species classification, indicating that it is difficult to use molecular data to separate the two species. In addition, sequencing of ribosomal ITS regions and AFLP analysis suggested that some species annotations in public culture collections may be inaccurate.

Aflatoxins↗

Yeast Protein database (YPD): a database for the complete proteome of Saccharomyces cerevisiae.

The Yeast Protein Database (YPD) is a database for the proteins of the budding yeast,Saccharomyces cerevisiae. YPD is the first annotated database for the complete proteome of any organism. Now that the complete genome sequence of yeast is available, YPD contains entries for each of the characterized proteins and for each of the uncharacterized proteins predicted from the sequence. Contained in YPD are the calculated properties of each protein such as molecular weight and isoelectric point, experimentally determined properties such as subcellular localization and post-translational modifications, and extensive annotations from the yeast literature. YPD contains 25 000 lines of textual annotation that describe the known functions, mutant phenotypes, interactions, and other properties for the approximately 6000 proteins in the yeast proteome. The information in YPD is updated daily, and it is available on the World Wide Web at http://www.proteome.com/YPDhome.html .

Amino Acid Sequence↗

Text-based analysis of genes, proteins, aging, and cancer.

The diverse nature of cancer- and aging-related genes presents a challenge for large-scale studies based on molecular sequence and profiling data. An underexplored source of data for modeling and analysis is the textual descriptions and annotations present in curated gene-centered biomedical corpora. Here, 450 genes designated by surveys of the scientific literature as being associated with cancer and aging were analyzed using two complementary approaches. The first, ensemble attribute profile clustering, is a recently formulated, text-based, semi-automated data interpretation strategy that exploits ideas from statistical information retrieval to discover and characterize groups of genes with common structural and functional properties. Groups of genes with shared and unique Gene Ontology terms and protein domains were defined and examined. Human homologs of a group of known Drosphila aging-related genes are candidates for genes that may influence lifespan (hep/MAPK2K7, bsk/MAPK8, puc/LOC285193). These JNK pathway-associated proteins may specify a molecular hub that coordinates and integrates multiple intra- and extracellular processes via space- and time-dependent interactions with proteins in other pathways. The second approach, a qualitative examination of the chromosomal locations of 311 human cancer- and aging-related genes, provides anecdotal evidence for a "phenotype position effect": genes that are proximal in the linear genome often encode proteins involved in the same phenomenon. Comparative genomics was employed to enhance understanding of several genes, including open reading frames, identified as new candidates for genes with roles in aging or cancer. Overall, the results highlight fundamental molecular and mechanistic connections between progenitor/stem cell lineage determination, embryonic morphogenesis, cancer, and aging. Despite diversity in the nature of the molecular and cellular processes associated with these phenomena, they seem related to the architectural hub of tissue polarity and a need to generate and control this property in a timely manner.

Aging↗

A comprehensive BAC resource.

The Human Genome Project has generated extensive map and sequence data for a large number of Bacterial Artificial Chromosome (BAC) clones. In order to maximize the efficient use of the data and to minimize the redundant work for the research community, The Institute for Genomic Research (TIGR) comprehensive BAC resource (cBACr) (http://www.tigr.org/tdb/BacResource/BAC_resourc e_intro. html) was built as an expansion of the TIGR human BAC ends database. This resource collects, integrates and reports the information on library, maps, sequence, annotation and functions for each human and mouse BAC. The current database contains 635 016 human BACs and 265 617 mouse BACs that were characterized by various approaches, among which 22 705 human clones and 1000 mouse clones have sequence and annotation data.

Animals↗

Transcription and histone modifications in the recombination-free region spanning a rice centromere.

Centromeres are sites of spindle attachment for chromosome segregation. During meiosis, recombination is absent at centromeres and surrounding regions. To understand the molecular basis for recombination suppression, we have comprehensively annotated the 3.5-Mb region that spans a fully sequenced rice centromere. Although transcriptional analysis showed that the 750-kb CENH3-containing core is relatively deficient in genes, the recombination-free region differs little in gene density from flanking regions that recombine. Likewise, the density of transposable elements is similar between the recombination-free region and flanking regions. We also measured levels of histone H4 acetylation and histone H3 methylation at 176 genes within the 3.5-Mb span. Active genes showed enrichment of H4 acetylation and H3K4 dimethylation as expected, including genes within the core. Our inability to detect sequence or histone modification features that distinguish recombination-free regions from flanking regions that recombine suggest that recombination suppression is an epigenetic feature of centromeres maintained by the assembly of CENH3-containing nucleosomes within the core. CENH3-containing centrochromatin does not appear to be distinguished by a unique combination of H3 and H4 modifications. Rather, the varied distribution of histone modifications might reflect the composition and abundance of sequence elements that inhabit centromeric DNA.

Centromere↗

A genomewide survey of developmentally relevant genes in Ciona intestinalis. IX. Genes for muscle structural proteins.

Ascidians are simple chordates that are related to, and may resemble, vertebrate ancestors. Comparison of ascidian and vertebrate genomes is expected to provide insight into the molecular genetic basis of chordate/vertebrate evolution. We annotated muscle structural (contractile protein) genes in the completely determined genome sequence of the ascidian Ciona intestinalis, and examined gene expression patterns through extensive EST analysis. Ascidian muscle protein isoform families are generally of similar, or lesser, complexity in comparison with the corresponding vertebrate isoform families, and are based on gene duplication histories and alternative splicing mechanisms that are largely or entirely distinct from those responsible for generating the vertebrate isoforms. Although each of the three ascidian muscle types - larval tail muscle, adult body-wall muscle and heart - expresses a distinct profile of contractile protein isoforms, none of these isoforms are strictly orthologous to the smooth-muscle-specific, fast or slow skeletal muscle-specific, or heart-specific isoforms of vertebrates. Many isoform families showed larval-versus-adult differential expression and in several cases numerous very similar genes were expressed specifically in larval muscle. This may reflect different functional requirements of the locomotor larval muscle as opposed to the non-locomotor muscles of the sessile adult, and/or the biosynthetic demands of extremely rapid larval development.

Amino Acid Sequence↗

Pseudogenes in metazoa: origin and features.

The complete genome sequences with their annotations are a considerable resource in biology, particularly in understanding the global structure of the genetic material at the molecular level. The reason why some eukaryotic genomes contain large quantities of apparently unnecessary DNA, namely pseudogenes, while others seem to invest in more efficient thinning processes or are equipped with protection systems against parasitic elements still remains a mystery. Several genome-wide surveys have been undertaken to identify pseudogenes in the completely sequenced genome, bringing to light some differences both in their amount and distribution. Since pseudogenes are important resources in evolutionary and comparative genomics - as 'molecular fossils' - in this paper, a survey on the origins, features, abundance and localisation of the different pseudogenes is reported. As an example of genes producing processed pseudogenes, some experimental data obtained in the authors' laboratories from the study of a nuclear gene coding for the mitochondrial transcription factor A (mtTFA), a key regulator of mitochondrial biogenesis, are also reported.

Animals↗

Gene ontology application to genomic functional annotation, statistical analysis and knowledge mining.

While a massive amount of biomolecular information is increasingly accumulating in different databanks, on the other hand high-throughput technologies are generating a great quantity of data that need to be annotated with the genomic information available, and interpreted. To this aim, the use of specific ontologies can greatly help either in integrating different information stored within heterogeneous databanks, or in identifying and clustering sequence data sharing common characteristics. In the molecular biology domain, the Gene Ontology (GO) is the most developed and widely used ontology. To demonstrate its great utility in the annotation and biological interpretation of gene sets obtained by means of high-throughput experiments, we implemented the web application here described. It enables functional annotations of a given gene set on a genomic scale and across different species. Within our application the annotations provided by the GO vocabulary allow either to easily bind several information from different resources, or to cluster annotated genes according to their biological characteristics. Through the GO structure it is also possible to represent biological concepts with different specificity levels, from very general to very precise concepts. Furthermore, the statistical evaluation of the categorizations provided by the GO annotations enables to highlight the most significant biological characteristics of a gene set, and therefore to mine knowledge from data. Our created tool meets the need to manage a vast quantity of biological data with a simple user interface adapt also for users with limited informatics knowledge, leading them to evaluate the functional significance of experiment's results with graphical views and statistical indexes in a well-known web browser user interface.

Genomics↗

The Gene Ontology (GO) database and informatics resource.

The Gene Ontology (GO) project (http://www. geneontology.org/) provides structured, controlled vocabularies and classifications that cover several domains of molecular and cellular biology and are freely available for community use in the annotation of genes, gene products and sequences. Many model organism databases and genome annotation groups use the GO and contribute their annotation sets to the GO resource. The GO database integrates the vocabularies and contributed annotations and provides full access to this information in several formats. Members of the GO Consortium continually work collectively, involving outside experts as needed, to expand and update the GO vocabularies. The GO Web resource also provides access to extensive documentation about the GO project and links to applications that use GO data for functional analyses.

Animals↗

Aphid biology: expressed genes from alate Toxoptera citricida, the brown citrus aphid.

The brown citrus aphid, Toxoptera citricida (Kirkaldy), is considered the primary vector of citrus tristeza virus, a severe pathogen which causes losses to citrus industries worldwide. The alate (winged) form of this aphid can readily fly long distances with the wind, thus spreading citrus tristeza virus in citrus growing regions. To better understand the biology of the brown citrus aphid and the emergence of genes expressed during wing development, we undertook a large-scale 5' end sequencing project of cDNA clones from alate aphids. Similar large-scale expressed sequence tag (EST) sequencing projects from other insects have provided a vehicle for answering biological questions relating to development and physiology. Although there is a growing database in GenBank of ESTs from insects, most are from Drosophila melanogaster and Anopheles gambiae, with relatively few specifically derived from aphids. However, important morphogenetic processes are exclusively associated with piercing-sucking insect development and sap feeding insect metabolism. In this paper, we describe the first public data set of ESTs from the brown citrus aphid, T. citricida. The cDNA library was derived from alate adults due to their significance in spreading viruses (e.g., citrus tristeza virus). Over 5180 cDNA clones were sequenced, resulting in 4263 high-quality ESTs. Contig alignment of these ESTs resulted in 2124 total assembled sequences, including both contiguous sequences and singlets. Approximately 33% of the ESTs currently have no significant match in either the non-redundant protein or nucleic acid databases. Sequences returning matches with an E-value of < or = -10 using BLASTX, BLASTN, or TBLASTX were annotated based on their putative molecular function and biological process using the Gene Ontology classification system. These data will aid research efforts in the identification of important genes within insects, specifically aphids and other sap feeding insects within the Order Hemiptera.

Animals↗

Phytome: a platform for plant comparative genomics.

Phytome is an online comparative genomics resource that can be applied to functional plant genomics, molecular breeding and evolutionary studies. It contains predicted protein sequences, protein family assignments, multiple sequence alignments, phylogenies and functional annotations for proteins from a large, phylogenetically diverse set of plant taxa. Phytome serves as a glue between disparate plant gene databases both by identifying the evolutionary relationships among orthologous and paralogous protein sequences from different species and by enabling cross-references between different versions of the same gene curated independently by different database groups. The web interface enables sophisticated queries on lineage-specific patterns of gene/protein family proliferation and loss. This rich dataset is serving as a platform for the unification of sequence-anchored comparative maps across taxonomic families of plants. The Phytome web interface can be accessed at the following URL: http://www.phytome.org. Batch homology searches and bulk downloads are available upon free registration.

Databases, Genetic↗

Comparative plant genomics resources at PlantGDB.

PlantGDB (http://www.plantgdb.org/) is a database of plant molecular sequences. Expressed sequence tag (EST) sequences are assembled into contigs that represent tentative unique genes. EST contigs are functionally annotated with information derived from known protein sequences that are highly similar to the putative translation products. Tentative Gene Ontology terms are assigned to match those of the similar sequences identified. Genome survey sequences are assembled similarly. The resulting genome survey sequence contigs are matched to ESTs and conserved protein homologs to identify putative full-length open reading frame-containing genes, which are subsequently provisionally classified according to established gene family designations. For Arabidopsis (Arabidopsis thaliana) and rice (Oryza sativa), the exon-intron boundaries for gene structures are annotated by spliced alignment of ESTs and full-length cDNAs to their respective complete genome sequences. Unique genome browsers have been developed to present all available EST and cDNA evidence for current transcript models (for Arabidopsis, see the AtGDB site at http://www.plantgdb.org/AtGDB/; for rice, see the OsGDB site at http://www.plantgdb.org/OsGDB/). In addition, a number of bioinformatic tools have been integrated at PlantGDB that enable researchers to carry out sequence analyses on-site using both their own data and data residing within the database.

Computational Biology↗

Strategies for the identification, the assembly and the classification of integrated biological systems in completely sequenced genomes.

The proteins involved in a single biological process may form a stable supra-molecular assembly or be transiently in interaction. Although, the first annotation steps of a complete genome may allow the identification of the different partners, their assembly in a functional system, referred to as an integrated system, is a domain where methodological effort has to be done. Indeed, the knowledge required to assemble partners of such systems should be explicitly included in annotation software. The availability of a complete genome, and therefore of all the proteins encoded by that genome, motivated the development of automated approaches through the coordinated combination of different bio-informatic methods allowing the identification of the different partners, their assembly and the classification of the reconstructed systems in functional categories. In this data flux, the identification of the sequence partners represents the principal bottleneck. Here, we describe and compare the results obtained with different classes of methods (BLASTP2, PSI-BLAST, MAST and META-MEME) applied to the identification in complete genomes of a given family of integrated systems: the ABC transporters. PSI-BLAST appears to significantly outperform motif-based methods, and the results are discussed according to the nature of the proteins and the structure of the sub-families.

Binding Sites↗

Deduction of functional peptide motifs in scorpion toxins.

Scorpion toxins are important physiological probes for characterizing ion channels. Molecular databases have limited functional annotation of scorpion toxins. Their function can be inferred by searching for conserved motifs in sequence signature databases that are derived statistically but are not necessarily biologically relevant. Mutation studies provide biological information on residues and positions important for structure-function relationship but are not normally used for extraction of binding motifs. 3D structure analyses also aid in the extraction of peptide motifs in which non-contiguous residues are clustered spatially. Here we present new, functionally relevant peptide motifs for ion channels, derived from the analyses of scorpion toxin native and mutant peptides.

Amino Acid Motifs↗

Metabolism and genetics of Helicobacter pylori: the genome era.

The publication of the complete sequence of Helicobacter pylori 26695 in 1997 and more recently that of strain J99 has provided new insight into the biology of this organism. In this review, we attempt to analyze and interpret the information provided by sequence annotations and to compare these data with those provided by experimental analyses. After a brief description of the general features of the genomes of the two sequenced strains, the principal metabolic pathways are analyzed. In particular, the enzymes encoded by H. pylori involved in fermentative and oxidative metabolism, lipopolysaccharide biosynthesis, nucleotide biosynthesis, aerobic and anaerobic respiration, and iron and nitrogen assimilation are described, and the areas of controversy between the experimental data and those provided by the sequence annotation are discussed. The role of urease, particularly in pH homeostasis, and other specialized mechanisms developed by the bacterium to maintain its internal pH are also considered. The replicational, transcriptional, and translational apparatuses are reviewed, as is the regulatory network. The numerous findings on the metabolism of the bacteria and the paucity of gene expression regulation systems are indicative of the high level of adaptation to the human gastric environment. Arguments in favor of the diversity of H. pylori and molecular data reflecting possible mechanisms involved in this diversity are presented. Finally, we compare the numerous experimental data on the colonization factors and those provided from the genome sequence annotation, in particular for genes involved in motility and adherence of the bacterium to the gastric tissue.

Gene Expression Regulation, Bacterial↗