PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Molecular Sequence Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

[DNA arrays: technological aspects and applications].

The Human Genome Project has allowed considerable progress in the construction of physical and genetic maps and the identification of genes involved in human sicknesses. The accelerated accumulation of biological information and knowledge is due in large part to the sequencing projects of other organisms, which in fact paved the way for the Human Genome Project. In parallel, recently developed techniques which take advantage of genomic sequences allow large scale molecular analyses resulting in the functional annotation of many of the proteins represented by these genes. This is the goal of functional genomics. These progresses are at the origin of the present revolution in biomedical research. DNA microarrays are playing a dominant role compared to the other developing technologies since they are relatively easy to make and use and are applicable to numerous scientific inquiries. They allow the simultaneous analysis of several thousands of genes in biological samples from sick or healthy tissues, at the genome or transcriptome level. The data obtained is expected to result in major advances in the health sciences. In addition to an improved understanding of the complex molecular interaction networks of healthy cells and tissues, a more precise genetic characterization of the molecular mechanisms involved in pathology should result in the identification of new therapeutic targets and the development of new medicines. The genetic profiles thus obtained should also permit the definition of new pathologic subclasses not recognizable by traditional clinical factors, as well as new markers for susceptibility to certain illnesses, and new prognostic markers or methods of predicting responses to treatment. In this article, we present the different approaches and potential applications of DNA microarray technology, in particular as applied to cancer research.

Chromosome Mapping↗

The role of protein structure in genomics.

The genome projects produce an enormous amount of sequence data that needs to be annotated in terms of molecular structure and biological function. These tasks have triggered additional initiatives like structural genomics. The intention is to determine as many protein structures as possible, in the most efficient way, and to exploit the solved structures for the assignment of biological function to hypothetical proteins. We discuss the impact of these developments on protein classification, gene function prediction, and protein structure prediction.

Databases, Factual↗

Motif identification neural design for rapid and sensitive protein family search.

The accelerated growth of the molecular sequencing data has generated a pressing need for advanced sequence annotation tools. This paper reports a new method, termed MOTIFIND (Motif Identification Neural Design), for rapid and sensitive protein family identification. The method is extended from our previous gene classification artificial neural system and employs two new designs to enhance the detection of distant relationships. These include an n-gram term weighting algorithm for extracting local motif patterns, and integrated neural networks for combining global and local sequence information. The system has been tested with three protein families of electron transferases, namely cytochrome c, cytochrome b and flavodoxin, with a 100% sensitivity and more than 99.6% specificity. The accuracy of MOTIFIND is comparable to the BLAST database search method, but its speed is more than 20 times faster. The system is much more robust than the PROSITE search which is based on simple signature patterns. MOTIFIND also compares favorably with the BLIMPS search of BLOCKS in detecting fragmentary sequences lacking complete motif regions. The method has the potential to become a full-scale database search and sequence analysis tool.

Algorithms↗

Experimental analysis of the annotation of promoters in the public database.

The ability to identify and examine promoter elements is important to researchers who wish to understand how gene expression is regulated in normal and pathological states. Unfortunately, the number of human promoters that have been directly experimentally defined is small. In order to determine if promoter sequences can be identified by simply aligning mRNA and genomic sequences, we have used a reporter gene assay to assess the promoter activity of the immediate 5' region flanking 38 mRNAs mapping to chromosome 21. For comparison, we have measured the activities of 19 sequences not thought to be promoters and 39 sequences taken from the Eukaryotic Promoter Database. Our results suggest that alignment of reference mRNAs to genomic sequence allows promoters to be identified for at least 75% of genes. These data provide the first empirical evidence that the current state of annotation of the genome is sufficient to allow molecular geneticists to correctly identify promoter sequences for most genes for which reference mRNA and genomic sequences are available.

Cell Line↗

WILMA-automated annotation of protein sequences.

Large-scale annotation of sets of proteins is a frequently occurring task in association with genome sequencing projects. Here, we present an automated platform for the functional annotation of large sets of protein sequences. Various bioinformatics tools are used to achieve a comprehensive description of protein sequences and to link these results to standard Gene Ontology descriptors for molecular function, biological processes and cellular components. Access to the annotation is provided via a web-interface and database queries. These interfaces allow to formulate proteome wide queries as well as the investigation of details of individual results. WILMA annotations of the proteomes of Homo sapiens, Mus musculus, Arabidopsis thaliana and Caenorhabditis elegans are accessible at http://www.came.sbg.ac.at/wilma/

Amino Acid Sequence↗

Molecular weight assessment of proteins in total proteome profiles using 1D-PAGE and LC/MS/MS.

BACKGROUND: The observed molecular weight of a protein on a 1D polyacrylamide gel can provide meaningful insight into its biological function. Differences between a protein's observed molecular weight and that predicted by its full length amino acid sequence can be the result of different types of post-translational events, such as alternative splicing (AS), endoproteolytic processing (EPP), and post-translational modifications (PTMs). The characterization of these events is one of the important goals of total proteome profiling (TPP). LC/MS/MS has emerged as one of the primary tools for TPP, but since this method identifies tryptic fragments of proteins, it has not generally been used for large-scale determination of the molecular weight of intact proteins in complex mixtures. RESULTS: We have developed a set of computational tools for extracting molecular weight information of intact proteins from total proteome profiles in a high throughput manner using 1D-PAGE and LC/MS/MS. We have applied this technology to the proteome profile of a human lymphoblastoid cell line under standard culture conditions. From a total of 1 x 10(7) cells, we identified 821 proteins by at least two tryptic peptides. Additionally, these 821 proteins are well-localized on the 1D-SDS gel. 656 proteins (80%) occur in gel slices in which the observed molecular weight of the protein is consistent with its predicted full-length sequence. A total of 165 proteins (20%) are observed to have molecular weights that differ from their predicted full-length sequence. We explore these molecular-weight differences based on existing protein annotation. CONCLUSION: We demonstrate that the determination of intact protein molecular weight can be achieved in a high-throughput manner using 1D-PAGE and LC/MS/MS. The ability to determine the molecular weight of intact proteins represents a further step in our ability to characterize gene expression at the protein level. The identification of 165 proteins whose observed molecular weight differs from the molecular weight of the predicted full-length sequence provides another entry point into the high-throughput characterization of protein modification.

Journal Article↗

Insights into the genetic organization of the Corynebacterium diphtheriae erythromycin resistance plasmid pNG2 deduced from its complete nucleotide sequence.

The complete nucleotide sequence of the erythromycin resistance plasmid pNG2 from the human pathogen Corynebacterium diphtheriae S601 was determined. The plasmid has a total size of 15,100 bp and contains at least 17 coding regions. Comparative genomics identified conserved motifs within replication initiator proteins of corynebacterial plasmids and a novel nucleotide sequence feature, termed 22-bp box, located downstream of the repA gene. The erythromycin resistance determinant erm(X) is flanked by inverted repeats of the novel insertion sequence IS3504, which may be responsible for a spontaneous deletion of the antibiotic resistance gene region. Furthermore, pNG2 encodes a putative conjugative relaxase, a membrane protein of the natural resistance-associated macrophage protein (Nramp) family and a protein with Nudix hydrolase signature. Expression of the predicted coding regions of pNG2 in Escherichia coli JM109 was demonstrated by reverse transcription-polymerase chain reaction (RT-PCR) assays. The detailed annotation of the entire pNG2 sequence provided genetic information regarding its molecular evolution and its role in dissemination of antibiotic resistance genes by horizontal gene transfer.

Amino Acid Sequence↗

Computationally efficient cluster representation in molecular sequence megaclassification.

Molecular sequence megaclassification is a technique for automated protein sequence analysis and annotation. Implementation of the method has been limited by the need to store and randomly access a database of all the sequence pair similarities. More than 80,000 protein sequences are now present in the public databases, and the pair similarity data table for the full protein sequence database requires over 1 gigabyte of storage. In this paper we present a computationally efficient representation of groups based on a graph theory approach where sequence clusters are described by a minimal spanning tree of highest scoring similarity pairs. This representation allows a classification of N proteins to be stored in order(N) memory. The use of this minimal spanning tree representation simplifies analysis of groups, the description of group characteristics and the manual correction of artifacts resulting from false hits. The new tree representation also introduces new possibilities for artifact generation in sequence classification. Methods for detecting and removing these artifacts are discussed.

Algorithms↗

Drosophila melanogaster as a model for studying protein-encoding genes that are resident in constitutive heterochromatin.

The organization of chromosomes into euchromatin and heterochromatin is one of the most enigmatic aspects of genome evolution. For a long time, heterochromatin was considered to be a genomic wasteland, incompatible with gene expression. However, recent studies--primarily conducted in Drosophila melanogaster--have shown that this peculiar genomic component performs important cellular functions and carries essential genes. New research on the molecular organization, function and evolution of heterochromatin has been facilitated by the sequencing and annotation of heterochromatic DNA. About 450 predicted genes have been identified in the heterochromatin of D. melanogaster, indicating that the number of active genes is higher than had been suggested by genetic analysis. Most of the essential genes are still unknown at the molecular level, and a detailed functional analysis of the predicted genes is difficult owing to the lack of mutant alleles. Far from being a peculiarity of Drosophila, heterochromatic genes have also been found in Saccharomyces cerevisiae, Schizosaccharomyces pombe, Oryza sativa and Arabidopsis thaliana, as well as in humans. The presence of expressed genes in heterochromatin seems paradoxical because they appear to function in an environment that has been considered incompatible with gene expression. In the future, genetic, functional genomic and proteomic analyses will offer powerful approaches with which to explore the functions of heterochromatic genes and to elucidate the mechanisms driving their expression.

Animals↗

An EST-based approach for identifying genes expressed in the intestine and gills of pre-smolt Atlantic salmon (Salmo salar).

BACKGROUND: The Atlantic salmon is an important aquaculture species and a very interesting species biologically, since it spawns in fresh water and develops through several stages before becoming a smolt, the stage at which it migrates to the sea to feed. The dramatic change of habitat requires physiological, morphological and behavioural changes to prepare the salmon for its new environment. These changes are called the parr-smolt transformation or smoltification, and pre-adapt the salmon for survival and growth in the marine environment. The development of hypo-osmotic regulatory ability plays an important part in facilitating the transition from rivers to the sea. The physiological mechanisms behind the developmental changes are largely unknown. An understanding of the transformation process will be vital to the future of the aquaculture industry. A knowledge of which genes are expressed prior to the smoltification process is an important basis for further studies. RESULTS: In all, 2974 unique sequences, consisting of 779 contigs and 2195 singlets, were generated for Atlantic salmon from two cDNA libraries constructed from the gills and the intestine, accession numbers [Genbank: CK877169-CK879929, CK884015-CK886537 and CN181112-CN181464]. Nearly 50% of the sequences were assigned putative functions because they showed similarity to known genes, mostly from other species, in one or more of the databases used. The Swiss-Prot database returned significant hits for 1005 sequences. These could be assigned predicted gene products, and 967 were annotated using Gene Ontology (GO) terms for molecular function, biological process and/or cellular component, employing an annotation transfer procedure. CONCLUSION: This paper describes the construction of two cDNA libraries from pre-smolt Atlantic salmon (Salmo salar) and the subsequent EST sequencing, clustering and assigning of putative function to 1005 genes expressed in the gills and/or intestine.

Animals↗

MMDB: Entrez's 3D-structure database.

Three-dimensional structures are now known within most protein families and it is likely, when searching a sequence database, that one will identify a homolog of known structure. The goal of Entrez's 3D-structure database is to make structure information and the functional annotation it can provide easily accessible to molecular biologists. To this end, Entrez's search engine provides several powerful features: (i) links between databases, for example between a protein's sequence and structure; (ii) pre-computed sequence and structure neighbors; and (iii) structure and sequence/structure alignment visualization. Here, we focus on a new feature of Entrez's Molecular Modeling Database (MMDB): Graphical summaries of the biological annotation available for each 3D structure, based on the results of automated comparative analysis. MMDB is available at: http://www.ncbi.nlm.nih.gov/Entrez/structure.html.

Animals↗

Low molecular weight proteins: a challenge for post-genomic research.

The EcoGene project involves the examination of Escherichia coli K-12 DNA sequences and accompanying annotation in the public databases in order to refine the representation and prediction of the entire set of E. coli K-12 chromosomally encoded protein sequences. The results of this ongoing effort have been deposited in the SWISSPROT protein sequence database as sequencing of the E. coli genome has progressed to completion in recent years. Through this continuing research, we have discovered that the prediction of low molecular weight (small) proteins, arbitrarily defined as protein sequences < or = 150 amino acids (aa) in length, is problematic and requires special attention. We describe the small protein subset of EcoGene and the approach used to derive this subset from the complete E. coli genome sequence and database annotations. These E. coli proteins have helped to identify new small genes in other organisms and to identify conserved residues (motifs) using database searches and multiple alignments. Two thirds of the E. coli small proteins have not been characterized experimentally. The careful application of computer and laboratory methods to the analysis of small proteins is needed for accurate prediction, verification and characterization. The problem of accurate protein sequence identification is not limited to small proteins or to E. coli; these problems are encountered to varying degrees throughout all sequence databases.

Amino Acid Sequence↗

Enhanced automated function prediction using distantly related sequences and contextual association by PFP.

The impetus for the recent development and emergence of automated function prediction methods is an exponentially growing flood of new experimental data, the interpretation of which is hindered by a shortage of reliable annotations for proteins that lack experimental characterization or significant homologs in current databases. Here we introduce PFP, an automated function prediction server that provides the most probable annotations for a query sequence in each of the three branches of the Gene Ontology: biological process, molecular function, and cellular component. Rather than utilizing precise pattern matching to identify functional motifs in the sequences and structures of these proteins, we designed PFP to increase the coverage of function annotation by lowering resolution of predictions when a detailed function is not predictable. To do this we extend a traditional PSI-BLAST search by extracting and scoring annotations (GO terms) individually, including annotations from distantly related sequences, and applying a novel data mining tool, the Function Association Matrix, to score strongly associated pairs of annotations. We show that PFP can correctly assign function using only weakly similar sequences with a significantly better accuracy and coverage than a standard PSI-BLAST search, improving it more than fivefold. The most descriptive annotations predicted by PFP (GO depth > or = 8) can identify a significant subgraph in the GO with > 60% accuracy and approximately 100% coverage for our benchmark set. We also provide examples of the superb performance of PFP in an assessment of automated function prediction servers at the Automated Function Prediction Special Interest Group meeting at ISMB 2005 (AFP-SIG '05).

Algorithms↗

Unexpected catalytic site variation in phosphoprotein phosphatase homologues of cofactor-dependent phosphoglycerate mutase.

The cofactor-dependent phosphoglycerate mutase (dPGM) superfamily contains, besides mutases, a variety of phosphatases, both broadly and narrowly substrate-specific. Distant dPGM homologues, conspicuously abundant in microbial genomes, represent a challenge for functional annotation based on sequence comparison alone. Here we carry out sequence analysis and molecular modelling of two families of bacterial dPGM homologues, one the SixA phosphoprotein phosphatases, the other containing various proteins of no known molecular function. The models show how SixA proteins have adapted to phosphoprotein substrate and suggest that the second family may also encode phosphoprotein phosphatases. Unexpected variation in catalytic and substrate-binding residues is observed in the models.

Amino Acid Sequence↗

Total sequence decomposition distinguishes functional modules, "molegos" in apurinic/apyrimidinic endonucleases.

BACKGROUND: Total sequence decomposition, using the web-based MASIA tool, identifies areas of conservation in aligned protein sequences. By structurally annotating these motifs, the sequence can be parsed into individual building blocks, molecular legos ("molegos"), that can eventually be related to function. Here, the approach is applied to the apurinic/apyrimidinic endonuclease (APE) DNA repair proteins, essential enzymes that have been highly conserved throughout evolution. The APEs, DNase-1 and inositol 5'-polyphosphate phosphatases (IPP) form a superfamily that catalyze metal ion based phosphorolysis, but recognize different substrates. RESULTS: MASIA decomposition of APE yielded 12 sequence motifs, 10 of which are also structurally conserved within the family and are designated as molegos. The 12 motifs include all the residues known to be essential for DNA cleavage by APE. Five of these molegos are sequentially and structurally conserved in DNase-1 and the IPP family. Correcting the sequence alignment to match the residues at the ends of two of the molegos that are absolutely conserved in each of the three families greatly improved the local structural alignment of APEs, DNase-1 and synaptojanin. Comparing substrate/product binding of molegos common to DNase-1 showed that those distinctive for APEs are not directly involved in cleavage, but establish protein-DNA interactions 3' to the abasic site. These additional bonds enhance both specific binding to damaged DNA and the processivity of APE1. CONCLUSION: A modular approach can improve structurally predictive alignments of homologous proteins with low sequence identity and reveal residues peripheral to the traditional "active site" that control the specificity of enzymatic activity.

Amino Acid Motifs↗

Analysis and prediction of functionally important sites in proteins.

The rapidly increasing volume of sequence and structure information available for proteins poses the daunting task of determining their functional importance. Computational methods can prove to be very useful in understanding and characterizing the biochemical and evolutionary information contained in this wealth of data, particularly at functionally important sites. Therefore, we perform a detailed survey of compositional and evolutionary constraints at the molecular and biological function level for a large set of known functionally important sites extracted from a wide range of protein families. We compare the degree of conservation across different functional categories and provide detailed statistical insight to decipher the varying evolutionary constraints at functionally important sites. The compositional and evolutionary information at functionally important sites has been compiled into a library of functional templates. We developed a module that predicts functionally important columns (FIC) of an alignment based on the detection of a significant "template match score" to a library template. Our template match score measures an alignment column's similarity to a library template and combines a term explicitly representing a column's residue composition with various evolutionary conservation scores (information content and position-specific scoring matrix-derived statistics). Our benchmarking studies show good sensitivity/specificity for the prediction of functional sites and high accuracy in attributing correct molecular function type to the predicted sites. This prediction method is based on information derived from homologous sequences and no structural information is required. Therefore, this method could be extremely useful for large-scale functional annotation.

Binding Sites↗

Information integration in molecular bioscience.

Integrating information in the molecular biosciences involves more than the cross-referencing of sequences or structures. Experimental protocols, results of computational analyses, annotations and links to relevant literature form integral parts of this information, and impart meaning to sequence or structure. In this review, we examine some existing approaches to integrating information in the molecular biosciences. We consider not only technical issues concerning the integration of heterogeneous data sources and the corresponding semantic implications, but also the integration of analytical results. Within the broad range of strategies for integration of data and information, we distinguish between platforms and developments. We discuss two current platforms and six current developments, and identify what we believe to be their strengths and limitations. We identify key unsolved problems in integrating information in the molecular biosciences, and discuss possible strategies for addressing them including semantic integration using ontologies, XML as a data model, and graphical user interfaces as integrative environments.

Animals↗

Genomic and bioinformatics analyses of HAdV-4vac and HAdV-7vac, two human adenovirus (HAdV) strains that constituted original prophylaxis against HAdV-related acute respiratory disease, a reemerging epidemic disease.

Vaccine strains of human adenovirus serotypes 4 and 7 (HAdV-4vac and HAdV-7vac) have been used successfully to prevent adenovirus-related acute respiratory disease outbreaks. The genomes of these two vaccine strains have been sequenced, annotated, and compared with their prototype equivalents with the goals of understanding their genomes for molecular diagnostics applications, vaccine redevelopment, and HAdV pathoepidemiology. These reference genomes are archived in GenBank as HAdV-4vac (35,994 bp; AY594254) and HAdV-7vac (35,240 bp; AY594256). Bioinformatics and comparative whole-genome analyses with their recently reported and archived prototype genomes reveal six mismatches and four insertions-deletions (indels) between the HAdV-4 prototype and vaccine strains, in contrast to the 611 mismatches and 130 indels between the HAdV-7 prototype and vaccine strains. Annotation reveals that the HAdV-4vac and HAdV-7vac genomes contain 51 and 50 coding units, respectively. Neither vaccine strain appears to be attenuated for virulence based on bioinformatics analyses. There is evidence of genome recombination, as the inverted terminal repeat of HAdV-4vac is initially identical to that of species C whereas the prototype is identical to species B1. These vaccine reference sequences yield unique genome signatures for molecular diagnostics. As a molecular forensics application, these references identify the circulating and problematic 1950s era field strains as the original HAdV-4 prototype and the Greider prototype, from which the vaccines are derived. Thus, they are useful for genomic comparisons to current epidemic and reemerging field strains, as well as leading to an understanding of pathoepidemiology among the human adenoviruses.

Acute Disease↗