PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Molecular Sequence Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Transcriptome analysis of Neotyphodium and Epichloë grass endophytes.

Large-scale gene discovery has been performed for the grass fungal endophytes Neotyphodium coenophialum, Neotyphodium lolii, and Epichloë festucae. The resulting sequences have been annotated by comparison with public DNA and protein sequence databases and using intermediate gene ontology annotation tools. Endophyte sequences have also been analysed for the presence of simple sequence repeat and single nucleotide polymorphism molecular genetic markers. Sequences and annotation are maintained within a MySQL database that may be queried using a custom web interface. Two cDNA-based microarrays have been generated from this genome resource. They permit the interrogation of 3806 Neotyphodium genes (Nchip microarray), and 4195 Neotyphodium and 920 Epichloë genes (EndoChip microarray), respectively. These microarrays provide tools for high-throughput transcriptome analysis, including genome-specific gene expression studies, profiling of novel endophyte genes, and investigation of the host grass-symbiont interaction. Comparative transcriptome analysis in Neotyphodium and Epichloë was performed.

Computational Biology↗

Hepatitis C databases, principles and utility to researchers.

Part of the effort to develop hepatitis C-specific drugs a nd vaccines is the study of genetic variability of allpublicly available HCV sequences. Three HCV databases are currently available to aid this effort and to provide additional insight into the basic biology, immunology, and evolution of the virus. The Japanese HCV database (http://s2as02.genes.nig.ac.jp) gives access to a genomic mapping of sequences as well as their phylogenetic relationships. The European HCV database (http://euhcvdb.ibcp.fr) offers access to a computer-annotated set of sequences and molecular models of HCV proteins and focuses on protein sequence, structure and function analysis. The HCV database at the Los Alamos National Laboratory in the United States (http://hcv.lanl.gov) provides access to a manually annotated sequence database and a database of immunological epitopes which contains concise descriptions of experimental results. In this paper, we briefly describe each of these databases and their associated websites and tools, and give some examples of their use in furthering HCV research.

Biomedical Research↗

The Annotator's Assistant: an expert system for direct submission of genetic sequence data.

As DNA sequencing technology improves and more rapid techniques become routine in molecular biology labs, researchers need to expedite the incorporation of information into genetic sequence databases, such as GenBank, by directly submitting sequence data. The Annotator's Assistant is an expert system that runs on an IBM PC and helps the molecular biologist, who may have little knowledge of the structure or content of a GenBank entry, to construct a complete and valid sequence submission file. This expert system uses a simple molecular biology knowledge base and a selection of customized screen entry forms to guide the user through the entry and annotation of a sequence and its biological features. The system compiles information about the contributor, journal references, physical and functional characteristics of the nucleic acid, source organism and features, and checks it to eliminate incomplete answers and simple errors. Users supply input by answering direct and multiple-choice questions, selecting menu items and completing entry forms; on-line help is available. Users may also enter new or unusual information using generic forms. Several modules of the expert system were converted into Prolog programs and compiled, decreasing the running time significantly. The expert system rules and the data entry forms are easy to modify, update and customize for specific sequence classes.

Base Sequence↗

A generalized affine gap model significantly improves protein sequence alignment accuracy.

Sequence alignment underpins common tasks in molecular biology, including genome annotation, molecular phylogenetics, and homology modeling. Fundamental to sequence alignment is the placement of gaps, which represent character insertions or deletions. We assessed the ability of a generalized affine gap cost model to reliably detect remote protein homology and to produce high-quality alignments. Generalized affine gap alignment with optimal gap parameters performed as well as the traditional affine gap model in remote homology detection. Evaluation of alignment quality showed that the generalized affine model aligns fewer residue pairs than the traditional affine model but achieves significantly higher per-residue accuracy. We conclude that generalized affine gap costs should be used when alignment accuracy carries more importance than aligned sequence length.

Algorithms↗

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype↗

Atlas - a data warehouse for integrative bioinformatics.

BACKGROUND: We present a biological data warehouse called Atlas that locally stores and integrates biological sequences, molecular interactions, homology information, functional annotations of genes, and biological ontologies. The goal of the system is to provide data, as well as a software infrastructure for bioinformatics research and development. DESCRIPTION: The Atlas system is based on relational data models that we developed for each of the source data types. Data stored within these relational models are managed through Structured Query Language (SQL) calls that are implemented in a set of Application Programming Interfaces (APIs). The APIs include three languages: C++, Java, and Perl. The methods in these API libraries are used to construct a set of loader applications, which parse and load the source datasets into the Atlas database, and a set of toolbox applications which facilitate data retrieval. Atlas stores and integrates local instances of GenBank, RefSeq, UniProt, Human Protein Reference Database (HPRD), Biomolecular Interaction Network Database (BIND), Database of Interacting Proteins (DIP), Molecular Interactions Database (MINT), IntAct, NCBI Taxonomy, Gene Ontology (GO), Online Mendelian Inheritance in Man (OMIM), LocusLink, Entrez Gene and HomoloGene. The retrieval APIs and toolbox applications are critical components that offer end-users flexible, easy, integrated access to this data. We present use cases that use Atlas to integrate these sources for genome annotation, inference of molecular interactions across species, and gene-disease associations. CONCLUSION: The Atlas biological data warehouse serves as data infrastructure for bioinformatics research and development. It forms the backbone of the research activities in our laboratory and facilitates the integration of disparate, heterogeneous biological sources of data enabling new scientific inferences. Atlas achieves integration of diverse data sets at two levels. First, Atlas stores data of similar types using common data models, enforcing the relationships between data types. Second, integration is achieved through a combination of APIs, ontology, and tools. The Atlas software is freely available under the GNU General Public License at: http://bioinformatics.ubc.ca/atlas/

Computational Biology↗

Inferring higher functional information for RIKEN mouse full-length cDNA clones with FACTS.

FACTS (Functional Association/Annotation of cDNA Clones from Text/Sequence Sources) is a semiautomated knowledge discovery and annotation system that integrates molecular function information derived from sequence analysis results (sequence inferred) with functional information extracted from text. Text-inferred information was extracted from keyword-based retrievals of MEDLINE abstracts and by matching of gene or protein names to OMIM, BIND, and DIP database entries. Using FACTS, we found that 47.5% of the 60,770 RIKEN mouse cDNA FANTOM2 clone annotations were informative for text searches. MEDLINE queries yielded molecular interaction-containing sentences for 23.1% of the clones. When disease MeSH and GO terms were matched with retrieved abstracts, 22.7% of clones were associated with potential diseases, and 32.5% with GO identifiers. A significant number (23.5%) of disease MeSH-associated clones were also found to have a hereditary disease association (OMIM Morbidmap). Inferred neoplastic and nervous system disease represented 49.6% and 36.0% of disease MeSH-associated clones, respectively. A comparison of sequence-based GO assignments with informative text-based GO assignments revealed that for 78.2% of clones, identical GO assignments were provided for that clone by either method, whereas for 21.8% of clones, the assignments differed. In contrast, for OMIM assignments, only 28.5% of clones had identical sequence-based and text-based OMIM assignments. Sequence, sentence, and term-based functional associations are included in the FACTS database (http://facts.gsc.riken.go.jp/), which permits results to be annotated and explored through web-accessible keyword and sequence search interfaces. The FACTS database will be a critical tool for investigating the functional complexity of the mouse transcriptome, cDNA-inferred interactome (molecular interactions), and pathome (pathologies).

Animals↗

Exact algorithms for computing pairwise alignments and 3-medians from structure-annotated sequences (extended abstract).

Given the problem of mutation saturation in ancient molecular sequences, there is great interest in inferring phylogenies from higher-order types of molecular data that change more slowly, such as genomic organization and the secondary and tertiary structures of ribosomal RNA and proteins. In this paper, we define edit distances based on two representations of RNA secondary structure, arc annotation and hierarchical string annotation, and give algorithms for computing these distances on pairs of annotated sequences, aligning pairs of annotated sequences, and computing 3-median annotated sequences from triples of annotated sequences. The 3-median algorithms can be used as part of a well-known iterative heuristic for inferring phylogenies. All given algorithms are adapted from algorithms for computing longest common annotated subsequences of pairs of annotated sequences.

Algorithms↗

Searchlight on domains.

In this issue of Structure, examine in detail the functions of selected domains within proteins both when they are alone and when in combination with others. Domain function is relevant to molecular evolution and to annotation of proteins known only by sequence.

Enzymes↗

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7 Mb with a contig N50 of 39.3 Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant↗

Whole-genome analysis of Brevibacterium sanguinis AZMABM HM27: a bacterial isolate from the sea anemone Radianthus magnifica and exhibiting promising multi-therapeutic properties.

BACKGROUND: The marine anemone Radianthus magnifica harbors symbiotic microbes with promising biomedical potential, yet their diversity and therapeutic properties remain underexplored. This study aimed to characterize a symbiotic bacterium isolated from R. magnifica collected from Samalona Island, Indonesia, and to evaluate its multi-therapeutic potential. METHODS: Strain AZMABM HM27 was characterized using whole-genome sequencing, functional annotation, biosynthetic gene cluster prediction, molecular docking, and in vitro bioactivity assays. RESULTS: Phylogenetic and genome-based analyses confirmed AZMABM HM27 as Brevibacterium sanguinis, with an OrthoANI value of 97.37% and a dDDH value of 76.50% against the type strain. The genome comprises a 3,834,082 bp chromosome encoding 3,362 protein-coding genes, including 95 genes involved in secondary metabolite biosynthesis. Five biosynthetic gene clusters were predicted, including those associated with ectoine, terpene, and siderophore production. The crude extract demonstrated antioxidant activity (IC₅₀ = 0.87 mg/mL), anti-inflammatory activity (up to 60% inhibition), antidiabetic activity through α-glucosidase inhibition (up to 40% inhibition), and dose-dependent antiproliferative activity against MCF-7 breast cancer cells (74.10% viability at 1 mg/mL). Molecular docking identified a lead compound, 8,9,9,10,10,11-hexafluoro-4,4-dimethyl-3,5-dioxatetracyclo [5.4.1.0(2,6)0.0(8,11)] dodecane, with strong binding affinities to selected therapeutic targets. CONCLUSIONS: B. sanguinis AZMABM HM27 represents a marine symbiotic strain associated with R. magnifica and a promising source of bioactive compounds with antioxidant, anti-inflammatory, antidiabetic, and antiproliferative potential. Further purification, structural elucidation, and in vivo studies are warranted to validate its therapeutic potential.

Animals↗

Extension and integration of the gene ontology (GO): combining GO vocabularies with external vocabularies.

Structured vocabulary development enhances the management of information in biological databases. As information grows, handling the complexity of vocabularies becomes difficult. Defined methods are needed to manipulate, expand and integrate complex vocabularies. The Gene Ontology (GO) project provides the scientific community with a set of structured vocabularies to describe domains of molecular biology. The vocabularies are used for annotation of gene products and for computational annotation of sequence data sets. The vocabularies focus on three concepts universal to living systems, biological process, molecular function and cellular component. As the vocabularies expand to incorporate terms needed by diverse annotation communities, species-specific terms become problematic. In particular, the use of species-specific anatomical concepts remains unresolved. We present a method for expansion of GO into areas outside of the three original universal concept domains. We combine concepts from two orthogonal vocabularies to generate a larger, more specific vocabulary. The example of mammalian heart development is presented because it addresses two issues that challenge GO; inclusion of organism-specific anatomical terms, and proliferation of terms and relationships. The combination of concepts from orthogonal vocabularies provides a robust representation of relevant terms and an opportunity for evaluation of hypothetical concepts.

Animals↗

SS-Wrapper: a package of wrapper applications for similarity searches on Linux clusters.

BACKGROUND: Large-scale sequence comparison is a powerful tool for biological inference in modern molecular biology. Comparing new sequences to those in annotated databases is a useful source of functional and structural information about these sequences. Using software such as the basic local alignment search tool (BLAST) or HMMPFAM to identify statistically significant matches between newly sequenced segments of genetic material and those in databases is an important task for most molecular biologists. Searching algorithms are intrinsically slow and data-intensive, especially in light of the rapid growth of biological sequence databases due to the emergence of high throughput DNA sequencing techniques. Thus, traditional bioinformatics tools are impractical on PCs and even on dedicated UNIX servers. To take advantage of larger databases and more reliable methods, high performance computation becomes necessary. RESULTS: We describe the implementation of SS-Wrapper (Similarity Search Wrapper), a package of wrapper applications that can parallelize similarity search applications on a Linux cluster. Our wrapper utilizes a query segmentation-search (QS-search) approach to parallelize sequence database search applications. It takes into consideration load balancing between each node on the cluster to maximize resource usage. QS-search is designed to wrap many different search tools, such as BLAST and HMMPFAM using the same interface. This implementation does not alter the original program, so newly obtained programs and program updates should be accommodated easily. Benchmark experiments using QS-search to optimize BLAST and HMMPFAM showed that QS-search accelerated the performance of these programs almost linearly in proportion to the number of CPUs used. We have also implemented a wrapper that utilizes a database segmentation approach (DS-BLAST) that provides a complementary solution for BLAST searches when the database is too large to fit into the memory of a single node. CONCLUSIONS: Used together, QS-search and DS-BLAST provide a flexible solution to adapt sequential similarity searching applications in high performance computing environments. Their ease of use and their ability to wrap a variety of database search programs provide an analytical architecture to assist both the seasoned bioinformaticist and the wet-bench biologist.

Algorithms↗

The RESID Database of Protein Modifications as a resource and annotation tool.

The RESID Database of Protein Modifications is a comprehensive collection of annotations and structures for protein modifications and cross-links including pre-, co-, and post-translational modifications. The database provides: systematic and alternate names, atomic formulas and masses, enzymatic activities that generate the modifications, keywords, literature citations, Gene Ontology (GO) cross-references, protein sequence database feature table annotations, structure diagrams, and molecular models. This database is freely accessible on the Internet through resources provided by the European Bioinformatics Institute (http://www.ebi.ac.uk/RESID), and by the National Cancer Institute--Frederick Advanced Biomedical Computing Center (http://www.ncifcrf.gov/RESID). Each RESID Database entry presents a chemically unique modification and shows how that modification is currently annotated in the protein sequence databases, Swiss-Prot and the Protein Information Resource (PIR). The RESID Database provides a table of corresponding equivalent feature annotations that is used in the UniProt project, an international effort to combine the resources of the Swiss-Prot, TrEMBL and PIR. As an annotation tool, the RESID Database is used in standardizing and enhancing modification descriptions in the feature tables of Swiss-Prot entries. As an Internet resource, the RESID Database assists researchers in high-throughput proteomics to search monoisotopic masses and mass differences and identify known and predicted protein modifications.

Databases, Factual↗

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62 Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52 Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals↗

BioEditor-simplifying macromolecular structure annotation.

SUMMARY: BioEditor is an application to enable scientists and educators to prepare and present structure annotations containing formatted text, graphics, sequence data, and interactive molecular views. It is intended to bridge the gap between printed journal articles and Internet presentation formats. BioEditor is relevant in the era of structural genomics, where annotation and publication could become the rate determining step in structure determination. AVAILABILITY: BioEditor is available at http://bioeditor.sdsc.edu. The Web site includes the latest version of the software for Microsoft Windows, including documentation, the opportunity to submit bug reports and suggestions, example documentaries prepared with BioEditor and a repository where users can submit documentaries for posting to the site.

Biopolymers↗

An annotated cDNA library and microarray for large-scale gene-expression studies in the ant Solenopsis invicta.

Ants display a range of fascinating behaviors, a remarkable level of intra-species phenotypic plasticity and many other interesting characteristics. Here we present a new tool to study the molecular mechanisms underlying these traits: a tentatively annotated expressed sequence tag (EST) resource for the fire ant Solenopsis invicta. From a normalized cDNA library we obtained 21,715 ESTs, which represent 11,864 putatively different transcripts with very diverse molecular functions. All ESTs were used to construct a cDNA microarray.

Animals↗

Finding errors in DNA sequences.

An algorithm is described that can detect certain errors within coding regions of DNA sequences. The algorithm is based on the idea that an insertion or deletion error within a coding sequence would interrupt the reading frame and cause the correct translation of a DNA sequence to require one or more frameshifts. If the coding sequence shows similarity to a known protein sequence then such errors can be detected by comparing the conceptual translations of DNA sequences in all six reading frames with every sequence in a protein sequence data base. We have incorporated these ideas into a computer program, called DETECT, that can serve as an aid to the experimentalist who is determining new DNA sequences so that obvious errors may be located and corrected. The program has been tested using raw experimental data and against sequences from the European Molecular Biology Laboratory data base, annotated as containing frameshifts. We have also tested it using unidentified open reading frames that flank known, annotated genes in the GenBank data base. Many potential errors are apparent and in some cases functions can be suggested for the "corrected" versions of these reading frames leading to the identification of new genes. As more sequences are determined the power of this method will increase substantially.

Adenylyl Cyclases↗