PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Molecular Sequence Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Deduction of functional peptide motifs in scorpion toxins.

Scorpion toxins are important physiological probes for characterizing ion channels. Molecular databases have limited functional annotation of scorpion toxins. Their function can be inferred by searching for conserved motifs in sequence signature databases that are derived statistically but are not necessarily biologically relevant. Mutation studies provide biological information on residues and positions important for structure-function relationship but are not normally used for extraction of binding motifs. 3D structure analyses also aid in the extraction of peptide motifs in which non-contiguous residues are clustered spatially. Here we present new, functionally relevant peptide motifs for ion channels, derived from the analyses of scorpion toxin native and mutant peptides.

Amino Acid Motifs↗

Metabolism and genetics of Helicobacter pylori: the genome era.

The publication of the complete sequence of Helicobacter pylori 26695 in 1997 and more recently that of strain J99 has provided new insight into the biology of this organism. In this review, we attempt to analyze and interpret the information provided by sequence annotations and to compare these data with those provided by experimental analyses. After a brief description of the general features of the genomes of the two sequenced strains, the principal metabolic pathways are analyzed. In particular, the enzymes encoded by H. pylori involved in fermentative and oxidative metabolism, lipopolysaccharide biosynthesis, nucleotide biosynthesis, aerobic and anaerobic respiration, and iron and nitrogen assimilation are described, and the areas of controversy between the experimental data and those provided by the sequence annotation are discussed. The role of urease, particularly in pH homeostasis, and other specialized mechanisms developed by the bacterium to maintain its internal pH are also considered. The replicational, transcriptional, and translational apparatuses are reviewed, as is the regulatory network. The numerous findings on the metabolism of the bacteria and the paucity of gene expression regulation systems are indicative of the high level of adaptation to the human gastric environment. Arguments in favor of the diversity of H. pylori and molecular data reflecting possible mechanisms involved in this diversity are presented. Finally, we compare the numerous experimental data on the colonization factors and those provided from the genome sequence annotation, in particular for genes involved in motility and adherence of the bacterium to the gastric tissue.

Gene Expression Regulation, Bacterial↗

Subunit composition of NDH-1 complexes of Synechocystis sp. PCC 6803: identification of two new ndh gene products with nuclear-encoded homologues in the chloroplast Ndh complex.

Cyanobacteria contain several genes, annotated ndh, whose products show sequence similarities to subunits found in complex I (NADH:ubiquinone oxidoreductase) of eubacteria and mitochondria. However, it is still unclear whether the cyanobacterial ndh gene products actually form a single large protein complex or exist as smaller independent complexes. To address this, we have constructed a strain of Synechocystis sp. PCC 6803 in which the C terminus of the NdhJ subunit was fused to an His(6) tag to aid isolation. Three major NdhJ-containing complexes were resolved by blue native polyacrylamide gel electrophoresis, with approximate apparent molecular masses of 460, 330, and 110 kDa. N-terminal sequencing and mass spectrometry revealed that the 460-kDa complex contained ten annotated ndh gene products. Detergent-induced fragmentation experiments indicated that the 460-kDa complex was composed of hydrophobic (150 kDa) and hydrophilic (110-130 kDa) modules similar to that found in the minimal form of complex I found in Escherichia coli, except that the electron input module was not conserved. The difference in size between the 460- and 330-kDa complexes is attributed to differences in the stoichiometry of the hydrophilic and hydrophobic modules in the complex, either 2:1 or 1:1, respectively. We have also detected the presence of two new Ndh subunits (slr1623 and sll1262) that are unrelated to subunits in the eubacterial complex I but which have homologues in the closely related chloroplast Ndh complex of maize (Funk, E., Schäfer, E., and Steinmüller, K. (1999) J. Plant Physiol. 154, 16-23). The presence of these additional subunits might reflect the use by the NDH-1 and Ndh complexes of a different, so far unidentified, electron input module.

Amino Acid Sequence↗

Structural genomics of minimal organisms and protein fold space.

The initial aim of the Berkeley Structural Genomics Center is to obtain a near-complete structural complement of two minimal organisms, closely related pathogens Mycoplasma genitalium and M. pneumoniae. The former has fewer than 500 genes and the latter fewer than 700 genes. To achieve this goal, the current protein targets have been selected starting with those predicted to be most tractable and likely to yield new structural and functional information. During the past 3 years, the semi-automated structural genomics pipeline has been set up from cloning, expression, purification, and ultimately to structural determination. The results from the pipeline substantially increased the coverage of the protein fold space of M. pneumoniae and M. genitalium. Furthermore, about 1/2 of the structures of 'unique' protein sequences revealed new and novel folds, and over 2/3 of the structures of previously annotated 'hypothetical proteins' inferred their molecular functions.

Bacterial Proteins↗

Advances in diagnostic tests for bacterial STDs.

Because of their asymptomatic nature and nonspecific symptoms, laboratory tests are often required to diagnose a sexually transmitted infection. Over the past few years, there have been advances in technology, such as the development of nucleic acid amplification assays, which have improved our ability to diagnose infections caused by Chlamydia trachomatis. The finding that nucleic acid amplification tests can detect more infected individuals and are useful in screening low prevalence populations, has led to the development of strategies designed to reduce the cost of these assays without significantly impacting their sensitivity. The development of new tests for the diagnosis of syphilis has gained momentum from the report of a synthetic VDRL antigen, which will result in better nontreponemal antibody tests for syphilis. In spite of the completion of the genome sequence of Treponema pallidum and its annotation, we are still unable to cultivate this microorganism in vitro. However, the molecular revolution has resulted in the development of PCR assays for detecting Treponema pallidum in various types of clinical specimens, and to the production of recombinant antigens for use in tests that detect treponemal-specific antibodies. Further research will improve the availability of low cost, sensitive tests for the diagnosis of sexually transmitted infections. The English version of this paper is available too at:http://www.insp.mx/salud/index.html.

Bacterial Infections↗

Phylogeny of Na+/Ca2+ exchanger (NCX) genes from genomic data identifies new gene duplications and a new family member in fish species.

The Na+/Ca2+ exchanger (NCX) is a member of the cation/Ca2+ antiporter (CaCA) family and plays a key role in maintaining cellular Ca2+ homeostasis in a variety of cell types. NCX is present in a diverse group of organisms and exhibits high overall identity across species. To date, three separate genes, i.e., NCX1, NCX2, and NCX3, have been identified in mammals. However, phylogenetic analysis of the exchanger has been hindered by the lack of nonmammalian NCX sequences. In this study, we expand and diversify the list of NCX sequences by identifying NCX homologs from whole-genome sequences accessible through the Ensembl Genome Browser. We identified and annotated 13 new NCX sequences, including 4 from zebrafish, 4 from Japanese pufferfish, 2 from chicken, and 1 each from honeybee, mosquito, and chimpanzee. Examination of NCX gene structure, together with construction of phylogenetic trees, provided novel insights into the molecular evolution of NCX and allowed us to more accurately annotate NCX gene names. For the first time, we report the existence of NCX2 and NCX3 in organisms other than mammals, yielding the hypothesis that two serial NCX gene duplications occurred around the time vertebrates and invertebrates diverged. In addition, we have found a putative new NCX protein, named NCX4, that is related to NCX1 but has been observed only in fish species genomes. These findings present a stronger foundation for our understanding of the molecular evolution of the NCX gene family and provide a framework for further NCX phylogenetic and molecular studies.

Amino Acid Sequence↗

Exploring biosynthetic potential of the endophytic Penicillium turbatum BLH34 using whole-genome sequence analysis and molecular networking.

An in-depth genomic and metabolomic investigation was conducted on the endophytic fungus Penicillium turbatum BLH34, isolated from Macleaya cordata. Hybrid sequencing (Illumina-Nanopore) generated a high-quality 27.9 Mb genome (GC 48.6%) encoding 9798 proteins, with functional annotation linking 5350 genes to the NCBI non-redundant database and 3404 to KEGG pathways. AntiSMASH analysis uncovered 35 biosynthetic gene clusters (BGCs), 23 of which lacked homology to known pathways, highlighting BLH34's potential for novel metabolite discovery. Molecular networking (GNPS) and LC-MS/MS identified 19 specialised metabolites, including antimicrobial polyketides. Bioassays demonstrated potent inhibition against Staphylococcus aureus (36 mm), Bacillus subtilis (28 mm) and Escherichia coli (24 mm), underscoring its pharmaceutical relevance.

Penicillium↗

Protein family alignment annotation.

For bioscientists studying protein structure and function, the Protein Family Alignment Annotation Tool (Pfaat) is a useful and simple program for annotating collections of proteins. This open-source software includes methods for viewing and aligning protein families, and for annotating sequence structure and residues with known functions. It offers new options to aid the study of proteins, and an extensible annotation tool for bioinformatics developers.

Amino Acid Sequence↗

Medium-chain dehydrogenases/reductases (MDR). Family characterizations including genome comparisons and active site modeling.

Completed eukaryotic genomes were screened for medium-chain dehydrogenases/reductases (MDR). In the human genome, 23 MDR forms were found, a number that probably will increase, because the genome is not yet fully interpreted. Partial sequences already indicate that at least three further members exist. Within the MDR superfamily, at least eight families were distinguished. Three families are formed by dimeric alcohol dehydrogenases (ADH; originally detected in animals/plants), cinnamyl alcohol dehydrogenases (originally detected in plants) and tetrameric alcohol dehydrogenases (originally detected in yeast). Three further families are centred around forms initially detected as mitochondrial respiratory function proteins, acetyl-CoA reductases of fatty acid synthases, and leukotriene B4 dehydrogenases. The two remaining families with polyol dehydrogenases (originally detected as sorbitol dehydrogenase) and quinone reductases (originally detected as zeta-crystallin) are also distinct but with variable sequences. The most abundant families in the human genome are the dimeric ADH forms and the quinone oxidoreductases. The eukaryotic patterns are different from those of Escherichia coli. The different families were further evaluated by molecular modelling of their active sites as to geometry, hydrophobicity and volume of substrate-binding pockets. Finally, sequence patterns were derived that are diagnostic for the different families and can be used in genome annotations.

Amino Acid Motifs↗

Molecular identification of a Drosophila G protein-coupled receptor specific for crustacean cardioactive peptide.

The Drosophila Genome Project website (www.flybase.org) contains the sequence of an annotated gene (CG6111) expected to code for a G protein-coupled receptor. We have cloned this receptor and found that its gene was not correctly predicted, because an annotated neighbouring gene (CG14547) was also part of the receptor gene. DNA corresponding to the corrected gene CG6111 was expressed in Chinese hamster ovary cells, where it was found to code for a receptor that could be activated by low concentrations of crustacean cardioactive peptide, which is a neuropeptide also known to occur in Drosophila and other insects (EC(50), 5.4 x 10(-10)M). Other known Drosophila neuropeptides, such as adipokinetic hormone, did not activate the receptor. The receptor is expressed in all developmental stages from Drosophila, but only very weakly in larvae. In adult flies, the receptor is mainly expressed in the head. Furthermore, we identified a gene sequence in the genomic database from the malaria mosquito Anopheles gambiae that very likely codes for a crustacean cardioactive peptide receptor.

Amino Acid Sequence↗

The evolutionary analysis of "orphans" from the Drosophila genome identifies rapidly diverging and incorrectly annotated genes.

In genome projects of eukaryotic model organisms, a large number of novel genes of unknown function and evolutionary history ("orphans") are being identified. Since many orphans have no known homologs in distant species, it is unclear whether they are restricted to certain taxa or evolve rapidly, either because of a lack of constraints or positive Darwinian selection. Here we use three criteria for the selection of putatively rapidly evolving genes from a single sequence of Drosophila melanogaster. Thirteen candidate genes were chosen from the Adh region on the second chromosome and 1 from the tip of the X chromosome. We succeeded in obtaining sequence from 6 of these in the closely related species D. simulans and D. yakuba. Only 1 of the 6 genes showed a large number of amino acid replacements and in-frame insertions/deletions. A population survey of this gene suggests that its rapid evolution is due to the fixation of many neutral or nearly neutral mutations. Two other genes showed "normal" levels of divergence between species. Four genes had insertions/deletions that destroy the putative reading frame within exons, suggesting that these exons have been incorrectly annotated. The evolutionary analysis of orphan genes in closely related species is useful for the identification of both rapidly evolving and incorrectly annotated genes.

Animals↗

A polyphosphate kinase (PPK2) widely conserved in bacteria.

Synthesis of inorganic polyphosphate (poly P) from the terminal phosphate of ATP is catalyzed reversibly by poly P kinase (PPK, now designated PPK1) initially isolated from Escherichia coli. PPK1 is highly conserved in many bacteria, including some of the major pathogens such as Pseudomonas aeruginosa. In a null mutant of P. aeruginosa lacking ppk1, we have discovered a previously uncharacterized PPK activity (designated PPK2) distinguished from PPK1 by the following: synthesis of poly P from GTP or ATP, a preference for Mn2+ over Mg2+, and a stimulation by poly P. The reverse reaction, a poly P-driven nucleoside diphosphate kinase synthesis of GTP from GDP, is 75-fold greater than the forward reaction, poly P synthesis from GTP. The gene encoding PPK2 (ppk2) was identified from the amino acid sequence of the protein purified near 1,000-fold, to homogeneity. The 5'-end is 177 bp upstream of the annotated genome sequence of a "conserved hypothetical protein"; ppk2 (1,074 bp) encodes a protein of 357 aa with a molecular mass of 40.8 kDa. Sequences homologous to PPK2 are present in two other proteins in P. aeruginosa, in two Archaea, and in 32 other bacteria (almost all with PPK1 as well); these include rhizobia, cyanobacteria, Streptomyces, and several pathogenic species. Distinctive features of the poly P-driven nucleoside diphosphate kinase activity and structural aspects of PPK2 are among the subjects of an accompanying report.

Amino Acid Sequence↗

Improving quality of expressed sequence tag (EST) databases: recovery of reversed, antisense cDNA sequences.

Expressed sequence tag (EST) databases contain a significant number (5-20%) of reversed, antisense, cDNA sequences that can be recognized by the label "reversed clone: similarity on wrong strand" in the annotations to the sequence. Despite this high number of altered sequences, no attempt has been made to explain the alteration in molecular terms, or to evaluate their effect on the quality of the information curated in EST databases. In this paper we try to explain the way these altered sequences are originated, and propose a plausible mechanism: a "double priming" of the first strand oligo-dT primer at both ends of nascent cDNAs. In this way, a symmetrical cDNA intermediate is generated, an intermediate that can be cloned after partial digestion with the restriction enzyme used for the directional cloning. Furthermore, when "secondary" priming takes place inside the cDNA, the chain synthesized is prone to be truncated prematurely, with the subsequent loss of upstream information. One of the most subtle effects of this cloning alteration is the generation of virtual open reading frames (ORFs) in sequences with no homologues available for comparison. Nevertheless, and according to our model and our data, the "double priming mechanism" does not shift the ORF effected, so antisense sequences should be considered as normal ones after a simple transformation in their inverse-complementary forms.

Artifacts↗

Identification of related proteins with weak sequence identity using secondary structure information.

Molecular modeling of proteins is confronted with the problem of finding homologous proteins, especially when few identities remain after the process of molecular evolution. Using even the most recent methods based on sequence identity detection, structural relationships are still difficult to establish with high reliability. As protein structures are more conserved than sequences, we investigated the possibility of using protein secondary structure comparison (observed or predicted structures) to discriminate between related and unrelated proteins sequences in the range of 10%-30% sequence identity. Pairwise comparison of secondary structures have been measured using the structural overlap (Sov) parameter. In this article, we show that if the secondary structures likeness is >50%, most of the pairs are structurally related. Taking into account the secondary structures of proteins that have been detected by BLAST, FASTA, or SSEARCH in the noisy region (with high E: value), we show that distantly related protein sequences (even with <20% identity) can be still identified. This strategy can be used to identify three-dimensional templates in homology modeling by finding unexpected related proteins and to select proteins for experimental investigation in a structural genomic approach, as well as for genome annotation.

Algorithms↗

Primer on medical genomics. Part IV: Expression proteomics.

Proteomics, simply defined, is the study of proteomes. More completely, proteomics is defined as the study of all proteins, including their relative abundance, distribution, posttranslational modifications, functions, and interactions with other macromolecules, in a given cell or organism within a given environment and at a specific stage in the cell cycle. Proteins carry out the biological functions encoded by genes; hence, once the initial stage of genome sequencing and gene discovery is completed, a study of the proteome must be undertaken to address fundamental biological questions. The 3 broad areas are expression proteomics, which catalogues the relative abundance of proteins; cell-mapping or cellular proteomics, which delineates functional protein-protein interactions and organelle-specific protein distribution; and structural proteomics, which characterizes the 3-dimensional structure of proteins. With these approaches, proteins are studied on a global scale using a synergistic combination of powerful, high-throughput technologies, including 2-dimensional polyacrylamide gel electrophoresis, mass spectrometry, multidimensional liquid chromatography, and bioinformatics. Mass spectrometry, which provides highly accurate molecular mass measurements, has emerged as the analytical technology of choice for protein identification, characterization, and sequencing. This task has been made considerably easier with the availability of complete, nonredundant, and annotated genome sequence databases for many organisms. This article reviews the area of expression proteomics.

Biotechnology↗

Genetic landscape of pediatric seizures in Southeast China: identification of a novel GLI3 frameshift variant through whole-exome sequencing.

BACKGROUND: Pediatric seizure disorders are clinically and genetically heterogeneous. Whole-exome sequencing has improved the detection of rare genetic variants in childhood epilepsy; however, data from pediatric populations in Southeast China remain limited. This study aimed to characterize the genetic landscape of pediatric seizure disorders in Southeast China and to evaluate the clinical diagnostic yield of whole-exome sequencing. MATERIALS AND METHODS: This retrospective observational study included 21 pediatric patients with seizure disorders who were recruited at the Fifth Hospital of Xiamen, Fujian, China, between January 2021 and June 2024. Clinical data were extracted from medical records. Whole-exome sequencing was performed on DNA extracted from peripheral blood. Sequence variants were annotated, filtered, and classified according to the guidelines of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Copy-number variants were evaluated using exome-based algorithms. Descriptive statistics were used because of the limited sample size. RESULTS: WES identified three clinically relevant, likely pathogenic findings in 3 of 21 patients, corresponding to a provisional diagnostic yield of 14.3%. The remaining 62 of 65 variants were of uncertain significance (VUS). The three retained variants included a GLI3 frameshift variant (exon 2: c.90_91insCAGATGTGAGC; p.Glu31Glnfs*3) and two copy-number variants (16p13.12-16p13.11 duplication and Xp22.31 deletion) with established clinical significance. Functional analysis of all 65 variants revealed that ion channel genes and neurodevelopmental genes were the most frequently affected categories. CONCLUSION: Whole-exome sequencing identified clinically relevant genetic findings in a subset of Southeast Chinese children with seizure disorders. The novel GLI3 frameshift variant may suggest an expansion of the GLI3-associated phenotypic spectrum, but further segregation, functional validation, and larger cohort studies are needed. The high proportion of variants of uncertain significance highlights the ongoing challenges of genetic interpretation in pediatric seizure disorders.

GLI3 frameshift variant↗

Cataloging transcription factor and major signaling molecule genes for functional genomic studies in Ciona intestinalis.

The ascidian Ciona intestinalis provides an excellent experimental system for functional genomic studies because (1) its genome has been sequenced, (2) the transcription factor genes and genes for major signal transduction molecules have been extensively screened and annotated on a genome-wide scale using the molecular phylogenetical method, and (3) their embryonic expression profiles have been almost completely determined. However, the entire genetic structure, including the 5' and 3' untranslated regions and the protein-coding regions, of most gene models used in these prior studies is not always supported by cDNA evidence, and thus, these gene models are potentially imprecise. To facilitate functional genomic studies based on precise gene structures, our present study determined 406 cDNA sequences for 357 transcription factor genes and 112 cDNA sequences for 107 signal transduction molecule genes, greatly improving the previous gene models and revealing transcript variants for 44 genes. Considering these data alongside those of previously characterized genes deposited in the DNA Data Bank of Japan/European Molecular Biology Laboratory/GENBANK databases, 95.6% of the catalogued transcription factor genes (373/390) and 98.3% of the catalogued signal transduction molecule genes (117/119) have now been verified by cDNA sequences. Thus, the present study greatly improves the resources available for functional genomic studies in C. intestinalis.

Animals↗

Survey of transcripts in the adult Drosophila brain.

BACKGROUND: Classic methods of identifying genes involved in neural function include the laborious process of behavioral screening of mutagenized flies and then rescreening candidate lines for pleiotropic effects due to developmental defects. To accelerate the molecular analysis of brain function in Drosophila we constructed a cDNA library exclusively from adult brains. Our goal was to begin to develop a catalog of transcripts expressed in the brain. These transcripts are expected to contain a higher proportion of clones that are involved in neuronal function. RESULTS: The library contains approximately 6.75 million independent clones. From our initial characterization of 271 randomly chosen clones, we expect that approximately 11% of the clones in this library will identify transcribed sequences not found in expressed sequence tag databases. Furthermore, 15% of these 271 clones are not among the 13,601 predicted Drosophila genes. CONCLUSIONS: Our analysis of this unique Drosophila brain library suggests that the number of genes may be underestimated in this organism. This work complements the Drosophila genome project by providing information that facilitates more complete annotation of the genomic sequence. This library should be a useful resource that will help in determining how basic brain functions operate at the molecular level.

Aging↗