PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Comparative and evolutionary analysis of the cytochrome b sequences in cyprinids with different ploidy levels derived from crosses.

The mitochondrial cyt b genes in the allotetraploid and triploid crucian carp as well as triploid common carp were isolated and completely sequenced. Their DNA sequences were compared with those derived from the cyt b genes of the red crucian carp, Japanese crucian carp, and common carp with MEGA 1.0 software. Phylogenetic analysis revealed the sister relationships between allotetraploid and diploid red crucian carp, between the triploid crucian carp and diploid Japanese crucian carp, and between triploid common carp and diploid common carp. Our results indicated the cyt b genes in the allotetraploid, triploid crucian carp, and triploid common carp were maternally inherited. Through maternal inheritance, the cyt b gene in the F11 tetraploid displayed extremely high similarity to that in the female parent red crucian carp after 11 generations (from F1 to F11 hybrids). Since the establishment of the new tetraploid stocks has great significance in analyzing evolutionary theory of vertebrate and in improving aquaculture industry, analysis of the cyt b gene and the elucidation of the variation of the cyt b gene DNA in different cyprinids prove that cyt b is a useful genetic marker to monitor the variations in the progeny of the crosses.

Amino Acid Sequence↗

Systematic identification of splice variants in human P/Q-type channel alpha1(2.1) subunits: implications for current density and Ca2+-dependent inactivation.

P/Q-type (Ca(v)2.1) calcium channels support a host of Ca2+-driven neuronal functions in the mammalian brain. Alternative splicing of the main alpha1A (alpha1(2.1)) subunit of these channels may thereby represent a rich strategy for tuning the functional profile of diverse neurobiological processes. Here, we applied a recently developed "transcript-scanning" method for systematic determination of splice variant transcripts of the human alpha1(2.1) gene. This screen identified seven loci of variation, which together have never been fully defined in humans. Genomic sequence analysis clarified the splicing mechanisms underlying the observed variation. Electrophysiological characterization and a novel analytical paradigm, termed strength-current analysis, revealed that one focus of variation, involving combinatorial inclusion and exclusion of exons 43 and 44, exerted a primary effect on current amplitude and a corollary effect on Ca2+-dependent channel inactivation. These findings significantly expand the anticipated scope of functional diversity produced by splice variation of P/Q-type channels.

Alternative Splicing↗

Estimation of the number of authentic orphan genes in bacterial genomes.

Genome annotation produces a considerable number of putative proteins lacking sequence similarity to known proteins. These are referred to as "orphans." The proportion of orphan genes varies among genomes, and is independent of genome size. In the present study, we show that the proportion of orphan genes roughly correlates with the isolation index of organisms (IIO), an indicator introduced in the present study, which represents the degree of isolation of a given genome as measured by sequence similarity. However, there are outlier genomes with respect to the linear correlation, consisting of those genomes that may contain excess amounts of orphan genes. Comparisons of genome sequences among closely related strains revealed that some of the annotated genes are not conserved, suggesting that they are ORFs occurring by chance. Exclusion of these non-conserved ORFs within closely related genomes improved the correlation between the proportion of orphan genes and the IIO values. Assuming that the correlation holds in general, this relationship was used to estimate the number of "authentic" orphan genes in a genome. Using this definition of authentic orphan genes, the anomalies arising from over-assignments, e.g., the percentages of structural annotations, were corrected for 16 genomes, including those of five archaea.

Amino Acid Sequence↗

[Bioinformatic analysis of adenoma-normal mucosa SSH library of colon].

We established a colonic adenoma-normal mucosa suppressive subtraction hybridization (SSH) library in 1999. In this study, we wanted to explore the expression profile of all candidate genes in this library. We developed an EST pipeline which contained two in-house software packages, nucleic acid analytical software and GetUni. The nucleic acid analytical software, an integrator of the universal bioinformatics tools including phred, phd2fasta, cross_match, repeatmasker and blast2.0, can blast sequences of differential clones with the downloaded non-redundant nucleotide (NR) database. GetUni can cluster these NR sequences into Unigene via matching with the downloaded Homo Sapiens UniGene database. Sixty-two candidate genes in A-N library were obtained via the high throughput automatic gene expression bioinformatics pipeline. Gene Ontology online analysis revealed that ribosome genes and immunity-regulating genes were the two most common categories in the KEGG or Biocarta Pathway. We also detected the expression of 2 genes with highest hits, Reg4 and FAM46A, by semi-quantitative RT-PCR. Both genes were up-regulated in 10 or 9 out of 10 adenomas in comparison with the paired normal mucosa, respectively. The candidate genes in A-N library would be of great significance in disclosing the molecular mechanism underlying in colonic adenoma initiation and progression.

Adenoma↗

Data mining of molecular dynamics trajectories of nucleic acids.

Analysis, storage, and transfer of molecular dynamic trajectories are becoming the bottleneck of computer simulations. In this paper we discuss different approaches for data mining and data processing of huge trajectory files generated from molecular dynamic simulations of nucleic acids.

Computer Simulation↗

IMGT, the international ImMunoGeneTics database: a new design for immunogenetics data access.

IMGT, the international ImMunoGeneTics database is an integrated database specializing in Immunoglobulins (Ig), T-cell receptors (TcR) and MHC molecules of all vertebrate species, created by Marie-Paule Lefranc, University of Montpellier, CNRS, Montpellier, France (Nucleic Acids Research, Database issue, Vol 26, January 1998). IMGT includes three databases: LIGM-DB (for Ig and TcR), MHC/HLA-DB and IMGT/PRIMER-DB (an Ig, TcR and MHC-related primer database), the last two in development. IMGT comprises expertly annotated sequences and alignment tables. LIGM-DB contains more than 24.000 Immunoglobulin and T cell Receptor sequences from 81 different species. MHC/HLA-DB contains class I and class II Human Leucocyte Antigen alignment tables. An IMGT tool, DNAPLOT, developed for Ig, TcR and MHC sequence analysis, is also available. IMGT goals are to establish a common data access to all immunogenetics data, including nucleotide and protein sequences, oligonucleotide primers, gene maps and other genetic data of Ig, TcR and MHC molecules, from all species, and to provide a graphical user friendly data access. IMGT has important implications in medical research (repertoire in autoimmune diseases, AIDS, leukemias, lymphomas), therapeutical approaches (antibody engineering), genome diversity and genome evolution studies. In this paper, we describe our approach for the data modelisation, the automation of the annotation procedure and control of data quality in LIGM-DB database. IMGT is freely available on the CNUSC WWW server at Montpellier: http://imgt.cnusc.fr: 8104 (contact: Denys.Chaume@cnusc.fr) and on the EBI servers: http://www.ebi.ac.uk/imgt (contact: malik@ebi.ac.uk) and ftp.ebi.ac.uk/pub/databases/imgt. LIGM-DB users are encouraged to report errors or suggestions to giudi@ligm.crbm.cnrs-mop.fr. IMGT initiator and coordinator: Marie-Paule Lefranc, lefranc@ligm.crbm.cnrs-mop.fr. (fax: +33(0)467040231).

Amino Acid Sequence↗

Comparative genomics of the sperm mitochondria-associated cysteine-rich protein gene.

The sperm mitochondrial cysteine-rich protein (SMCP) is a rapidly evolving cysteine- and proline-rich protein that is localized in the mitochondrial capsule and enhances sperm motility. The sequences of the SMCP protein, gene, and mRNA in a variety of mammals have been compared to understand their evolution and regulation. SMCP can now be reliably identified by its tripartite structure including a short amino-terminal segment; a central segment containing short tandem repeats rich in cysteine, proline, glutamine, and lysine; and a C-terminal segment containing no repeats, few cysteines, and a C-terminal lysine. The SMCP gene is located in the epidermal differentiation complex (EDC), a large gene cluster that functions in forming epithelial barriers. Similarities in chromosomal location, molecular function, intron-exon structure, and protein organization argue that SMCP originated from an EDC gene and acquired spermatogenic cell-specific transcriptional and translational regulation and a novel cellular function in sperm motility. The SMCP 5' UTR and 3' UTR contain conserved elements and uORFs that may function in cytoplasmic regulation of gene expression, and the levels of SMCP mRNA in human are much lower than in other mammals, a feature of male-biased expression. The evolution of SMCP has been accompanied by changes in the sequence, number, and length of repeat units, including three alleles in dogs. The major proteins associated with the mitochondrial capsule, SMCP and phospholipid hydroperoxide glutathione peroxidase, provide outstanding examples of changes in cellular function driven by selective pressures on sperm motility, an important determinant of male reproductive success.

3' Untranslated Regions↗

Molecular analysis of human Siglec-8 orthologs relevant to mouse eosinophils: identification of mouse orthologs of Siglec-5 (mSiglec-F) and Siglec-10 (mSiglec-G).

We recently identified a novel human sialic acid binding immunoglobulin-like lectin, Siglec-8, using mRNA from human eosinophils. To search for a mouse Siglec (mSiglec) ortholog of Siglec-8 and other mouse Siglec paralogs, we conducted public database searches with cDNA sequences of human Siglec-5 to -10 and identified two novel mSiglecs. One has significant sequence identity to human Siglec-5 and is a splice variant of mSiglec-F. The other has greatest sequence identity to human Siglec-10 (mSiglec-G). Both mSiglecs have extracellular Ig-like domains and intracellular tyrosine-based motifs. To determine whether these mSiglecs were relevant to mouse eosinophils, RT-PCR and Northern blot analysis were performed. We detected expression of mSiglec-5 (or -F), -10, and -E mRNA in purified mouse eosinophils, but Northern blot data comparing expression in tissues from normal, IL-5 transgenic, and allergen-sensitized and -challenged mice suggest that mSiglec-10 is probably most relevant to mouse eosinophils.

Amino Acid Sequence↗

Suffix-tree analyser (STAN): looking for nucleotidic and peptidic patterns in chromosomes.

SUMMARY: We have developed STAN (suffix-tree analyser), a tool to search for nucleotidic and peptidic patterns within whole chromosomes. Pattern syntax uses a string variable grammar-like formalism which allows the description of complex patterns including ambiguities, insertions/deletions, gaps, repeats and palindromes. STAN is based on a reduction to multipart matching on a suffix-tree data structure and can handle large DNA sequences, whether assembled or not.

Amino Acid Sequence↗

A DNA repair system specific for thermophilic Archaea and bacteria predicted by genomic context analysis.

During a systematic analysis of conserved gene context in prokaryotic genomes, a previously undetected, complex, partially conserved neighborhood consisting of more than 20 genes was discovered in most Archaea (with the exception of Thermoplasma acidophilum and Halobacterium NRC-1) and some bacteria, including the hyperthermophiles Thermotoga maritima and Aquifex aeolicus. The gene composition and gene order in this neighborhood vary greatly between species, but all versions have a stable, conserved core that consists of five genes. One of the core genes encodes a predicted DNA helicase, often fused to a predicted HD-superfamily hydrolase, and another encodes a RecB family exonuclease; three core genes remain uncharacterized, but one of these might encode a nuclease of a new family. Two more genes that belong to this neighborhood and are present in most of the genomes in which the neighborhood was detected encode, respectively, a predicted HD-superfamily hydrolase (possibly a nuclease) of a distinct family and a predicted, novel DNA polymerase. Another characteristic feature of this neighborhood is the expansion of a superfamily of paralogous, uncharacterized proteins, which are encoded by at least 20-30% of the genes in the neighborhood. The functional features of the proteins encoded in this neighborhood suggest that they comprise a previously undetected DNA repair system, which, to our knowledge, is the first repair system largely specific for thermophiles to be identified. This hypothetical repair system might be functionally analogous to the bacterial-eukaryotic system of translesion, mutagenic repair whose central components are DNA polymerases of the UmuC-DinB-Rad30-Rev1 superfamily, which typically are missing in thermophiles.

Amino Acid Sequence↗

Delila system tools.

We introduce three new computer programs and associated tools of the Delila nucleic-acid sequence analysis system. The first program, Module, allows rapid transportation of new sequence analysis tools between scientists using different computers. The second program, DBpull, allows efficient access to the large nucleic-acid sequence databases being collected in the United States and Europe. The third program, Encode, provides a flexible way to process sequence data for analysis by other programs.

Base Sequence↗

The FxRxHrS motif: a conserved region essential for DNA binding of the VirR response regulator from Clostridium perfringens.

The VirSR two-component signal transduction pathway regulates virulence and toxin production in Clostridium perfringens, the causative agent of gas gangrene. The response regulator, VirR, binds to repeat sequences located upstream of the promoter and is directly responsible for the transcriptional activation of pfoA, the structural gene for the cholesterol-dependent cytolysin, perfringolysin O. Comparative sequence analysis of the 236 amino acid residue VirR protein revealed a two-domain structure: a typical N-terminal response regulator domain and an uncharacterised C-terminal domain. Database searching revealed that over 40 other proteins, many of which appeared to be response regulators or transcriptional activators, had homology with the VirR C-terminal domain (VirRc). Multiple sequence alignment of this VirRc family revealed a highly conserved region that was designated the FxRxHrS motif. By deletion analysis this motif was shown to be essential for the functional integrity of the VirR protein. Alanine scanning mutagenesis and subsequent phenotypic analysis indicated that conserved residues located within the motif were required for activity. These residues extended from L179 to N194. More detailed site-directed mutagenesis showed that amino acid residues R186, H188 and S190 were essential for activity since even conservative substitutions in these positions resulted in non-functional proteins. Three of the mutant proteins, R186K, S190A and S190C, were purified and shown by in vitro gel shift analysis to be unable to bind to the specific target DNA with the same efficiency as the wild-type protein. These data reveal for the first time that VirRc functions as a DNA binding domain in which the highly conserved FxRxHrS motif has a functional role. These studies have important implications for this new family of transcriptional factors since they imply that the conserved FxRxHrS motif may be involved in DNA binding in all of these proteins, irrespective of their biological role.

Amino Acid Motifs↗

Characterization and comparative analysis of the EGLN gene family.

Rat Sm-20 is a homologue of the Caenorhabditis elegans gene egl-9 and has been implicated in the regulation of growth, differentiation and apoptosis in muscle and nerve cells. Null mutants in egl-9 result in a complete tolerance to an otherwise lethal toxin produced by Pseudomonas aeruginosa. This study describes the conserved Egl-Nine (EGLN) gene family of which rat SM-20 and C. elegans Egl-9 are members and characterizes the mouse and human homologues. Each of the human genes (EGLN1, EGLN2 and EGLN3) are of a conserved genomic structure consisting of five coding exons. Phylogenetic analysis and domain organization show that EGLN1 represents the ancestral form of the gene family and that EGLN3 is the human orthologue of rat Sm-20. The previously observed mitochondrial targeting of rat SM-20 is unlikely to be a general feature of the protein family and may be a feature specific to rats. An EGLN gene is unexpectedly found in the genome of P. aeruginosa, a bacterium known to produce a toxin that acts through the Egl-9 protein. The pathogenic bacterium Vibrio cholerae is also shown to have an EGLN gene suggesting that it is an important pathogenicity factor. These results provide new insights into host-pathogen interactions and a basis for further functional characterization of the gene family and resolve discrepancies in annotation between gene family members.

Amino Acid Sequence↗

Mutations in the fukutin-related protein gene (FKRP) cause a form of congenital muscular dystrophy with secondary laminin alpha2 deficiency and abnormal glycosylation of alpha-dystroglycan.

The congenital muscular dystrophies (CMD) are a heterogeneous group of autosomal recessive disorders presenting in infancy with muscle weakness, contractures, and dystrophic changes on skeletal-muscle biopsy. Structural brain defects, with or without mental retardation, are additional features of several CMD syndromes. Approximately 40% of patients with CMD have a primary deficiency (MDC1A) of the laminin alpha2 chain of merosin (laminin-2) due to mutations in the LAMA2 gene. In addition, a secondary deficiency of laminin alpha2 is apparent in some CMD syndromes, including MDC1B, which is mapped to chromosome 1q42, and both muscle-eye-brain disease (MEB) and Fukuyama CMD (FCMD), two forms with severe brain involvement. The FCMD gene encodes a protein of unknown function, fukutin, though sequence analysis predicts it to be a phosphoryl-ligand transferase. Here we identify the gene for a new member of the fukutin protein family (fukutin related protein [FKRP]), mapping to human chromosome 19q13.3. We report the genomic organization of the FKRP gene and its pattern of tissue expression. Mutations in the FKRP gene have been identified in seven families with CMD characterized by disease onset in the first weeks of life and a severe phenotype with inability to walk, muscle hypertrophy, marked elevation of serum creatine kinase, and normal brain structure and function. Affected individuals had a secondary deficiency of laminin alpha2 expression. In addition, they had both a marked decrease in immunostaining of muscle alpha-dystroglycan and a reduction in its molecular weight on western blot analysis. We suggest these abnormalities of alpha-dystroglycan are caused by its defective glycosylation and are integral to the pathology seen in MDC1C.

Adult↗

Comparative genomics on BMP4 orthologs.

Bone morphogenetic proteins (BMPs) are implicated in cell-fate determination of embryonic stem (ES) cells and cancer cells. GREM1 (CKTSF1B1 or DAND2) and CER1 (Cerberus 1 or DAND4) are cysteine knot superfamily proteins, functioning as secreted-type BMP antagonists. BMP4 is preferentially expressed in diffuse-type gastric cancer cells. Here, vertebrate BMP4 orthologs were identified and characterized by using bioinformatics for comparative proteomics and comparative genomics analyses. Baboon BMP4 gene within AC153751.2 genome sequence encoded a 408-aa protein, showing A152V and S298P amino-acid substitutions compared with human BMP4. Cow Bmp4, bat Bmp4 and zebrafish bmp4 genes were located within AC149774.2, AC156788.2 and CR391996.2 genome sequences, respectively. Human BMP4 showed 99.5%, 98.0%, 97.8%, 97.1%, 96.3%, 83.3% and 71.1% total-amino-acid identity with baboon BMP4, cow Bmp4, bat Bmp4, mouse Bmp4, rat Bmp4, chicken bmp4 and zebrafish bmp4, respectively. Human BMP4 gene was found consisting of six exons, including novel exon 1C, and known exons 1 (1A or I), 1B (II), 2 (III), 3 (IV) and 4 (V). Forty human BMP4 ESTs started from exon 1, seven from intron 1 (5'-flanking region of exon 2), and two from exon 1C. Fourteen mouse Bmp4 ESTs started from exon 1, and one from intron 1. The 5'-flanking region of exon 1 and exon 1 itself, but not exons 1C and 1B, were well conserved between human BMP4 and rodent Bmp4 genes. The major promoter region of human BMP4 and rodent Bmp4 genes were located within the 5'-flanking region of exon 1. FOXA2, OLF1, and MYC-binding sites were conserved among the major promoter region of human, baboon, cow, bat, mouse and rat BMP4 orthologs.

Amino Acid Sequence↗

A new method based on entropy theory for genomic sequence analysis.

We have refined entropy theory to explore the meaning of the increasing sequence data on nucleic acids and proteins more conveniently. The concept of selection constraint was not introduced, only the analyzed sequences themselves were considered. The refined theory serves as a basis for deriving a method to analyze non-coding regions (NCRs) as well as coding regions. Positions with maximal entropy might play the most important role in genome functions as opposed to positions with minimal entropy. This method was tested in the well-characterized coding regions of 12 strains of Classical Swine Fever Virus (CSFV) and non-coding regions of 20 strains of CSFV. It is suitable to analyze nucleic acid sequences of a complete genome and to detect sensitive positions for mutagenesis. As such, the method serves to formulate the basis for elucidating the functional mechanism.

Algorithms↗

Improving functional annotation of non-synonomous SNPs with information theory.

Automated functional annotation of nsSNPs requires that amino-acid residue changes are represented by a set of descriptive features, such as evolutionary conservation, side-chain volume change, effect on ligand-binding, and residue structural rigidity. Identifying the most informative combinations of features is critical to the success of a computational prediction method. We rank 32 features according to their mutual information with functional effects of amino-acid substitutions, as measured by in vivo assays. In addition, we use a greedy algorithm to identify a subset of highly informative features. The method is simple to implement and provides a quantitative measure for selecting the best predictive features given a set of features that a human expert believes to be informative. We demonstrate the usefulness of the selected highly informative features by cross-validated tests of a computational classifier, a support vector machine (SVM). The SVM's classification accuracy is highly correlated with the ranking of the input features by their mutual information. Two features describing the solvent accessibility of "wild-type" and "mutant" amino-acid residues and one evolutionary feature based on superfamily-level multiple alignments produce comparable overall accuracy and 6% fewer false positives than a 32-feature set that considers physiochemical properties of amino acids, protein electrostatics, amino-acid residue flexibility, and binding interactions.

Analysis of Variance↗

Identification and characterization of a human DNA glycosylase for repair of modified bases in oxidatively damaged DNA.

8-oxoguanine (8-oxoG), ring-opened purines (formamidopyrimidines or Fapys), and other oxidized DNA base lesions generated by reactive oxygen species are often mutagenic and toxic, and have been implicated in the etiology of many diseases, including cancer, and in aging. Repair of these lesions in all organisms occurs primarily via the DNA base excision repair pathway, initiated with their excision by DNA glycosylase/AP lyases, which are of two classes. One class utilizes an internal Lys residue as the active site nucleophile, and includes Escherichia coli Nth and both known mammalian DNA glycosylase/AP lyases, namely, OGG1 and NTH1. E. coli MutM and its paralog Nei, which comprise the second class, use N-terminal Pro as the active site. Here, we report the presence of two human orthologs of E. coli mutM nei genes in the human genome database, and characterize one of their products. Based on the substrate preference, we have named it NEH1 (Nei homolog). The 44-kDa, wild-type recombinant NEH1, purified to homogeneity from E. coli, excises Fapys from damaged DNA, and oxidized pyrimidines and 8-oxoG from oligodeoxynucleotides. Inactivation of the enzyme because of either deletion of N-terminal Pro or Histag fusion at the N terminus supports the role of N-terminal Pro as its active site. The tissue-specific levels of NEH1 and OGG1 mRNAs are distinct, and S phase-specific increase in NEH1 at both RNA and protein levels suggests that NEH1 is involved in replication-associated repair of oxidized bases.

Amino Acid Sequence↗