PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Integrating mutation data and structural analysis of the TP53 tumor-suppressor protein.

TP53 encodes p53, which is a nuclear phosphoprotein with cancer-inhibiting properties. In response to DNA damage, p53 is activated and mediates a set of antiproliferative responses including cell-cycle arrest and apoptosis. Mutations in the TP53 gene are associated with more than 50% of human cancers, and 90% of these affect p53-DNA interactions, resulting in a partial or complete loss of transactivation functions. These mutations affect the structural integrity and/or p53-DNA interactions, leading to the partial or complete loss of the protein's function. We report here the results of a systematic automated analysis of the effects of p53 mutations on the structure of the core domain of the protein. We found that 304 of the 882 (34.4%) distinct mutations reported in the core domain can be explained in structural terms by their predicted effects on protein folding or on protein-DNA contacts. The proportion of "explained" mutations increased to 55.6% when substitutions of evolutionary conserved amino acids were included. The automated method of structural analysis developed here may be applied to other frequently mutated gene mutations such as dystrophin, BRCA1, and G6PD.

Amino Acid Substitution↗

Diversification of ftsZ during early land plant evolution.

The plastid division proteins FtsZ are encoded by a small nuclear gene family in land plants. Although it has been shown for some of the gene products that they are imported into plastids and function in plastid division, the evolution and function of this gene family and their products remain to be unraveled. Here we present two new ftsZ genes from the moss Physcomitrella patens and compare the genomic structure of members of the two plant ftsZ gene families. Comparison of sequence features and phylogenetic analyses confirm the presence of two clusters of paralogues in land plants and demonstrate that these genes were duplicated before the divergence of mosses, ferns and seed plants.

Amino Acid Sequence↗

Defensin gene family in Medicago truncatula: structure, expression and induction by signal molecules.

A large gene family encoding the putative cysteine-rich defensins was discovered in Medicago truncatula. Sixteen members of the family were identified by screening a cloned seed defensin from M. sativa (Gao et al. 2000) against the Institute for Genomic Research's (TIGR) M. truncatula gene index (MtGI version 7). Based on the comparison of their amino acid sequences, M. truncatula defensins fell arbitrarily into three classes displaying extensive sequence divergence outside of the eight canonical cysteine residues. The presence of Class II defensins is reported for the first time in a legume plant. In silico as well as Northern blot and RT-PCR analyses indicated these genes were expressed in a variety of tissues including leaves, flowers, developing pods, mature seed and roots. The expression of these genes was differentially induced in response to a variety of biotic and abiotic stimuli. For the first time, a defensin gene (TC77480) was shown to be induced in roots in response to infection by the mycorrhizal fungus, Glomus versiforme. Northern blot analysis indicated that the tissue-specific expression patterns of the cloned Def1 and Def2 genes differed substantially between M. truncatula and M. sativa. Furthermore, the induction profiles of the Def1 and Def2 genes in response to the signaling molecules methyl jasmonate, ethylene and salicylic acid differed markedly between these two legumes.

Acetates↗

Cloning of alkaline sphingomyelinase from rat intestinal mucosa and adjusting of the hypothetical protein XP_221184 in GenBank.

Intestinal alkaline sphingomyelinase (alk-SMase) digests sphingomyelin and the process may influence colonic tumorigenesis and cholesterol absorption. We recently identified the gene of human alk-SMase and cloned the cDNA. Cross-species screening of homology in GenBank found a hypothetical rat protein, XP_221184, with 491 amino acid residues, which shares 73% identity with human alk-SMase. Based on the cDNA sequence of this protein, we cloned a cDNA from rat intestinal mucosa by RT-PCR. The cloned cDNA encodes a protein with 439 amino acid residues and higher (85%) identity with human alk-SMase. The cloned cDNA differed from the XP_221184 cDNA in splice sites linking exons 2 and 3, and exons 3 and 4, respectively. In the sequence of the cloned protein, the predicted activity motif, sphingomyelin binding sites, and potential glycosylation sites in human alk-SMase are all conserved. To confirm the cloned protein is the real form of alk-SMase, native alk-SMase was purified from rat intestine and subjected to proteolytic digestion followed by matrix-assisted laser desorption/ionization (MALDI) mass spectrometry and electrospray ionization (ESI) tandem mass spectrometry. Seven tryptic peptides were found to match the cloned protein sequence. Transient expression of the cloned cDNA linked with a myc tag in COS-7 cells demonstrated high SMase activity, with an optimal pH at 9.0 and a specific dependence on taurocholate and taurochenodeoxycholate. The expressed protein reacted with both anti-myc and anti-human alk-SMase antibodies. Northern blotting of rat tissues revealed high levels of mRNA in jejunum but not in other tissues. In conclusion, we cloned rat alk-SMase cDNA from rat intestine, adjusted the putative rat alk-SMase protein in GenBank, and confirmed the specific expression of the gene in the small intestine.

Amino Acid Sequence↗

Regulation of novel members of the Arabidopsis thaliana CCAAT-binding nuclear factor Y subunits.

Nuclear factor Y (NF-Y) is a highly conserved trimeric activator that recognizes with high specificity and affinity the widespread CCAAT box promoter element. We previously cloned the genes of 23 NF-Y genes of Arabidopsis thaliana (Gene 264 (2001) 173). Now that the Arabidopsis genome sequencing project is complete, we present the cloning, alignments and expression profiles of the remaining six genes coding for the three NF-Y subunits. Consistent with our previous reports, most of the new members of the three subunits show a unique tissue-specific pattern, while another AtNF-YC9 is rather ubiquitous.

Amino Acid Sequence↗

Using proteomics to mine genome sequences.

We present a method for mining unannotated or annotated genome sequences with proteomic data to identify open reading frames. The region of a genome coding for a protein sequence is identified by using information from the analysis of proteins and peptides with MALDI-TOF mass spectrometry. The raw genome sequence or any unassembled contigs of an organism are theoretically cleaved into a number of equal sized but overlapping fragments, and these are then translated in all six frames into a series of virtual proteins. Each virtual protein is then subjected to a theoretical enzymatic digestion. Standard proteomic sample preparation methods are used to separate, array, and digest the proteins of interest to peptides. The masses of the resulting peptides are measured using mass spectrometry and compared to the theoretical peptide masses of the virtual proteins. The region of the genome responsible for coding for a particular protein can then be identified when there are a large number of hits between peptides from the protein and peptides from the virtual protein. The method makes no assumptions about the location of a protein in a particular gene sequence or the positions or types of start and stop codons. To illustrate this approach, all 773 proteins of Pseudomonas aeruginosa contained in SWISS-PROT were used to theoretically test the method and optimize parameters. Increasing the size of the virtual proteins results in an overall improvement in the ability to detect the coding region, at the cost of decreasing the sensitivity of the method for smaller proteins. Increasing the minimum number of matching peptides, lowering the mass error tolerance, or increasing the signal-to-noise ratio of the simulated mass spectrum, improves the ability to detect coding regions. The method is further demonstrated on experimental data from Mycobacterium tuberculosis and is also shown to work with eukaryotic organisms (e.g., Homo sapiens).

Amino Acid Sequence↗

Identification of a novel cell cycle regulated gene, HURP, overexpressed in human hepatocellular carcinoma.

An analytic strategy was followed to identify putative regulatory genes during the development of human hepatocellular carcinoma (HCC). This strategy employed a bioinformatics analysis that used a database search to identify genes, which are differentially expressed in human HCC and are also under cell cycle regulation. A novel cell cycle regulated gene (HURP) that is overexpressed in HCC was identified. Full-length cDNAs encoding the human and mouse HURP genes were isolated. They share 72 and 61% identity at the nucleotide level and amino-acid level, respectively. Endogenous levels of HURP mRNA were found to be tightly regulated during cell cycle progression as illustrated by its elevated expression in the G(2)/M phase of synchronized HeLa cells and in regenerating mouse liver after partial hepatectomy. Immunofluorescence studies revealed that hepatoma up-regulated protein (HURP) localizes to the spindle poles during mitosis. Overexpression of HURP in 293T cells resulted in an enhanced cell growth at low serum levels and at polyhema-based, anchorage-independent growth assay. Taken together, these results strongly suggest that HURP is a potential novel cell cycle regulator that may play a role in the carcinogenesis of human cancer cells.

Amino Acid Sequence↗

Understanding human disease mutations through the use of interspecific genetic variation.

Data on replacement mutations in genes of disease patients exist in a variety of online resources. In addition, genome sequencing projects and individual gene sequencing efforts have led to the identification of disease gene homologs in diverse metazoan species. The availability of these two types of information provides unique opportunities to investigate factors that are important in the development of genetically based disease by contrasting long and short-term molecular evolutionary patterns. Therefore, we conducted an analysis of disease-associated human genetic variation for seven disease genes: the cystic fibrosis transmembrane conductance regulator, glucose-6-phosphate dehydrogenase, the neural cell adhesion molecule L1, phenylalanine hydroxylase, paired box 6, the X-linked retinoschisis gene and TSC2/tuberin. Our analyses indicate that disease mutations show definite patterns when examined from an evolutionary perspective. Human replacement mutations resulting in disease are overabundant at amino acid positions most conserved throughout the long-term history of metazoans. In contrast, human polymorphic replacement mutations and silent mutations are randomly distributed across sites with respect to the level of conservation of amino acid sites within genes. Furthermore, disease-causing amino acid changes are of types usually not observed among species. Using Grantham's chemical difference matrix, we find that amino acid changes observed in disease patients are far more radical than the variation found among species and in non-diseased humans. Overall, our results demonstrate the usefulness of evolutionary analyses for understanding patterns of human disease mutations and underscore the biomedical significance of sequence data currently being generated from various model organism genome sequencing projects.

Amino Acids↗

Evolutionary pressures on apicoplast transit peptides.

Malaria parasites (species of the genus Plasmodium) harbor a relict chloroplast (the apicoplast) that is the target of novel antimalarials. Numerous nuclear-encoded proteins are translocated into the apicoplast courtesy of a bipartite N-terminal extension. The first component of the bipartite leader resembles a standard signal peptide present at the N-terminus of secreted proteins that enter the endomembrane system. Analysis of the second portion of the bipartite leaders of P. falciparum, the so-called transit peptide, indicates similarities to plant transit peptides, although the amino acid composition of P. falciparum transit peptides shows a strong bias, which we rationalize by the extraordinarily high AT content of P. falciparum DNA. 786 plastid transit peptides were also examined from several other apicomplexan parasites, as well as from angiosperm plants. In each case, amino acid biases were correlated with nucleotide AT content. A comparison of a spectrum of organisms containing primary and secondary plastids also revealed features unique to secondary plastid transit peptides. These unusual features are explained in the context of secondary plastid trafficking via the endomembrane system.

Amino Acids↗

Update of AMmtDB: a database of multi-aligned Metazoa mitochondrial DNA sequences.

The AMmtDB database (http://bighost.area.ba.cnr.it/mitochondriome) has been updated by collecting the multi-aligned sequences of Chordata and Invertebrata mitochondrial genes coding for proteins and tRNAs. Links to the multi-aligned mtDNA intraspecies variants, collected in VarMmtDB at the Mitochondriome web site, have been introduced. The genes coding for proteins are multi-aligned based on the translated sequences and both the nucleotide and amino acid multi-alignments are provided. AMmtDB data selected through SRS can be viewed and managed using GeneDoc or other programs for the management of multi-aligned data depending on the user's operative system. The multiple alignments have been produced with CLUSTALW and PILEUP programs and then carefully optimized manually.

Amino Acid Sequence↗

Comparative genomics on mammalian Fgf6-Fgf23 locus.

CCND2-C12orf5-FGF23-FGF6 locus at human chromosome 12p13.32 and CCND1-ORAOV1-FGF19-FGF4 locus at human chromosome 11q13.3 are paralogous regions (paralogons) within the human genome. FGF23 is the causative factor for tumor-induced osteomalacia (TIO), a paraneoplastic disorder characterized by hypophosphatemia and skeletal undermineralization, and also for autosomal dominant hypophosphatemic rickets (ADHR). Here, rat Fgf6 and Fgf23 complete coding sequences were determined by using bioinformatics. Rat Fgf6 and Fgf23 genes, consisting of three exons, were located within AC103292.6 rat genome sequence. Rat Fgf6 and Fgf23 genes were clustered in tail-to-head manner with an interval of about 52 kb. Human FGF6 and FGF23 genes were clustered in tail-to-head manner with an interval of about 54 kb. Intergenic conserved region (IGCR) within the FGF6-FGF23 gene cluster was identified based on the evolutionary conservation. Human FGF6-FGF23 IGCR (nucleotide position 111648-112242 of AC008012.8 genome sequence) and rat Fgf6-Fgf23 IGCR (nucleotide position 156318-156894 of AC103292.6 genome sequence) showed 77.6% total nucleotide identity. CP2, E47, CREB and PAX4 binding sites were conserved among human FGF6, rat Fgf6, and mouse Fgf6 promoters. GATA and E47 binding sites were conserved among human FGF23, rat Fgf23, and mouse Fgf23 promoters. Because mouse Fgf23 mRNA was expressed in dendritic cells and activated spleen, tumor infiltrating dendritic cells are candidate sources of FGF23 secretion in TIO patients. This is the first report on comparative genomics analyses on human FGF6-FGF23 gene cluster and rodents Fgf6-Fgf23 gene cluster.

Amino Acid Sequence↗

An in silico mining for simple sequence repeats from expressed sequence tags of zebrafish, medaka, Fundulus, and Xiphophorus.

Teleost fish genome projects involving model species are resulting in a rapid accumulation of genomic and expressed DNA sequences in public databases. The expressed sequence tags (ESTs) collected in the databases can be mined for the analysis of both structural and functional genomics. In this study, we in silico analyzed 49,430 unigenes representing a total of 692,654 ESTs from four model fish for their potential use in developing simple sequence repeats (SSRs), or microsatellites. After bioinformatical mining, a total of 3,018 EST derived SSRs (EST-SSRs) were identified for 2,335 SSR containing ESTs (SSR-ESTs). The frequency of identified SSR-ESTs ranged from 1.5% for Xiphophorus to 7.3% for zebrafish. The dinucleotide repeat motif is the most abundant SSR, accounting for 47%, 52%, 64%, and 78% for medaka, Fundulus, zebrafish, and Xiphophorus, respectively. Simulation analysis suggests that a majority of these EST-SSRs have sufficient flanking sequences for polymerase chain reaction (PCR) primer design. Comparative DNA sequence analyses of SSR-ESTs identified several cross-species SSRs and sequences that may be used as cross-reference genes in comparative studies. For example, the flanking sequences of one SSR (CTG)n within the pituitary tumor-transforming gene (PTTG) 1 interacting protein (PTTGIP), showed conservation spanning the medaka, Fundulus, human, and mouse genomes. This study provides a large body of information on EST-SSRs that can be useful for the development of polymorphic markers, gene mapping, and comparative genome analysis. Functional analysis of these SSR-ESTs may reveal their role in metabolism and gene evolution of these model species.

Amino Acid Sequence↗

Identification and expression of the amphioxus Cks1 gene.

A full-length Cks1 homologue gene, AmphiCks1, was identified in amphioxus, Branchiostoma belcheri tsingtauense. Sequence characteristics, phylogeny and patterns of expression during embryonic and larval development were established. The protein predicted from AmphiCks1 showed high sequence identity with vertebrate and invertebrate homologues. Protein structural studies and phylogenetic analysis suggested that Cks homologues are evolutionarily conserved. The AmphiCks1 transcript was detected in most early developmental stages by northern blotting and whole-mount in situ hybridization, suggesting a role for the gene in cell division.

Amino Acid Sequence↗

Identification of a novel class of annexin genes.

The annexins are a family of calcium- and phospholipid-binding proteins that have been widely studied in animals. Investigation of annexins in the fungus Aspergillus fumigatus identified a novel annexin-like gene (ANXC4) as well as two conventional annexins (ANXC3.1 and ANXC3.2). The genes were initially identified by bioinformatics, and sequences were then determined experimentally. Reverse transcription polymerase chain reaction indicated that all three genes were expressed. ANXC4 lacked calcium-binding consensus sequences and had a 553 residue N-terminal tail. However, bioinformatics indicated that ANXC4 is an annexin and homologues were identified in other filamentous fungi. ANXC4 therefore represents a new grouping within the annexin family.

Amino Acid Sequence↗

Molecular phylogenetics of the RrmJ/fibrillarin superfamily of ribose 2'-O-methyltransferases.

Recent analyses identified a putative catalytic tetrad K-D-K-E common to several families of site-specific methyltransferases (MTases) that modify 2'-hydroxyl groups of ribose in mRNA, rRNA and tRNA (designated the RrmJ class after one of the structurally characterized members; 1eiz in Protein Data Bank) [Genome Biol. 2(9) (2001) 38]. Subsequently, three residues of the tetrad (K-D-K) were shown to be essential for catalysis in RrmJ [J. Biol. Chem. 277 (2002) 41978]. Here, we report identification of a similar conserved tetrad (K-D-K-H) in the family of snoRNA-guided ribose 2'-O-MTases related to fibrillarin (represented by the Mj0697 protein structure; 1fbn in PDB). The corresponding functional groups of putative catalytic tetrads of RrmJ and Mj0697 may be superimposed in space. However, one of the invariant residues (K(164) in RrmJ and K(179) in Mj0697) is observed in two distinct locations in the primary sequence, suggesting an interesting case of 'migration' of the conserved side chain within the framework of the active site. RrmJ and Mj0697 sequences were used as starting points to carry out comprehensive sequence database searches, resulting in identification of a similar conserved tetrad (and hence, prediction of a ribose 2'-O-specificity) in several families of putative MTases, including TlyA hemolysins, novel proteins from Trypanosoma, and large multidomain proteins from Flaviviriruses, Nidoviruses, and Alphaviruses. The results of our analysis of phylogenetic relationships in the RrmJ/fibrillarin superfamily provide insight into the evolution of site-specific and snoRNA-guided ribose 2'-O-MTases from a common ancestor.

Amino Acid Sequence↗

The human SPANX multigene family: genomic organization, alignment and expression in male germ cells and tumor cell lines.

Multigenicity is one of the features of cancer/testis-associated genes. In the present study we analyzed the number and expression of genes of the SPANX(CTp11) family of cancer/testis-associated genes. Genomic database analysis, next to the four previously described SPANX genes, revealed the presence of a novel gene: SPANXE. Moreover, we detected an allelic variant of SPANXB resulting in one amino acid substitution in the encoded protein: SPANXB'. Most SPANX genes are present on contig NT_011574 located at Xq26.3-Xq27.1. Based on expressed sequence tag databases and RT-PCR analysis three additional novel SPANX sequences were identified, though not represented so far in the human genome sequence. Sequence alignments justify a subdivision of this gene family based on the absence (SPANXA-likes) or presence (SPANXB) of an 18 base pair sequence stretch in the open reading frame. The alignments also reveal an unusually high level (99%) of intron homology. Furthermore, the nucleotide variations in the open reading frame almost all lead to amino acid substitutions. Southern blot and database analyses indicate that SPANX sequences are exclusively present in primates. With RT-PCR analysis on human sperm cell precursors and tumor cell lines most family members could be detected. SPANXB was only found in sperm cell precursors and could not be detected in the tumor cell lines tested. Overall SPANXA was the most frequently expressed SPANX variant in melanoma and glioblastoma cell lines.

Amino Acid Sequence↗

Analysis of expressed sequence tags (EST) obtained from common carp, Cyprinus carpio L., head kidney cells after stimulation by two mitogens, lipopolysaccharide and concanavalin-A.

A representative cDNA library from mRNA obtained from lipopolysaccharide and concanavalin-A-induced head kidney cells of carp, Cyprinus carpio, was constructed. Two hundred single pass and partially sequenced clones (AU183343 to AU183542) were generated from expressed sequence tags (ESTs) and these were searched for homology in the DDBJ/GENBANK with blastN and blastX programs. Clones matching known genes were classified according to their function and distribution. One hundred and twenty-nine genes showed homology with known genes in databases, whereas 71 (35.5%) clones did not show any significant homology to sequences in the public database. Known genes also showed homology to fish genes deposited in the database. Twenty-two clones (11%), encoding 16 different sequences, were identified as putative biodefense and oncogenes, associated with an immune response. High expression of lysozyme (3%) was detected. Putatively identified biodefense-related sequences such as Lectin type 2, MHC class II invariant chain, mcl-1a and lysozyme were aligned with known homologues from the database and the percentage identity determined. A time course evaluation of gene expression due to mitogen stimulation by RT-PCR revealed the above mentioned gene homologues were switched on early during the cell proliferation.

Amino Acid Sequence↗