PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Functional genomic analysis of Arabidopsis thaliana glycoside hydrolase family 1.

In plants, Glycoside Hydrolase (GH) Family 1 beta -glycosidases are believed to play important roles in many diverse processes including chemical defense against herbivory, lignification, hydrolysis of cell wall-derived oligosaccharides during germination, and control of active phytohormone levels. Completion of the Arabidopsis thaliana genome sequencing project has enabled us, for the first time, to determine the total number of Family 1 members in a higher plant. Reiterative database searches revealed a multigene family of 48 members that includes eight probable pseudogenes. Manual reannotation and analysis of the entire family were undertaken to rectify existing misannotations and identify phylogenetic relationships among family members. Forty-seven members (designated BGLU1 through BGLU47 ) share a common evolutionary origin and were subdivided into approximately 10 subfamilies based on phylogenetic analysis and consideration of intron-exon organizations. The forty-eighth member of this family ( At3g06510; sfr2 ) is a beta -glucosidase-like gene that belongs to a distinct lineage. Information pertaining to expression patterns and potential functions of Arabidopsis GH Family 1 members is presented. To determine the biological function of all family members, we intend to investigate the substrate specificity of each mature hydrolase after its heterologous expression in the Pichia pastoris expression system. To test the validity of this approach, the BGLU44 -encoded hydrolase was expressed in P. pastoris and purified to homogeneity. When tested against a wide range of natural and synthetic substrates, this enzyme showed a preference for beta -mannosides including 1,4- beta -D-mannooligosaccharides, suggesting that it may be involved in A. thaliana in degradation of mannans, galactomannans, or glucogalactomannans. Supporting this notion, BGLU44 shared high sequence identity and similar gene organization with tomato endosperm beta -mannosidase and barley seed beta -glucosidase/ beta -mannosidase BGQ60.

Arabidopsis↗

Ammonium transporter genes in Chlamydomonas: the nitrate-specific regulatory gene Nit2 is involved in Amt1;1 expression.

Ammonium transport is a key process in nitrogen metabolism. In the green alga Chlamydomonas, we have characterized molecularly the largest family of ammonium transporters (AMT1) so far described consisting of eight members. CrAmt1 genes have an interesting transcript structure with some very small exons. Differential expression patterns were found for each CrAmt1 gene in response to the nitrogen source by using Real Time PCR. These expression patterns were similar under high and low CO2 atmosphere. CrAmt1;1 expression was characterized in detail. It was repressed in both ammonium and nitrate medium, and strongly expressed in nitrogen-free media. Treatment with a Glutamine synthetase inhibitor released partially repression in ammonium and nitrate suggesting that ammonium and its derivatives participate in the observed repressing effects. By studying CrAmt1;1 expression in mutants deficient at different steps of the nitrate assimilation pathway, it has been shown that nitrate has a double negative effect on this gene expression; one related to its reduction to ammonium, and a second one by itself. This second effect of nitrate was dependent on the functionality of the regulatory gene Nit2, specific for nitrate assimilation. Thus, NIT2 would have a dual role on gene expression: the well-known positive one on nitrate assimilation and a novel negative one on Amt1;1 regulation.

Algal Proteins↗

Transcriptional co-regulation of secondary metabolism enzymes in Arabidopsis: functional and evolutionary implications.

The combined knowledge of the Arabidopsis genome and transcriptome now allows to get an integrated view of the dynamics and evolution of metabolic pathways in plants. We used publicly available sets of microarray data obtained in a wide range of different stress and developmental conditions to investigate the co-expression of genes encoding enzymes of secondary metabolism pathways, in particular indoles, phenylpropanoids, and flavonoids. We performed hierarchical clustering of gene expression profiles and found that major enzymes of each pathway display a clear and robust co-expression throughout all the conditions studied. Moreover, detailed analysis evidenced that some genes display co-regulation in particular physiological conditions only, certainly reflecting their modular recruitment into stress- or developmentally regulated biosynthetic pathways. The combination of these microarray data with sequence analysis allows to draw very precise hypotheses on the function of otherwise uncharacterized genes. To illustrate this approach, we focused our analysis on secondary metabolism glycosyltransferases (UGTs), a multigenic family involved in the conjugation of small molecules to sugars like glucose. We propose that UGT74B1 and UGT74C1 may be involved in aromatic and aliphatic glucosinolates synthesis, respectively. We also suggest that UGT75C1 may function as an anthocyanin-5-O-glucosyltransferase in planta. Therefore, this data-mining approach appears very powerful for the functional prediction of unknown genes, and could be transposed to virtually any other gene family. Finally, we suggest that analysis of expression pattern divergence of duplicated genes also provides some insight into the mechanisms of metabolic pathway evolution.

Arabidopsis↗

Transcription factors in rice: a genome-wide comparative analysis between monocots and eudicots.

It is not known how representative the Arabidopsis thaliana complement of transcription factors (TFs) is of other plants. The availability of rice (Oryza sativa) genome sequences makes possible a comparative analysis of TFs between monocots and eudicots, the two major monophyletic groups of angiosperms. Here, we identified 1611 TF genes that belong to 37 gene families in rice, comparable to the 1510 in Arabidopsis. Several gene subfamilies, but no families, were found to be lineage-specific. Phylogenetic analyses indicated that nearly half of the TF genes form clear orthologous pairs or groups, which were derived from 383 ancestral genes in the common ancestor of rice and Arabidopsis. Investigating gene duplication mechanisms revealed twelve pairs of large intragenomic duplicated blocks, which account for more than 40% of the rice genome. About 60% of the duplicated TF genes have been retained on duplicated segments. Functional conservation and diversification of TFs across monocot and eudicot lineages are discussed.

Arabidopsis↗

From single genes to co-expression networks: extracting knowledge from barley functional genomics.

The paper reports an 'in silico' approach to gene expression analysis based on a barley gene co-expression network resulting from the study of several publicly available cDNA libraries. The work is an application of Systems Biology to plant science: at the end of the computational step we identified groups of potentially related genes. The communities of co-expressed genes constructed from the network are remarkably characterized from the functional point of view, as shown by the statistical analysis of the Gene Ontology annotations of their members. Experimental, lab-based testing has been carried out to check the relationship between network and biological properties and to identify and suggest effective strategies of information extraction from the network-derived data.

Computational Biology↗

A large-scale collection of phenotypic data describing an insertional mutant population to facilitate functional analysis of rice genes.

In order to facilitate the functional analysis of rice genes, we produced about 50,000 insertion lines with the endogenous retrotransposon Tos17. Phenotypes of these lines in the M2 generation were observed in the field and characterized based on 53 phenotype descriptors. Nearly half of the lines showed more than one mutant phenotype. The most frequently observed phenotype was low fertility, followed by dwarfism. Phenotype data with photographs of each line are stored in the Tos17 mutant panel web-based database with a dataset of sequences flanking Tos17 insertion points in the rice genome (http://tos.nias.affrc.go.jp/). This combination of phenotypic and flanking sequence data will stimulate the functional analysis of rice genes.

DNA Transposable Elements↗

[Multiple template switches on LINE-directed reverse transcription: the most probable formation mechanism for the double and triple chimeric retroelements in mammals].

It was shown that the shuffling mechanism for transcribed genome components, which involves a template switch on the RNA reverse transcription using the L1 retroelement enzymatic machinery, is common in mammals. The occurrence frequency of the resulting chimeric retroelements in the genomes of rodents is twice as high as in the DNA of primates. Moreover, we proved that not only single but also double switches may occur in vivo, which result in the fusion of copies of three different transcripts. Many of the identified chimeras are transcribed in mammals.

Animals↗

Genetic regulation by non-coding RNAs.

Large scale cDNA sequencing and genome tiling array studies have shown that around 50% of genomic DNA in humans is transcribed, of which 2% is translated into proteins and the remaining 98% is non-coding RNAs (ncRNAs). There is mounting evidence that these ncRNAs play critical roles in regulating DNA structure, RNA expression, protein translation and protein functions through multiple genetic mechanisms, and thus affect normal development of organisms at all levels. Today, we know very little about the regulatory mechanisms and functions of these ncRNAs, which is clearly essential knowledge for understanding the secret of life. To promote this emerging research subject of critical importance, in this paper we review (1) ncRNAs' past and present, (2) regulatory mechanisms and their functions, (3) experimental strategies for identifying novel ncRNAs, (4) experimental strategies for investigating their functions, and (5) methodologies and examples of the application of ncRNAs.

Databases, Nucleic Acid↗

Tumor-specific gene expression patterns with gene expression profiles.

Gene expression profiles of 14 common tumors and their counterpart normal tissues were analyzed with machine learning methods to address the problem of selection of tumor-specific genes and analysis of their differential expressions in tumor tissues. First, a variation of the Relief algorithm, "RFE_Relief algorithm" was proposed to learn the relations between genes and tissue types. Then, a support vector machine was employed to find the gene subset with the best classification performance for distinguishing cancerous tissues and their counterparts. After tissue-specific genes were removed, cross validation experiments were employed to demonstrate the common deregulated expressions of the selected gene in tumor tissues. The results indicate the existence of a specific expression fingerprint of these genes that is shared in different tumor tissues, and the hallmarks of the expression patterns of these genes in cancerous tissues are summarized at the end of this paper.

Algorithms↗

Ethics, genomics, and information retrieval.

The union of genomics and computational information retrieval raises a number of ethical issues, including data sharing, database accuracy, group and subgroup stigma, and privacy and confidentiality. These issues are introduced and assigned a preliminary analysis which, it is hoped, may be of use in more sustained efforts to identify issues, solutions and potential guidelines, to stimulate education, and to strike the most appropriate balance between the rights of individuals and the needs of researchers and society.

Computer Security↗

Base composition analysis of human mitochondrial DNA using electrospray ionization mass spectrometry: a novel tool for the identification and differentiation of humans.

In traditional approaches, mitochondrial DNA (mtDNA) variation is exploited for forensic identity testing by sequencing the two hypervariable regions of the human mtDNA control region. To reduce time and labor, single nucleotide polymorphism (SNP) assays are being sought to possibly replace sequencing. However, most SNP assays capture only a portion of the total variation within the desired regions, require a priori knowledge of the position of the SNP in the genome, and are generally not quantitative. Furthermore, with mtDNA, the clustering of SNPs complicates the design of SNP extension primers or hybridization probes. This article describes an automated electrospray ionization mass spectrometry method that can detect a number of clustered SNPs within an amplicon without a priori knowledge of specific SNP positions and can do so quantitatively. With this technique, the base composition of a PCR amplicon, less than 140 nucleotides in length, can be calculated. The difference in base composition between two samples indicates the presence of an SNP. Therefore, no post-PCR analytical construct needs to be developed to assess variation within a fragment. Of the 2754 different mtDNA sequences in the public forensic mtDNA database, nearly 90% could be resolved by the assay. The mass spectrometer is well suited to characterize and quantitate heteroplasmic samples or those containing mixtures. This makes possible the interpretation of mtDNA mixtures (as well as mixtures when assaying other SNPs). This assay can be expanded to assess genetic variation in the coding region of the mtDNA genome and can be automated to facilitate analysis of a large number of samples such as those encountered after a mass disaster.

Automation↗

Identification of a novel enhancer of brain expression near the apoE gene cluster by comparative genomics.

Comparative analysis of the human and mouse genomic sequences downstream of the apolipoprotein E gene (APOE) revealed a highly conserved element with previously undefined function. In reporter gene transfection studies, this element which is located approximately 42 kb distal to APOE was found to have silencer activity in a subset of cell lines examined. Analysis of transgenic mice containing a fusion construct linking this distal 631 bp conserved element to a reporter gene comprised of the human APOE gene with its proximal promoter resulted in robust brain expression of the transgenic human apoE mRNA in three independent transgenic lines, supporting the identification of a novel brain controlling region (BCR). Further studies using immunohistochemistry revealed widespread human apoE localization throughout the brains of the BCR-apoE transgenic mice with prominent expression in the cortex and diencephalon. In addition, double-label immunofluorescence performed on brain sections and cultures of primary cortical cells localized human apoE protein to cortical neurons and microglia. These studies demonstrate that comparative sequence analysis is a successful strategy to predict candidate regulatory regions in vivo, although they do not imply that this element controls apoE expression physiologically.

Animals↗

The role of informatics in glycobiology research with special emphasis on automatic interpretation of MS spectra.

This paper reviews the current status of bioinformatics applications and databases in glycobiology, which are based on bioinformatics approaches as well as informatics for glycobiology where an explicit encoding of glycan structures is required. The availability of the complete sequence of the human genome has accelerated the systematic identification of so far unidentified glycogenes considerably in many areas of glycobiology using well-established bioinfomatics tools. Although there has been an immense development of new glyco-related data collections as well as informatics tools and several efforts have been started to cross-link and reference the various data deposited in distributed databases, informatics for glycobiology and glycomics is still poorly developed compared to the genomics and proteomics area. The development of algorithms for the automatic interpretation of MS spectra - currently, a severe bottleneck, which hampers the rapid and reliable interpretation of MS data in high-throughput glycomics projects - is reviewed. A comprehensive list of web resources is given. Several lines of progression are discussed. There is an urgent need for the development of decentralised input facilities of experimentally determined glycan structures. Simultaneously, agreements of standards for the structural description of glycans as well as formats for the related data have to be established. The integration of glycomics with genomics/proteomics has to increase.

Computational Biology↗

Bioinformatics for comprehensive finding and analysis of glycosyltransferases.

Bioinformatics is a very powerful tool in the field of glycoproteomics as well as genomics and proteomics. As a part of the Glycogene Project (GG project), we have developed a novel bioinformatics system for the comprehensive identification and in silico cloning of human glycogenes. Using our system, a total of 105 candidate human glycogenes were identified and then engineered for heterologous expression. Of these candidates, 38 recombinant proteins were successfully identified for their enzyme activity and substrate specificity. We also classified 47 out of 60 carbohydrate-active enzyme glycosyltransferase families into 4 superfamilies using the profile Hidden Markov Model method. On the basis of our classification and the relationship between glycosylation pathways and superfamilies, we propose the evolution of glycosyltransferases.

Algorithms↗

Proteomic analysis of Bacillus anthracis Sterne vegetative cells.

Mass spectrometry and proteomics have found increasing use as tools for the rapid detection of pathogenic bacteria, even when they are in a mixture of non-pathogenic relatives. The success of this technique is greatly augmented by the availability of publicly accessible proteomic databases for specific pathogenic bacteria. To aid proteomic detection analyses for the causative agent of anthrax, we have constructed a comprehensive proteomic catalogue of vegetative Bacillus anthracis Sterne cells using liquid chromatography tandem-mass spectrometry. Proteins were separated by molecular weight or isoelectric point prior to tryptic digestion. Alternatively, the whole protein extract was digested and tryptic peptides were separated by cation exchange chromatography prior to Reverse Phase-LC-MS/MS. The use of three complementary, pre-analytical separation techniques resulted in the identification of 1048 unique proteins, including 694 cytosolic, 153 membrane (including 27 cell wall), and 30 secreted proteins, accounting for 19% of the total predicted proteome. Each identified protein was functionally categorized using the gene attribute database from TIGR CMR. These results provide a large proteomic catalogue of vegetative B. anthracis cells and, coupled with the recent proteomic catalogue of B. anthracis spore proteins, form a thorough summary of proteins expressed in the active and dormant stages of this organism.

Bacillus anthracis↗

Skewed X-inactivation in cloned mice.

In female mammals, dosage compensation for X-linked genes is accomplished by inactivation of one of two X chromosomes. The X-inactivation ratio (a percentage of the cells with inactivated maternal X chromosomes in the whole cells) is skewed as a consequence of various genetic mutations, and has been observed in a number of X-linked disorders. We previously reported that phenotypically normal full-term cloned mouse fetuses had loci with inappropriate DNA methylation. Thus, cloned mice are excellent models to study abnormal epigenetic events in mammalian development. In the present study, we analyzed X-inactivation ratios in adult female cloned mice (B6C3F1). Kidneys of eight naturally produced controls and 11 cloned mice were analyzed. Although variations in X-inactivation ratio among the mice were observed in both groups, the distributions were significantly different (Ansary-Bradley test, P<0.01). In particular, 2 of 11 cloned mice showed skewed X-inactivation ratios (19.2% and 86.8%). Similarly, in intestine, 1 of 10 cloned mice had a skewed ratio (75.7%). Skewed X-inactivation was observed to various degrees in different tissues of different individuals, suggesting that skewed X-inactivation in cloned mice is the result of secondary cell selection in combination with stochastic distortion of primary choice. The present study is the first demonstration that skewed X-inactivation occurs in cloned animals. This finding is important for understanding both nuclear transfer technology and etiology of X-linked disorders.

Animals↗