PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

[Cloning, sequencing and bioinformatics analysis of a new tumor suppressor gene ndr2 from mouse].

BACKGROUND & OBJECTIVE: Ndr2 (N-myc down stream regulator) gene in human is a new gene cloned with the human adult whole brain cDNA as template in 1999, which accession number is AF159092 in GenBank. Locating backward position of the N-myc gene in human chromosome, this gene was named Ndr2 gene. The previous experimental results showed Ndr2 gene probably is a tumor suppressor gene. To research the function of Ndr2 gene, the authors cloned the genomic sequence of ndr2 from mouse. METHODS: To clone Ndr2 genomic sequence by reverse transcription-polymerase chain reaction(RT-PCR) with the mouse genome library as template; automatic sequencing was performed using 310 Genetic Analyzer; homogeneous analysis was made using GenBank BLAST; open reading fragment(ORF) analysis was made using PC Gene and ORF Finder; domain analysis was made using ProDom system. RESULTS: A fragment (about 3310bp,identified by agarose gel electrophoresis) was obtained using RT-PCR with the mouse genome library as template. The fragment was cloned in pMD18-T vector. BLAST analysis showed that the sequence was highly homogeneous (with the homogeneity rate of 91.4%) with Ndr2 gene in human and non-homogeneous with genomic sequence database in mouse. ORF analysis showed that there was a complete coding region in it, which including 8 extrons and 7 introns; it can interpret a protein containing about 200 amino acid residuals. ProDom analysis showed there was a domain like acyl carrier protein(ACP) in it. CONCLUSION: The authors cloned Ndr2 gene in mouse and proved that the sequence is a new genome sequence in mouse genomic sequence database. At present, the genome sequence has been submitted to GenBank(the accession number: AY151387).

Adaptor Proteins, Signal Transducing↗

[Bioinformatic analysis of dengue virus cDNAs and design of oligonucleotide probes for microarray detection of the virus].

OBJECTIVE: To design oligonucleotide probes for microarray detection of dengue virus. METHODS: By analyzing the cDNAs of dengue viruses of 4 different serotypes with BLAST program, a group of specific sequences for the candidate probes was acquired. Oligo6.0 software was applied to analyze the candidates to select the probes with high specificity, identical length and similar melting temperature (Tm). RESULT: Altogether 48 oligonucleotide probes were designed, and deposited on oligonucleotide chips as the microarray for dengue virus detection. CONCLUSION: BLAST program and Oligo6.0 software are simple and effective means for designing the oligonucleotide probes.

Base Sequence↗

A bioinformatics analysis of alternative exon usage in human genes coding for extracellular matrix proteins.

Alternative splicing increases protein diversity through the generation of different mRNA molecules from the same gene. Although alternative splicing seems to be a widespread phenomenon in the human transcriptome, it is possible that different subgroups of genes present different patterns, related to their biological roles. Analysis of a subgroup may enhance common features of its members that would otherwise disappear amidst a heterogeneous population. Extracellular matrix (ECM) proteins are a good set for such analyses since they are structurally and functionally related. This family of proteins is involved in a large variety of functions, probably achieved by the combinatorial use of protein domains through exon shuffling events. To determine if ECM genes have a different pattern of alternative splicing, we compared clusters of expressed sequences of ECM to all other genes regarding features related to the most frequent type of alternative splicing, alternative exon usage (AEU), such as: the number of alternative exon-intron structures per cluster, the number of AEU events per exon-intron structure, the number of exons per event, among others. Although we did not find many differences between the two sets, we observed a higher frequency of AEU events involving entire protein domains in the ECM set, a feature that could be associated with their multi-domain nature. As other subgroups or even the ECM set in different tissues could present distinct patterns of AEU, it may be premature to conclude that alternative splicing is homogeneous among groups of related genes.

Alternative Splicing↗

[Polymorphisms and bioinformatics analysis of chicken prolactin gene].

Four chicken breeds (White Leghorn, Yangshan, Taihe Silkies, White Recessive Rocks) with different reproduction were applied to screen potential SNPs related to laying performance in the 5' flanking region, exon region and partial intron region of chicken prolactin (cPRL) gene. Totally almost 4500 bp were screened rapidly based on DNA pooling and sequencing, and thirteen single nucleotide polymorphisms (SNPs) and two indels (24 bp and 15 bp) were found, including nine SNPs and two indels in the 5' flanking region, one SNP in Exon 2, two SNPs in Exon 5 and one SNP in Intron 2 respectively. Furthermore, 5' flanking region of cPRL gene was analyzed by the website of http://motif.genome.ad.jp/. A possible Evi-1 binding site (score 93) was found in White Leghorn cPRL gene because of the 24 bp insertion, another possible C/EBPbeta binding site (score 94) was found in Yangshan cPRL gene because of the variation of C-2402T. Further studies need to be carried out to verify their effects on the expression of cPRL gene, the broodiness and laying performance of chickens.

5' Flanking Region↗

[Bioinformatic analysis of the 14-3-3 gene family in rice].

Using two-step HMM (hidden markov model) scan strategy,eight 14-3-3-like proteins were identified by searching the Oryza sativa L. ssp. japonica protein database. From them four genes were newly detected in this study. We confined the genes expressing in Nipponbare by EST search. Expression analysis also showed each gene expressed diversely within any individual,this tends suggested specific function of particular gene. Alignment of amino acid sequences suggested that there could be isoform function of the specific residues. The analyses of gene structure and chromosome location indicated that rice genome contains both epsilon and no-epsilon forms of 14-3-3 proteins. In addition,we analyzed the evolution of the rice 14-3-3 protein family.

14-3-3 Proteins↗

[Cloning and bioinformatics analysis of a thrombin-like enzyme gene from Agkistrodon acutus].

The venoms of Viperidae and Crotalidae snakes contain a large variety of proteins and peptides affecting the hemostatic system, which classified as coagulant, anticoagulant and fibrinolytic factors. To obtaind the thrombin-like enzyme gene of snake venoms, primers 1 5' ATGGTGCTGATCAGAGTGCTAGC 3' and 2 5' CTCCTCTTAA-CTTTTTCAAAAGTTT 3' were designed according to the snake venom thrombin-like enzyme highly conserved regions of 5' and 3'. Total RNA was prepared from the venom glands of a D. acutus specimen collected from Guangxi province of China, RT-PCR was conducted to amplify the gene of the venom thrombin-like enzyme (TLE). A 0.8 kb DNA fragment was specifically amplified, inserted into the pMD18-T vector and transformed into Escherichia coli strain DH5alpha, then identified by PCR and sequencing. The results showed that this cDNA shared great sequence homology (98.5%) with the published snake TLE cDNA sequence, the deduced amino acid sequence of this TLE encoded by the 783 bp consisted of 260 amino acids, which included a signal peptide of 24 amino acids and a matured peptide of 236 amino acids. In conclusion, a new cDNA encoding snake TLE was obtained by amplificantion.

Agkistrodon↗

Coordinated regulation of the Neisseria gonorrhoeae-truncated denitrification pathway by the nitric oxide-sensitive repressor, NsrR, and nitrite-insensitive NarQ-NarP.

Neisseria gonorrhoeae survives anaerobically by reducing nitrite to nitrous oxide catalyzed by the nitrite and nitric oxide reductases, AniA and NorB. P(aniA) is activated by FNR (regulator of fumarate and nitrate reduction), the two-component regulatory system NarQ-NarP, and induced by nitrite; P(norB) is induced by NO independently of FNR by an uncharacterized mechanism. We report the results of microarray analysis, bioinformatic analysis, and chromatin immunoprecipitation, which revealed that only five genes with readily identified NarP-binding sites are differentially expressed in narP(+) and narP strains. These include three genes implicated in the truncated gonococcal denitrification pathway: aniA, norB, and narQ. We also report that (i) nitrite induces aniA transcription in a narP mutant; (ii) nitrite induction involves indirect inactivation by nitric oxide of a gonococcal repressor, NsrR, identified from a multigenome bioinformatic study; (iii) in an nsrR mutant, aniA, norB, and dnrN (encoding a putative reactive nitrogen species response protein) were expressed constitutively in the absence of nitrite, suggesting that NsrR is the only NO-sensing transcription factor in N. gonorrhoeae; and (iv) NO rather than nitrite is the ligand to which NsrR responds. When expressed in Escherichia coli, gonococcal NarQ and chimaeras of E. coli and gonococcal NarQ are ligand-insensitive and constitutively active: a "locked-on" phenotype. We conclude that genes involved in the truncated denitrification pathway of N. gonorrhoeae are key components of the small NarQP regulon, that NarP indirectly regulates P(norB) by stimulating NO production by AniA, and that NsrR plays a critical role in enabling gonococci to evade NO generated as a host defense mechanism.

Antigens, Bacterial↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

Identification of glycosylphosphatidylinositol-anchored proteins in Arabidopsis. A proteomic and genomic analysis.

In a recent bioinformatic analysis, we predicted the presence of multiple families of cell surface glycosylphosphatidylinositol (GPI)-anchored proteins (GAPs) in Arabidopsis (G.H.H. Borner, D.J. Sherrier, T.J. Stevens, I.T. Arkin, P. Dupree [2002] Plant Physiol 129: 486-499). A number of publications have since demonstrated the importance of predicted GAPs in diverse physiological processes including root development, cell wall integrity, and adhesion. However, direct experimental evidence for their GPI anchoring is mostly lacking. Here, we present the first, to our knowledge, large-scale proteomic identification of plant GAPs. Triton X-114 phase partitioning and sensitivity to phosphatidylinositol-specific phospholipase C were used to prepare GAP-rich fractions from Arabidopsis callus cells. Two-dimensional fluorescence difference gel electrophoresis and one-dimensional sodium dodecyl sulfate-polyacrylamide gel electrophoresis demonstrated the existence of a large number of phospholipase C-sensitive Arabidopsis proteins. Using liquid chromatography-tandem mass spectrometry, 30 GAPs were identified, including six beta-1,3 glucanases, five phytocyanins, four fasciclin-like arabinogalactan proteins, four receptor-like proteins, two Hedgehog-interacting-like proteins, two putative glycerophosphodiesterases, a lipid transfer-like protein, a COBRA-like protein, SKU5, and SKS1. These results validate our previous bioinformatic analysis of the Arabidopsis protein database. Using the confirmed GAPs from the proteomic analysis to train the search algorithm, as well as improved genomic annotation, an updated in silico screen yielded 64 new candidates, raising the total to 248 predicted GAPs in Arabidopsis.

Arabidopsis↗

Automated tissue analysis--a bioinformatics perspective.

OBJECTIVES: Recent progress in automated tissue analysis (tissomics) provides reproducible phenotypical characterization of histological specimens. We introduce informatics tools to cluster and correlate quantitative tissue profiles with gene expression data. The great potential of synergies between tissue analysis and bioinformatics and its perspectives in medical research and computational diagnostics are discussed. METHODS: Key enablers in microscopic imaging and machine vision are reviewed to perform a high-throughput tissue analysis. Methodologies are described and results are demonstrated that support a combined analysis of tissue with gene expression profiles whereby the consideration of individual responses is key. RESULTS: Comprehensive histomorphometric profiles, extracted using machine vision, provide information regarding the components and heterogeneity of a tissue in a reproducible format amenable to data mining and analysis. Tissue quantitative information can be placed in synergetic context with bioinformatics data, such as gene expression profiles, for a more comprehensive stratification of individual responses. From a bioinformatics point of view tissue data are co-variants that support the identification of candidate genes relevant in tissue injury or disease. CONCLUSIONS: Progress in automated analytics enables the generation of quantitative data about tissue previously limited to visual histopathology. Such reproducible data sets can be statistically correlated and clustered throughout the continuum of bioinformatics. The combined approach supports a system-wide view of biology and has a potential to accelerate developments for a personalized computational diagnosis.

Automation↗

Predicting HIV-1 coreceptor usage with sequence analysis.

Bioinformatics approaches are increasingly being used to identify and understand the genetic variation underlying changes in HIV-1 biological phenotype. The variable regions of the viral envelope are the major determinant of virus coreceptor usage and cell tropism. Specifically, amino acids 11 and 25 in the 3rd variable (V3) loop have been found to strongly influence viral syncytium inducing capacity and coreceptor usage. Many additional V3 loop changes, however, as well as changes elsewhere in Env, are thought to contribute to phenotype. In this review we describe several recently developed methods to analyze this variability and their use to predict biological phenotype based on sequence information. These approaches have identified changes in the V3 loop, in addition to the known changes at positions 11 and 25, that affect phenotype and significantly enhance our ability to predict phenotype from genotype. Besides improving phenotype prediction, methods that score V3 sequences on a continuous scale can also assist in the interpretation of evolutionary information about shifts in phenotype, and the relationship between that evolution and pathogenesis. Several examples and potential practical applications of this scoring are discussed. We conclude that advances in computational approaches have enhanced both our ability to predict and to understand HIV-1 biological phenotype evolution. Further development of these methods, by extending analysis to regions outside the V3 loop and to clades beyond subtype B, will extend our understanding of HIV-1 pathogenesis and inform treatment strategies.

Computational Biology↗

Crystal structure of a putative methyltransferase from Mycobacterium tuberculosis: misannotation of a genome clarified by protein structural analysis.

Bioinformatic analyses of whole genome sequences highlight the problem of identifying the biochemical and cellular functions of many gene products that are at present uncharacterized. The open reading frame Rv3853 from Mycobacterium tuberculosis has been annotated as menG and assumed to encode an S-adenosylmethionine (SAM)-dependent methyltransferase that catalyzes the final step in menaquinone biosynthesis. The Rv3853 gene product has been expressed, refolded, purified, and crystallized in the context of a structural genomics program. Its crystal structure has been determined by isomorphous replacement and refined at 1.9 A resolution to an R factor of 19.0% and R(free) of 22.0%. The structure strongly suggests that this protein is not a SAM-dependent methyltransferase and that the gene has been misannotated in this and other genomes that contain homologs. The protein forms a tightly associated, disk-like trimer. The monomer fold is unlike that of any known SAM-dependent methyltransferase, most closely resembling the phosphohistidine domains of several phosphotransfer systems. Attempts to bind cofactor and substrate molecules have been unsuccessful, but two adventitiously bound small-molecule ligands, modeled as tartrate and glyoxalate, are present on each monomer. These may point to biologically relevant binding sites but do not suggest a function. In silico screening indicates a range of ligands that could occupy these and other sites. The nature of these ligands, coupled with the location of binding sites on the trimer, suggests that proteins of the Rv3853 family, which are distributed throughout microbial and plant species, may be part of a larger assembly binding to nucleic acids or proteins.

Amino Acid Sequence↗

Analysis of differentially expressed genes in schizophrenia based on bioinformatics and corresponding mRNA expression levels.

OBJECTIVE: This study aimed to use bioinformatics analysis to identify differentially expressed genes (DEGs) involved in the pathogenesis of schizophrenia and validate their mRNA expression levels through real-time quantitative PCR (qPCR). MATERIAL/METHODS: Datasets from the publicly available Gene Expression Omnibus (GEO) database were analyzed using R software to identify DEGs. Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, were conducted. A protein-protein interaction (PPI) network was constructed using Cytoscape software to identify key genes with notable expression changes. The expression levels of these key genes were subsequently validated in schizophrenia patients using qPCR to assess potential susceptibility genes. RESULTS: In total, 813 DEGs were identified, with six key genes highlighted through GO analysis and PPI network screening. Among these, HDAC1, UBA52, and FYN demonstrated statistically significant differences in mRNA expression between schizophrenia patients and healthy controls (P&#xa0;<&#xa0;0.05). CONCLUSIONS: This study identified several DEGs potentially linked to the pathogenesis of schizophrenia, suggesting that HDAC1, UBA52, and FYN could serve as candidate susceptibility genes and diagnostic biomarkers. These findings provide new insights and directions for future schizophrenia research.

Humans↗

Rare germ line CHEK2 variants identified in breast cancer families encode proteins that show impaired activation.

Germ line mutations in CHEK2, the gene that encodes the Chk2 serine/threonine kinase activated in response to DNA damage, have been found to confer an increased risk of some cancers. We have previously reported the presence of the common deleterious 1100delC and four rare CHEK2 mutations in inherited breast cancer. Here, we report that predictions made by bioinformatic analysis on the rare mutations indicate that two of these, delE161 (483-485delAGA) and R117G, are likely to be deleterious. We show that the proteins encoded by 1100delC and delE161 are both unstable and inefficiently phosphorylated at Thr68 in response to DNA damage, a step necessary for the oligomerization of Chk2. Oligomerization is in turn necessary for additional phosphorylation and full activation of the protein. A second rare mutation, R117G, is phosphorylated at Thr68 but fails to show a mobility shift on DNA damage, suggesting that it fails to become further phosphorylated and hence fully activated. Our results indicate that delE161 and R117G encode nonfunctional proteins and are therefore likely to be pathogenic. The findings from the biochemical analysis correlate well with predictions made by bioinformatics analysis. In addition, the results imply that these mutations, as well as 1100delC, cannot act in a dominant-negative manner to cause cancer, and tumorigenesis in association with these mutations may be due to haploinsufficiency.

Amino Acid Sequence↗

[Molecular cloning, characterization, chromosomal assignment, genomic organization and verification of SFRS12(SRrp508), a novel member of human SR protein superfamily and a human homolog of rat SRrp86].

We have identified and characterized a novel human serine-arginine-rich (SR) splicing regulatory protein 508 (SRrp508) gene that is related to other members of the growing SR superfamily, but only homologous to rat (Rattus norvegicus) serine-arginine-rich splicing regulatory protein 86 (SRrp86) gene. The full-length cDNA of 3811 bp for human SRrp508 was cloned through a blast search of public databases following the identification of a cDNA contig of 658 bp obtained by EST assembly with full robotization in supercomputer in large-scale. Structurally, human SRrp508 encodes a polypeptide of 508 amino acids, which contains a single amino-terminal RNA recognition motif (RRM) and two carboxy-terminal domains rich in serine-arginine dipeptides that are highly conserved among other members of the SR superfamily. The conserved SR and RRM domains emphasize the biological importance of this gene. The SRrp508 gene, which contains 12 exons ranging from 0.096 to 2.093 kb and 11 introns ranging from 0.14 to 5.153 kb, is mapped to the human cytogenetic region 5q11.2-q12.1 using the bioinformatic analysis, and it does not link to any other genes. Furthermore, we have experimentally cloned and sequenced a cDNA fragment of 1680 bp containing the full-length ORF of 1527 bp in this novel human gene by RT-PCR from the single-stranded human pancreas cDNA library (Clontech), which is fully identical with that of the in silico cloning determined by the nucleotide sequencing. Thus, we in silico cloned his gene with GenBank accession number of AF459094 identified solely by bioinformatic analysis of the nucleotide and protein. This novel gene has promotors, TATA-box, several stop codons in the upstream of ORF, and PolyA signal in the downstream of ORF. Based on the above results, it can be concluded that we have obtained a complete novel human gene. The gene sequence exhibits good overall homology to that of rat SRrp86 gene, with 84% and 86% identity over the full-length nucleotide and protein, respectively, and with 96% and 86% identity over the serine-rich domain (RS) or arginine-rich domain (RA), respectively. The full-length sequence exhibits little overall homology to any other known protein at either the nucleotide or the amino acid level. The other two most closely related proteins, with 34% and 35% identity over the full-length protein, respectively, or with 51% and 54% identity over the full-length nucleotide of ORF, respectively, are drosophila serine-arginine-rich protein 54 (SRp54) and human arginine-rich nuclear protein 54 (p54). When comparisons are restricted to the RS or RA domains, the percent identity increased for both SRp54 and p54 are 44% and 54% or 38% and 43%, respectively. These results well demonstrate that only the novel human protein of 508 amino acids cloned is the human homolog of rat SRrp86, thus correcting the standpoint made by Barnard and Patton (Barnard DC, Patton JG. Identification and Characterization of a Novel Serine-Arginine-Rich Splicing Regulatory Protein. Molecular and Cellular Biology, 2000, 20(9): 3049-3057) that human arginine-rich nuclear protein 54 (p54) is the human homolog of the rat SRrp86, and suggesting that human SRrp508 is a new member of this growing superfamily of SR proteins. SRrp508 has an extensive expression profile, and may be a transcriptional factor. On the basis of its sequence and functional properties, we have named this protein SRrp508 for SR-related splicing regulatory protein of 508 amino acids. In summary, by combining bioinformatic analysis with experimental verification, we have successfully cloned the human cDNA homolog of rat SRrp86, which is verified by a series of theoretical and experimental evidence. The HGNC has just given SRrp508 gene entry the nomenclature information containing APPROVED SYMBOL: SFRS12; NAME: splicing factor, arginine/serine-rich 12; and ALIAS: DKFZp564B176, SRrp86. We have cloned this gene for near one year with no person landing the GenBank for registering the same gene. Our newly-established technique line will be helpful in discovering much more novel human genes.

Amino Acid Sequence↗

A microarray study of MPP+-treated PC12 Cells: Mechanisms of toxicity (MOT) analysis using bioinformatics tools.

BACKGROUND: This paper describes a microarray study including data quality control, data analysis and the analysis of the mechanism of toxicity (MOT) induced by 1-methyl-4-phenylpyridinium (MPP+) in a rat adrenal pheochromocytoma cell line (PC12 cells) using bioinformatics tools. MPP+ depletes dopamine content and elicits cell death in PC12 cells. However, the mechanism of MPP+-induced neurotoxicity is still unclear. RESULTS: In this study, Agilent rat oligo 22K microarrays were used to examine alterations in gene expression of PC12 cells after 500 muM MPP+ treatment. Relative gene expression of control and treated cells represented by spot intensities on the array chips was analyzed using bioinformatics tools. Raw data from each array were input into the NCTR ArrayTrack database, and normalized using a Lowess normalization method. Data quality was monitored in ArrayTrack. The means of the averaged log ratio of the paired samples were used to identify the fold changes of gene expression in PC12 cells after MPP+ treatment. Our data showed that 106 genes and ESTs (Expressed Sequence Tags) were changed 2-fold and above with MPP+ treatment; among these, 75 genes had gene symbols and 59 genes had known functions according to the Agilent gene Refguide and ArrayTrack-linked gene library. The mechanism of MPP+-induced toxicity in PC12 cells was analyzed based on their genes functions, biological process, pathways and previous published literatures. CONCLUSION: Multiple pathways were suggested to be involved in the mechanism of MPP+-induced toxicity, including oxidative stress, DNA and protein damage, cell cycling arrest, and apoptosis.

1-Methyl-4-phenylpyridinium↗

Genomic analysis of anaerobic respiration in the archaeon Halobacterium sp. strain NRC-1: dimethyl sulfoxide and trimethylamine N-oxide as terminal electron acceptors.

We have investigated anaerobic respiration of the archaeal model organism Halobacterium sp. strain NRC-1 by using phenotypic and genetic analysis, bioinformatics, and transcriptome analysis. NRC-1 was found to grow on either dimethyl sulfoxide (DMSO) or trimethylamine N-oxide (TMAO) as the sole terminal electron acceptor, with a doubling time of 1 day. An operon, dmsREABCD, encoding a putative regulatory protein, DmsR, a molybdopterin oxidoreductase of the DMSO reductase family (DmsEABC), and a molecular chaperone (DmsD) was identified by bioinformatics and confirmed as a transcriptional unit by reverse transcriptase PCR analysis. dmsR, dmsA, and dmsD in-frame deletion mutants were individually constructed. Phenotypic analysis demonstrated that dmsR, dmsA, and dmsD are required for anaerobic respiration on DMSO and TMAO. The requirement for dmsR, whose predicted product contains a DNA-binding domain similar to that of the Bat family of activators (COG3413), indicated that it functions as an activator. A cysteine-rich domain was found in the dmsR gene, which may be involved in oxygen sensing. Microarray analysis using a whole-genome 60-mer oligonucleotide array showed that the dms operon is induced during anaerobic respiration. Comparison of dmsR+ and DeltadmsR strains by use of microarrays showed that the induction of the dmsEABCD operon is dependent on a functional dmsR gene, consistent with its action as a transcriptional activator. Our results clearly establish the genes required for anaerobic respiration using DMSO and TMAO in an archaeon for the first time.

Anaerobiosis↗

Functional annotation and analysis of Korean patented biological sequences using bioinformatics.

A recent report of the Korean Intellectual Property Office (KIPO) showed that the number of biological sequence-based patents is rapidly increasing in Korea. We present biological features of Korean patented sequences though bioinformatic analysis. The analysis is divided into two steps. The first is an annotation step in which the patented sequences were annotated with the Reference Sequence (RefSeq) database. The second is an association step in which the patented sequences were linked to genes, diseases, pathway, and biological functions. We used Entrez Gene, Online Mendelian Inheritance in Man (OMIM), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Gene Ontology (GO) databases. Through the association analysis, we found that nearly 2.6% of human genes were associated with Korean patenting, compared to 20% of human genes in the U.S. patent. The association between the biological functions and the patented sequences indicated that genes whose products act as hormones on defense responses in the extra-cellular environments were the most highly targeted for patenting. The analysis data are available at http://www.patome.net.

Base Sequence↗