PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

The hppA gene of Helicobacter pylori encodes the class C acid phosphatase precursor.

Screening of the Helicobacter pylori genomic library with sera from infected humans and from immunized rabbits resulted in identification of the 25 kDa protein cell envelope (HppA) which exhibits acid phosphatase activity. Enzyme activity was demonstrated by specific enzymatic assays with whole-cell protein preparations of H. pylori strain N6 and from Escherichia coli carrying the hppA gene (pUWM192). HppA showed optimum activity at pH 5.6 and was resistant to inhibition by EDTA. Bioinformatics analysis and site-directed mutagenesis of two putative active site residues (D73 and D192) provide further insight into the sequence-structure-function relationships of HppA as a member of the DDDD phosphohydrolase superfamily.

Acid Phosphatase↗

Molecular identification of wheat endoxylanase inhibitor TAXI-I1, member of a new class of plant proteins.

Triticum aestivum endoxylanase inhibitors (TAXIs) are wheat proteins that inhibit family 11 endoxylanases commonly used in different (bio)technological processes. Here, we report on the identification of the TAXI-I gene which encodes a mature protein of 381 amino acids with a calculated molecular mass of 38.8 kDa. When expressed in Escherichia coli, the recombinant protein had the specificity and inhibitory activity of natural TAXI-I, providing conclusive evidence that the isolated gene encodes an endoxylanase inhibitor. Bioinformatical analysis indicated that no conserved domains nor motifs common to other known proteins are present. Sequence analysis revealed similarity with a glycoprotein of carrot and with gene families in Arabidopsis thaliana and rice, all with unknown functions. Our data indicate that TAXI-I belongs to a newly identified class of plant proteins for which a molecular function as glycoside hydrolase inhibitor can now be suggested.

Amino Acid Sequence↗

A thermostable manganese-containing superoxide dismutase from pathogen Chlamydia pneumoniae.

The gene CP0718 encoding a putative manganese-containing superoxide dismutase of Chlamydia pneumoniae AR39 was cloned and expressed in Escherichia coli. Characterization showed that the expressed protein with a monomeric molecular mass of 23.1 kDa had superoxide dismutase (SOD) activity and the cofactor of CpSOD was a bivalent manganese cation. It is unexpected that this enzyme was hyperthermostable, and maintained about 90% activity after incubation at 70 degrees C for 60 min. Manganese binding residues found in the SOD sequences from different species are conserved in CpSOD. Bioinformatics analysis compared with Propionibacterium shermanii MnSOD was performed to elucidate the CpSOD hyperthermostability based on amino acid sequences.

Amino Acid Sequence↗

Learning from the genome sequence of Mycobacterium tuberculosis H37Rv.

Mycobacterium tuberculosis, the scourge of humanity, is one of the most successful and scientifically challenging pathogens of all time. To catalyse the conception of new prophylactic and therapeutic interventions against tuberculosis, and to enhance our understanding of the biology of the tubercle bacillus, the complete genome sequence of the most widely used strain, H37Rv, has been determined. Bioinformatic analysis led to the identification of approximately 4000 genes in the 4.41 Mb genome sequence and provided fresh insight into the biochemistry, physiology. genetics and immunology of this much-feared bacterium. Genomic information is centralised in TubercuList (http://www.pasteur.fr/Bio/TubercuList/).

DNA, Bacterial↗

Characterization of the omega class of glutathione transferases.

The Omega class of cytosolic glutathione transferases was initially recognized by bioinformatic analysis of human sequence databases, and orthologous sequences were subsequently discovered in mouse, rat, pig, Caenorhabditis elegans, Schistosoma mansoni, and Drosophila melanogaster. In humans and mice, two GSTO genes have been recognized and their genetic structures and expression patterns identified. In both species, GSTO1 mRNA is expressed in liver and heart as well as a range of other tissues. GSTO2 is expressed predominantly in the testis, although moderate levels of expression are seen in other tissues. Extensive immunohistochemistry of rat and human tissue sections has demonstrated cellular and subcellular specificity in the expression of GSTO1-1. The crystal structure of recombinant human GSTO1-1 has been determined, and it adopts the canonical GST fold. A cysteine residue in place of the catalytic tyrosine or serine residues found in other GSTs was shown to form a mixed disulfide with glutathione. Omega class GSTs have dehydroascorbate reductase and thioltransferase activities and also catalyze the reduction of monomethylarsonate, an intermediate in the pathway of arsenic biotransformation. Other diverse actions of human GSTO1-1 include modulation of ryanodine receptors and interaction with cytokine release inhibitory drugs. In addition, GSTO1 has been linked to the age at onset of both Alzheimer's and Parkinson's diseases. Several polymorphisms have been identified in the coding regions of the human GSTO1 and GSTO2 genes. Our laboratory has expressed recombinant human GSTO1-1 and GSTO2-2 proteins, as well as a number of polymorphic variants. The expression and purification of these proteins and determination of their enzymatic activity is described.

Amino Acid Sequence↗

Interleukin 20: discovery, receptor identification, and role in epidermal function.

A structural, profile-based algorithm was used to identify interleukin 20 (IL-20), a novel IL-10 homolog. Chromosomal localization of IL-20 led to the discovery of an IL-10 family cytokine cluster. Overexpression of IL-20 in transgenic (TG) mice causes neonatal lethality with skin abnormalities including aberrant epidermal differentiation. Recombinant IL-20 protein stimulates a signal transduction pathway through STAT3 in a keratinocyte cell line, demonstrating a direct action of this ligand. An IL-20 receptor was identified as a heterodimer of two orphan class II cytokine receptor subunits. Both receptor subunits are expressed in skin and are dramatically upregulated in psoriatic skin. Taken together, these results demonstrate a role in epidermal function and psoriasis for IL-20, a novel cytokine identified solely by bioinformatics analysis.

Animals↗

Distinct replication requirements for the two Vibrio cholerae chromosomes.

Studies of prokaryotic chromosome replication have focused almost exclusively on organisms with one chromosome. We defined and characterized the origins of replication of the two Vibrio cholerae chromosomes, oriCI(vc) and oriCII(vc). OriCII(vc) differs from the origin assigned by bioinformatic analysis and is unrelated to oriCI(vc). OriCII(vc)-based replication requires an internal 12 base pair repeat and two hypothetical genes that flank oriCII(vc). One of these genes is conserved among diverse genera of the family Vibrionaceae and encodes an origin binding protein. The other gene codes for an RNA and not a protein. OriCII(vc)- but not oriCI(vc)-based replication is negatively regulated by a DNA sequence adjacent to oriCII(vc). There is an unprecedented requirement for DNA adenine methyltransferase in both oriCI(vc)- and oriCII(vc)-based replication. Our studies of replication in V. cholerae indicate that microorganisms having multiple chromosomes may utilize unique mechanisms for the control of replication.

Bacterial Proteins↗

Characterisation and distribution of a cryptic Salmonella typhi plasmid pHCM2.

pHCM2 is a 106 kbp cryptic plasmid harboured by Salmonella typhi CT18, originally isolated from a typhoid patient in Vietnam. The genome of S. typhi CT18, including pHCM2, has recently been completely sequenced and annotated. Bioinformatic analysis revealed that 57% of the coding sequences (CDSs) encoded on pHCM2 display over 97% DNA sequence identity to the virulence-associated plasmid of Yersinia pestis, pFra. pHCM2 encodes no obvious virulence-associated determinants or antibiotic resistance genes but does encode a wide array of putative genes directly related to DNA metabolism and replication. PCR analysis of a series of S. typhi isolates from Vietnam detected pHCM2-related DNA sequences in some S. typhi isolated before, but not after, 1994. Similar pHCM2-related sequences were also detected in S. typhi isolated from other regions of South East Asia and Pakistan but not elsewhere in the world.

Base Sequence↗

CAP5.5, a life-cycle-regulated, cytoskeleton-associated protein is a member of a novel family of calpain-related proteins in Trypanosoma brucei.

The cell shape of African trypanosomes is determined by the presence of an extensive subpellicular microtubule cytoskeleton. Other possible functions of the cytoskeleton, such as providing a potential framework for signalling proteins transducing information from the intracellular and extracellular environment, have not yet been investigated in trypanosomes. In this study, we have identified a novel cytoskeleton-associated protein in Trypanosoma brucei. CAP5.5 is the first member of a new family of proteins in trypanosomes, characterised by their similarity to the catalytic region of calpain-type proteases. CAP5.5 is only expressed in procyclic, but not in bloodstream, trypanosomes. Furthermore, CAP5.5 has been shown to be both myristoylated and palmitoylated, suggesting a stable interaction with the cell membrane. A bioinformatics analysis of the trypanosome genome revealed a diverse family of calpain-related proteins with primary structures similar to CAP5.5, but of varying length. We suggest a nomenclature for this new family of proteins in T. brucei.

Amino Acid Sequence↗

mymA operon of Mycobacterium tuberculosis: its regulation and importance in the cell envelope.

Mycobacterium tuberculosis faces various stressful conditions inside the host and responds to them through a coordinated regulation of gene expression. We had previously reported identification of the virS gene of M. tuberculosis (Rv3082c) belonging to the AraC family of transcriptional regulators. In the current study, we show that the seven genes (Rv3083-Rv3089) which are present divergently to virS (Rv3082c) constitute an operon designated the mymA operon. Further investigation on the regulation of this operon showed that transcription of the mymA operon is dependent on the presence of VirS protein. A four-fold induction of the mymA operon promoter occurs specifically in wild-type M. tuberculosis and not in the virS mutant of M. tuberculosis (MtbDeltavirS) when exposed to acidic pH. Expression of the mymA operon was also induced in infected macrophages by 10-fold over a 6-day period. To gain an insight into the function of the proteins encoded by this operon, we carried out a bioinformatic analysis, which suggested the involvement of these proteins in the modification of fatty acids required for cell envelope. This was supported by altered colony morphology and cell envelope structure displayed by the virS mutant of M. tuberculosis (MtbDeltavirS).

Bacterial Proteins↗

Genomic sequence and expression analyses of human chromatin assembly factor 1 p150 gene.

Chromatin assembly factor-1 (CAF-1) plays essential roles in eukaryotic chromatin assembly during DNA replication (Smith and Stillman, 1989. Cell 58, 15-25), (Krude, 1999. Eur. J. Biochem. 263, 1-5). Its p150 subunit, involved in interaction with histone H3 and H4, is critical to the CAF-1 nucleosome assembly activity. In this study, we sequenced a 96-kb genomic DNA region that includes a 42.8-kb CAF-1 p150 subunit gene (CHAF1A), and a 41.1-kb EEN gene. A scripted bioinformatics analysis pipeline (research agent) has been set up to annotate the BAC sequence with a set of integrated algorithms. The CAF-1 p150 subunit gene contains 15 exons and 14 introns. The promoter region is characterized by deletional analyses, revealing a potential repressor. Tissue-correlated alternative splicing forms of the transcript was initially identified by EST clustering analysis, then confirmed by RT-PCR which resulted more splicing forms than computational prediction.

3T3 Cells↗

The human cDNA for a homologue of the plant enzyme 1-aminocyclopropane-1-carboxylate synthase encodes a protein lacking that activity.

The sequences of genes encoding homologues of 1-aminocyclopropane-1-carboxylate (ACC) synthase, the first enzyme in the two-step biosynthetic pathway of the important plant hormone ethylene, have recently been found in Fugu rubripes and Homo sapiens (Peixoto et al., Gene 246 (2000) 275). ACC synthase (ACS) catalyzes the formation of ACC from S-adenosyl-L-methionine. ACC is oxidized to ethylene in the second and final step of ethylene biosynthesis. Profound physiological questions would be raised if it could be demonstrated that ACC is formed in animals, because there is no known function for ethylene in these organisms. We describe the cloning of the putative human ACS (PHACS) cDNA that encodes a 501 amino acid protein that exhibits 58% sequence identity to the putative Fugu ACS and approximately 30% sequence identity to plant ACSs. Purified recombinant PHACS, expressed in Pichia pastoris, contains bound pyridoxal-5'-phosphate (PLP), but does not catalyze the synthesis of ACC. PHACS does, however, catalyze the deamination of L-vinylglycine, a known side-reaction of apple ACS. Bioinformatic analysis indicates that PHACS is a member of the alpha-family of PLP-dependent enzymes. Molecular modeling data illustrate that the conservation of residues between PHACS and the plant ACSs is dispersed throughout its structure and that two active site residues that are important for ACS activity in plants are not conserved in PHACS.

Amino Acid Sequence↗

Characterization and in silico mapping of a novel murine zinc finger transcription factor.

Transcription factors play important roles in development and homeostasis. We have completed an embryonic stem cell-based neural differentiation screen, which was carried out with a view to isolating early regulators of neurogenesis. Fifty eight of the expressed sequence tags isolated from this screen represent known transcription factors or sequences containing transcription factor motifs. We have determined the full-length sequence of a novel mouse zinc finger-containing gene (ZFEND; also known as Mus musculus zinc finger protein 358 (Zfp358)) that was identified from this screen. ZFEND has 87% nucleotide and 86% amino acid identity to a previously identified human cDNA, FLJ10390, which is moderately similar to zinc finger protein 135. Northern blotting and RPAs demonstrate highest expression of ZFEND during mid-late mouse embryogenesis. Expression is also observed in several adult tissues with highest expression in heart, brain, and liver. Whole-mount in situ hybridization studies reveal apparent ubiquitous expression of ZFEND during mid-gestation stages (embryonic days 11.5, 12.5), while sections of whole-mount embryos reveal much higher expression levels in the neural folds during neural tube closure and at the boundary between the forelimb buds and the body wall. Bioinformatic analysis maps ZFEND to mouse chromosome 8pter, while FLJ10390 resides on 19p13.3-p13.2, a gene-rich region to which a number of disorders have been mapped. More precise mapping indicates that the involvement of FLJ10390 in atherogenic lipoprotein phenotype, familial febrile convulsions 2, and psoriasis susceptibility cannot be ruled out.

Amino Acid Sequence↗

Cysteine and tyrosine-rich 1 (CYYR1), a novel unpredicted gene on human chromosome 21 (21q21.2), encodes a cysteine and tyrosine-rich protein and defines a new family of highly conserved vertebrate-specific genes.

A novel human gene has been identified by in-depth bioinformatics analysis of chromosome 21 segment 40/105 (21q21.1), with no coding region predicted in any previous analysis. Brain-derived DNA complementary to RNA (cDNA) sequencing predicts a 154-amino acid product with no similarity to any known protein. The gene has been named cysteine and tyrosine-rich protein 1 gene (symbol cysteine and tyrosine-rich 1, CYYR1). The CYYR1 messenger RNA was found by Northern blot analysis in a broad range of tissues (two transcripts of 3.4 and 2.2 kb). The gene consists of four exons and spans about 107 kb, including a very large intron of 85.8 kb. Analysis of expressed sequence tags shows high CYYR1 expression in cells belonging to the amine precursor uptake and decarboxylation system. We also cloned the cDNA of the murine ortholog Cyyr1, which was mapped by a radiation hybrid panel on chromosome 16 within the region corresponding to that containing the respective human homolog on chromosome 21. Sequence and phylogenetic analysis led to identification of several genes encoding CYYR1 homologous proteins. The most prominent feature identified in the protein family is a central, unique cysteine and tyrosine-rich domain, which is strongly conserved from lower vertebrates (fishes) to humans but is absent in bacteria and invertebrates.

Amino Acid Sequence↗

The human SLC8A3 gene and the tissue-specific Na+/Ca2+ exchanger 3 isoforms.

We have identified the human gene for member 3 of Solute Carrier family 8 (SLC8A3) by bioinformatic analysis of human genomic sequences. The gene is located on chromosome 14q24.2, and spans a region of about 150 kb. The full-length DNA complementary to RNA encoding the Na(+)/Ca(2+) exchanger isoform 3 (NCX3), amplified by reverse transcriptase-polymerase chain reaction (RT-PCR) from the human neuroblastoma SH-SY5Y RNA, includes seven exons and encodes a protein of about 100 kDa. RT-PCR analysis was performed in different tissues to determine the exon composition in the region encoding the large intracellular loop of the protein. The region underwent modifications by alternative tissue-specific splicing. NCX3.2, including exon 4 but not exon 5, was found in human brain and in the neuroblastoma cell line. In human skeletal muscle two additional isoforms were identified: NCX3.3, including exons 4 and 5, and a truncated isoform (NCX3.4) produced by the skipping of both exons 3 and 4. The skipping causes a frame shift downstream of the exon 2 sequence. The new coding sequence of 25 amino acids terminates with a stop codon in exon 6. The NCX3.4 isoform (68 kDa) is truncated in the C-terminal portion of the domain first found in Drosophila Na(+)/Ca(2+) exchanger domain (Calxbeta) and lacks the C-terminal hydrophobic segments.

Alternative Splicing↗

Distribution pattern of Notch3 mutations suggests a gain-of-function mechanism for CADASIL.

Mutations in Notch3 cause the syndrome CADASIL (cerebral autosomal dominant arteriopathy with subcortical infarcts and leukoencephalopathy). The mechanism by which these mutations result in a CADASIL phenotype has been widely speculated upon. A first step toward understanding a disease mechanism is to learn whether the mutations result in the loss of Notch3 function, in particular, its role in signaling or in the gain of a novel function. Notch3 genomic sequences were analyzed for sites of conservation across species. We present here a bioinformatic analysis of the Notch paralogs and orthologs that suggest that CADASIL mutations result in a gain of function. This finding diminishes the likelihood that a Notch3 signaling deficit is responsible for the phenotype and increases the likelihood that CADASIL joins the growing list of neurological diseases with protein deposits due to misfolding and aggregation.

Animals↗

Cloning and characterization of a human lysyl oxidase-like 3 gene (hLOXL3).

Using the PCR primers generated from human expressed sequence tag (EST), the cDNA of lysyl oxidase-like gene 3 (LOXL3), a new member of human lysyl oxidases gene family, was cloned from the human fetal brain mRNA. The predicted amino acid sequence of the hLOXL3 gene was highly homologous to mLOR2. Bioinformatics analysis shows that hLOXL3 protein is also a member of the scavenger receptor cysteine-rich family, which contains a 25 amino acids signal peptide. The hLOXL3 gene was mapped to human 2p13 locus by BLAST search and at least 14 exons were found. Expression of the hLOXL3 gene was detected in several human tissues and especially high in spleen and testis.

Amino Acid Oxidoreductases↗

A genomic perspective on human proteases as drug targets.

Of the approximately 400 known human proteases, approximately 14% are under investigation as drug targets. Although the total is certain to rise during the finishing phase of the human genome project, the initial annotation of the approximately 30,000 human proteome set includes approximately 500 proteases. Bioinformatic analysis can now be performed on complete human protease families and will soon include comparisons with mice and fish. New sequences will require evaluation of their function in normal physiology and human disease. By revealing details such as splice variants and population polymorphisms, genomic sequence information will have a central role in the validation of protease drug targets.

Journal Article↗