PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Bioinformatic analysis of the genomes of the cyanobacteria Synechocystis sp. PCC 6803 and Synechococcus elongatus PCC 7942 for the presence of peroxiredoxins and their transcript regulation under stress.

The genomes of the cyanobacteria Synechocystis sp. PCC 6803 and Synechococcus elongatus PCC 7942 encode five and six open reading frames (ORFs), respectively, with similarity to peroxide-detoxifying peroxiredoxins (Prx). In addition to one highly conserved gene each for 2-Cys Prx and 1-Cys Prx, the Synechocystis sp. PCC 6803 genome contains one TypeII Prx and two PrxQ-like ORFs, while Synechococcus elongatus PCC 7942 has four PrxQ-like ORFs. The transcript regulation of all these bioinformatically identified genes was analysed under selected stress conditions, i.e. light limitation and light stress, hydrogen peroxide, methylviologen, salinity, as well as nitrogen- and iron-deficiency. The results on specific time- and stress-dependent regulation of transcript amounts suggest conserved as well as variable functions of these putative Prx-s in antioxidant defence. The results are discussed in the context of evolution and physiological function, particularly in relation to photosynthesis.

Amino Acid Sequence↗

Molecular cloning and bioinformatics analysis of a novel spliced variant of survivin from human breast cancer cells.

Survivin gene and its two alternatively spliced variants, survivin-2 B and survivin- Delta Ex3 gene were cloned from human breast cancer cell lines B-cap37 firstly. A new gene designated as survivin-image (SI) was cloned from above cell lines, which has not been reported yet to clone from any cell lines. It was found that the novel gene 507 bp comprises partial survivin gene (345 bp), partial image gene (155 bp) of eye cancer and other insertion of 7 bp by analyzing with a series of recent bioinformatics software at the level of nucleotide and protein deduced. Predicted 3-D structures of the new molecule showed greatly similar to that of survivin in N-terminal containing BIR by homology modeling. These results suggested SI gene (GenBank accession No.AY830084) might be a novel alternatively spliced isoforms of the survivin gene involved in other functional significances related to tumorigenesis.

Alternative Splicing↗

Serological identification and bioinformatics analysis of immunogenic antigens in multiple myeloma.

Identifying appropriate tumor antigens is critical to the development of successful specific cancer immunotherapy. Serological analysis of tumor antigens by a recombinant cDNA expression library (SEREX) allows the systematic cloning of tumor antigens recognized by the spontaneous autoantibody repertoire of cancer patients. We applied SEREX to the cDNA expression library of cell line HMy2, which led to the isolation of six known characterized genes and 12 novel genes. Known genes, including ring finger protein 167, KLF10, TPT1, p02 protein, cDNA FLJ46859 fis, and DNMT1, were related to the development of different tumors. Bioinformatics was performed to predict 12 novel MMSA (multiple myeloma special antigen) genes. The prediction of tumor antigens provides potential targets for the immunotherapy of patients with multiple myeloma (MM) and help in the understanding of carcinogenesis. Crude lysate ELISA methodology indicated that the optical density value of MMSA-3 and MMSA-7 were significantly higher in MM patients than in healthy donors. Furthermore, SYBR Green real-time PCR showed that MMSA-1 presented with a high number of copy messages in MM. In summary, the antigens identified in this study may be potential candidates for diagnosis and targets for immunotherapy in MM.

Antigens, Neoplasm↗

[Immunological screening for multiple myeloma-associated antigens and their bioinformatics analysis].

This study was aimed to screen the cell cDNA expression library of multiple myeloma HMy2 (MM HMy2) by using "serological analysis of cDNA expression library (SEREX)" technique. The obtained 30 positive clones were all sequenced, and analyzed by BLAST (basic local alignment search tool). The results indicated that 6 known genes and 12 new MM-associated genes were obtained, part of which sequences were spliced by EST (expressed sequence tag) splicing. 6 known genes such as for ring finger protein 167, KLF10, TPT1 protein, p02 protein, cDNA FLJ46859 fis, DNMT1 methyltrasferase etc. have been demonstrated a certain relationship with other tumor's formation, progress and prognosis. The structures and functions of the new genes preliminarily analyzed and predicted by means of bioinformatics showed that MMSA-3, MMSA-8 and MMSA-11 encoding 215, 160 and 122 amino acid residues respectively had the full open reading frames (ORF). All the new genes might be located at euchromosomes but MMSA-1 at sex chromosome. MMSA-4 was highly similar to the protein controlling the transcription of tumor antigen, MMSA-5 might take part in cell phagocytosis, MMSA-7 might inactivated NF-kappaB, and MMSA-12 might be a lymphocytic cytoplasmic protein. The specificity of new genes such as MMSA-3 and MMSA-7 were higher, by a preliminary analysis using CrELISA. It is concluded that tumor antigens screened by this study can be used for early immunological diagnosis, surveillance of minor residual foci, assessment of prognosis, and preparation of tumor vaccine and so on.

Antigens, Neoplasm↗

Subunit E of mitochondrial ATP synthase: a bioinformatic analysis reveals a phosphopeptide binding motif supporting a multifunctional regulatory role and identifies a related human brain protein with the same motif.

The mitochondrial adenosine triphosphate (ATP) synthase is located in the inner membrane and consists of at least 16 subunit types in animals, one of which is subunit e, the function of which is not clearly defined. A highly homologous protein is located in the nucleus and named progesterone receptor binding protein (RBF), to designate its role in this organelle. In addition, the expression level of subunit e in mammalian cells fluctuates greatly and is induced by certain carcinogens and elevated in liver cancers. Because these previous observations suggested to us that subunit e may play multifunctional regulatory roles, we employed a bioinformatic approach to test this view. First, from sequence alignment studies, secondary structure analyses, and basic local alignment search tool (BLAST) searches, we concluded that mitochondrial subunit e and the homologous nuclear protein RBF are most likely the same protein. Second, we examined the known sequence and structure of one of the most common multifunctional cell regulatory proteins, the 14-3-3 protein, involved in phosphopeptide binding, and deduced that it has an apparent binding motif (-KX(6)R---RY-). Third, from careful examination of the conserved residues within all subunit e sequences in the database, we discovered that this protein has a comparable binding motif (-RY---KX(6)R-). Finally, in a BLAST search for additional homologs of subunit e, we found a human brain protein, KIAA1578, the C-terminal 30 amino acids of which are identical to those of human subunit e. This protein also contains a potential phosphopeptide binding motif. In summary, these studies provide support for the view that subunit e is a multifunctional cell regulator involved in cell signaling, and implicate the involvement of the KIAA1578 protein in cell signaling as well. These studies suggest also that, while functioning as a subunit of mitochondrial ATP synthases, subunit e may help regulate these complexes by binding to phosphopeptides within one or more of the other subunit types.

14-3-3 Proteins↗

Bioinformatic analysis of the link between gene composition and expressivity in Saccharomyces cerevisiae and Schizosaccharomyces pombe.

The compositional non-randomness was studied in genes of Saccharomyces cerevisiae and Schizosaccharomyces pombe. In both species, codon usage is well correlated with expressivity (measured as the codon adaptation index). Both species generally display higher nucleotide non-randomness in the group of highly expressed genes than in the lowly expressed genes. The highly expressed genes in both species are furthermore characterized by marked peaks in non-randomness at N=3 upstream of start codons, N=2 downstream of start codons and at N=1 and N=7 downstream of stop codons, indicating that these nucleotides may be key elements in translational regulation. Intragenic variation in codon usage was also observed to be linked to expressivity. It is suggested that the firm link between expressivity and codon usage calls for codon optimization. Based on bioinformatic calculations, examples of proteins are given for which codon optimizations might be relevant.

Base Sequence↗

Bioinformatics analysis of ferroptosis in frozen shoulder.

OBJECTIVES: Frozen shoulder is a common shoulder disease that significantly affects the patient's life and work. Ferroptosis is a new type of programmed cell death, which is involved in many diseases. However, there have been no studies reporting the relationship between frozen shoulders and ferroptosis. This study identified potential molecular markers of ferroptosis in frozen shoulders to provide more effective strategies for the treatment of frozen shoulders. METHODS: GSE238053 was downloaded from the Gene Expression Omnibus (GEO) dataset and intersected with ferroptosis genes to obtain differentially expressed genes (DEGs). The signaling pathways and biological functions of DEGs were performed by WebGestalt and Metascape. The interactions related to these DEGs and the key genes between frozen shoulders and ferroptosis was performed by STRING and Cytoscape. A frozen shoulders rat model was used to validate our predicted genes, Western Blot and qRT-PCR was used to assess the expression levels of our genes of interest. RESULTS: A total of 34 DEGs between GSE238053 and Ferroptosis Database were obtained, most of which were involved in the HIF-1 signaling pathway and inflammatory response. A protein-protein interaction network was obtained by Cytoscape and the key genes (IL-6, HMOX1 and TLR4) were screened by MCODE. Our results of Western Blot showed that the protein expression level of TLR4 and HMOX1 were elevated, and the protein level of IL-6 decreased in frozen shoulders rat model. The mRNA level after frozen shoulders showed that IL-6 was upregulated, whereas TLR4 and HMOX1were downregulated. CONCLUSIONS: The results demonstrated that ferroptosis may affect the pathological process of frozen shoulders through these signaling pathways and genes. The identification of IL-6, HMOX1 and TLR4 genes can provide new therapeutic targets for frozen shoulders.

Ferroptosis↗

Gene profiling and bioinformatic analysis of Schwann cell embryonic development and myelination.

To elucidate the molecular mechanisms involved in Schwann cell development, we profiled gene expression in the developing and injured rat sciatic nerve. The genes that showed significant changes in expression in developing and dedifferentiated nerve were validated with RT-PCR, in situ hybridisation, Western blot and immunofluorescence. A comprehensive approach to annotating micro-array probes and their associated transcripts was performed using Biopendium, a database of sequence and structural annotation. This approach significantly increased the number of genes for which a functional insight could be found. The analysis implicates agrin and two members of the collapsin response-mediated protein (CRMP) family in the switch from precursors to Schwann cells, and synuclein-1 and alphaB-crystallin in peripheral nerve myelination. We also identified a group of genes typically related to chondrogenesis and cartilage/bone development, including type II collagen, that were expressed in a manner similar to that of myelin-associated genes. The comprehensive function annotation also identified, among the genes regulated during nerve development or after nerve injury, proteins belonging to high-interest families, such as cytokines and kinases, and should therefore provide a uniquely valuable resource for future research.

Agrin↗

Bioinformatics Analysis and Experimental Validation of Key Genes Associated With Hypoxia and Ischemia in Myocardial Infarction.

BACKGROUND: This study aimed to screen and identify core hypoxia-ischemia-related genes associated with myocardial infarction (MI). METHOD: Two transcriptomic datasets, GSE97320 and GSE48060, were retrieved from the Gene Expression Omnibus (GEO) database. After data integration and batch effect elimination, differential expression analysis was performed to screen differentially expressed genes (DEGs), and the corresponding visualization analysis was conducted. Hypoxia-ischemia-related genes were acquired from the GeneCards database; hypoxia-ischemia related genes (HIRGs) were subsequently identified by intersecting the retrieved genes with screened DEGs. Gene Ontology (GO) functional enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were implemented to explore the biological functions and underlying signaling pathways of HIRGs. A combination of protein-protein interaction (PPI) network analysis and random forest (RF) algorithm was applied to screen hub genes from HIRGs. The external GEO dataset GSE66360 was utilized to validate the expression patterns of candidate hub genes. Furthermore, an acute myocardial infarction (AMI) mouse model was established, and quantitative real-time polymerase chain reaction (qPCR) was performed to detect the mRNA expression levels of hub genes in myocardial tissues for in&#xa0;vivo validation. RESULTS: A total of 633 DEGs and 308 hypoxia-ischemia-related genes were screened in the present study, among which 21 overlapping HIRGs were obtained. PLAUR and IL1B were finally identified as two hub genes from HIRGs based on PPI network and random forest algorithm. The qPCR results revealed that the expression levels of PLAUR and IL1B were significantly upregulated in the AMI group compared with the sham operation group (p&#x2009;<&#x2009;0.05). CONCLUSION: The present findings demonstrated that PLAUR and IL1B serve as pivotal genes involved in the pathological hypoxia-ischemia process of AMI. These two genes may act as novel biomarkers and promising therapeutic targets for the recognition and clinical intervention of hypoxia-ischemia injury following AMI.

Myocardial Infarction↗

Proteomic and bioinformatic analysis of iron- and sulfur-oxidizing Acidithiobacillus ferrooxidans using immobilized pH gradients and mass spectrometry.

A comparative analysis of the protein composition of Acidithiobacillus ferrooxidans cells grown on elemental sulfur and ferrous iron was performed. A newly developed protocol involving immobilized pH gradients, improved protein reduction, mass spectrometry protein identification and full genome sequence information was applied. This approach resulted in more than 1300 protein spots displayed in broad and basic pH ranges, the best A. ferrooxidans proteome resolution to date. A comparative image analysis revealed that the proteome was significantly influenced by the growth type, and allowed for the detection of many physiologically important proteins. Among them were sulfate adenylyltransferase and sulfide dehydrogenase, which are involved in sulfate assimilation and sulfide metabolism, respectively. Many other proteins were related to important processes like cell attachment and electron transport. Co-migration of phosphate and sulfate transport proteins was also observed.

Acidithiobacillus↗

Bioinformatic analysis of primary endothelial cell gene array data illustrated by the analysis of transcriptome changes in endothelial cells exposed to VEGF-A and PlGF.

We recently published a review in this journal describing the design, hybridisation and basic data processing required to use gene arrays to investigate vascular biology (Evans et al. Angiogenesis 2003; 6: 93-104). Here, we build on this review by describing a set of powerful and robust methods for the analysis and interpretation of gene array data derived from primary vascular cell cultures. First, we describe the evaluation of transcriptome heterogeneity between primary cultures derived from different individuals, and estimation of the false discovery rate introduced by this heterogeneity and by experimental noise. Then, we discuss the appropriate use of Bayesian t-tests, clustering and independent component analysis to mine the data. We illustrate these principles by analysis of a previously unpublished set of gene array data in which human umbilical vein endothelial cells (HUVEC) cultured in either rich or low-serum media were exposed to vascular endothelial growth factor (VEGF)-A165 or placental growth factor (PlGF)-1(131). We have used Affymetrix U95A gene arrays to map the effects of these factors on the HUVEC transcriptome. These experiments followed a paired design and were biologically replicated three times. In addition, one experiment was repeated using serial analysis of gene expression (SAGE). In contrast to some previous studies, we found that VEGF-A and PlGF consistently regulated only small, non-overlapping and culture media-dependant sets of HUVEC transcripts, despite causing significant cell biological changes.

Cells, Cultured↗

Bioinformatic analysis of the TonB protein family.

TonB is a protein prevalent in a large number of Gram-negative bacteria that is believed to be responsible for the energy transduction component in the import of ferric iron complexes and vitamin B(12) across the outer membrane. We have analyzed all the TonB proteins that are currently contained in the Entrez database and have identified nine different clusters based on its conserved 90-residue C-terminal domain amino acid sequence. The vast majority of the proteins contained a single predicted cytoplasmic transmembrane domain; however, nine of the TonB proteins encompass a approximately 290 amino acid N-terminal extension homologous to the MecR1 protein, which is composed of three additional predicted transmembrane helices. The periplasmic linker region, which is located between the N-terminal domain and the C-terminal domain, is extremely variable both in length (22-283 amino acids) and in proline content, indicating that a Pro-rich domain is not a required feature for all TonB proteins. The secondary structure of the C-terminal domain is found to be well preserved across all families, with the most variable region being between the second alpha-helix and the third beta-strand of the antiparallel beta-sheet. The fourth beta-strand found in the solution structure of the Escherichia coli TonB C-terminal domain is not a well conserved feature in TonB proteins in most of the clusters. Interestingly, several of the TonB proteins contained two C-terminal domains in series. This analysis provides a framework for future structure-function studies of TonB, and it draws attention to the unusual features of several TonB proteins.

Amino Acid Sequence↗

Molecular modeling and bioinformatical analysis of the antibacterial target enzyme MurA from a drug design perspective.

The enzyme MurA (UDP-N-acetylglucosamine enolpyruvyl transferase) catalyzes the first cytoplasmatic step in the synthesis of murein precursors. This function is of vital relevance for bacteria, and the enzyme therefore represents an important target protein for the development of novel antibacterial compounds. Several X-ray structures of liganded and un-liganded MurA have been published, which may be used for rational drug design. MurA, however, contains a highly flexible surface loop, which is involved in substrate and inhibitor binding. In the available X-ray structures, the conformation of this surface loop varies, depending on the presence or absence of ligands or substrate and probably also on the crystal packing. The uncertainty of the low-energy, or "resting state" conformation of this surface loop hampers the application of rational drug design to this class of enzymes. We have therefore performed an extensive molecular dynamics study of the enzyme in order to identify one or several low-energy conformers. The results indicate that, at least in some of the X-ray structures, the conformation of the flexible surface loop is influenced by crystallographic contacts. Furthermore, three partially helical foldamers of the surface loop are identified which may resemble the resting states of the enzyme or intermediate states that are "traversed" during the substrate binding process. Another, very important aspect for the development of novel antibacterial compounds is the inter- and intra-species variability of the target structure. We present a comparison of MurA sequences from 163 organisms which were analyzed under the aspects of enzyme mechanism, structure and drug design. The results allow us to identify the most promising binding sites for inhibitor interaction, which are present in MurA enzymes of most species and are expected to be insusceptible to resistance-inducing mutations.

Alkyl and Aryl Transferases↗

Structure, circadian regulation and bioinformatic analysis of the unique sigma factor gene in Chlamydomonas reinhardtii.

In higher plants, the transcription of plastid genes is mediated by at least two types of RNA polymerase (RNAP); a plastid-encoded bacterial RNAP in which promoter specificity is conferred by nuclear-encoded sigma factors, and a nuclear-encoded phage-like RNAP. Green algae, however, appear to possess only the bacterial enzyme. Since transcription of much, if not most, of the chloroplast genome in Chlamydomonas reinhardtii is regulated by the circadian clock and the nucleus, we sought to identify sigma factor genes that might be responsible for this regulation. We describe a nuclear gene (RPOD) that is predicted to encode an 80 kDa protein that, in addition to a predicted chloroplast transit peptide at the N-terminus, has the conserved motifs (2.1- 4.2) diagnostic of bacterial sigma-70 factors. We also identified two motifs not previously recognized for sigma factors, adjacent PEST sequences and a leucine zipper, both suggested to be involved in protein-protein interactions. PEST sequences were also found in approximately 40% of sigma factors examined, indicating they may be of general significance. Southern blot hybridization and BLAST searches of the genome and EST databases suggest that RPODmay be the only sigma factor gene in C. reinhardtii. The levels of RPODmRNA increased 2- 3-fold in the mid-to-late dark period of light-dark cycling cells, just prior to, or coincident with, the peak in chloroplast transcription. Also, the dark-period peak in RPOD mRNA persisted in cells shifted to continuous light or continuous dark for at least one cycle, indicating that RPODis under circadian clock control. These results suggest that regulation of RPODexpression contributes to the circadian clock's control of chloroplast transcription.

Journal Article↗

Expression and prognosis of CXCL13 in uterine corpus endometrial carcinoma based on bioinformatics analysis.

OBJECTIVE: The biological significance of the chemokine ligand C-X-C motif chemokine ligand 13 (CXCL13) may play a significant role in the pathogenesis of uterine corpus endometrial carcinoma (UCEC). This study aims to identify and verify CXCL13 with predictive value for prognosis in UCEC. METHODS: CXCL13 mRNA expression differences were analyzed using R software in three independent datasets: one each from The Cancer Genome Atlas (TCGA) and two from the Gene Expression Omnibus (GEO), namely GSE17025 and GSE106191. The correlation between CXCL13 expression and prognosis was evaluated by Kaplan-Meier analysis. Univariate and multivariate Cox analyses were utilized to construct a prognostic nomogram. Tumor Immune Estimation Resource (TIMER) and the Tumor and Immune System Interaction Database (TISIDB) were employed to assess the relationship between CXCL13 and tumor immune infiltration. Coexpressed genes with CXCL13 were identified by the Spearman correlation analysis. A CXCL13 protein-protein interaction (PPI) network was constructed with the STRING website tool and hub genes were screened out. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genome (KEGG) analyses were performed with the "clusterProfiler" R package. Gene set enrichment analysis (GSEA) was used to identify underlying biological mechanisms. A drug-gene interaction network was constructed in the Comparative Toxicogenomics Database (CTD). RESULTS: High CXCL13 mRNA expression were validated in UCEC in the above three independent datasets. High CXCL13 expression was associated with favorable prognosis in UCEC. A nomogram for predicting the 1-, 3-, and 5-year survival probability in UCEC was construct based on CXCL13 expression and other clinical parameters. The use of Spearman correlation indicated certain correlation between CXCL13 and immune cells and immune checkpoint (ICP) genes. Seven hub genes were upregulated in UCEC, namely CXCL9, IFNG, CXCL10, CXCL11, GBP5, CCL18, and GZMB. The expression and prognostic relevance of CXCL9, IFNG, GBP5, and GZMB were in accordance with CXCL13. The main biological processes enriched were cytokine-cytokine receptor interaction and chemokine signaling pathway. CONCLUSIONS: The above comprehensive analyses suggest that CXCL13 may serve as a potential prognostic biomarker for UCEC, specifically for early-stage UCEC.

CXCL13↗

Isolation, characterization, and bioinformatic analysis of calmodulin-binding protein cmbB reveals a novel tandem IP22 repeat common to many Dictyostelium and Mimivirus proteins.

A novel calmodulin-binding protein cmbB from Dictyostelium discoideum is encoded in a single gene. Northern analysis reveals two cmbB transcripts first detectable at 4 h during multicellular development. Western blotting detects an approximately 46.6 kDa protein. Sequence analysis and calmodulin-agarose binding studies identified a "classic" calcium-dependent calmodulin-binding domain (179IPKSLRSLFLGKGYNQPLEF198) but structural analyses suggest binding may not involve classic alpha-helical calmodulin-binding. The cmbB protein is comprised of tandem repeats of a newly identified IP22 motif ([I,L]Pxxhxxhxhxxxhxxxhxxxx; where h = any hydrophobic amino acid) that is highly conserved and a more precise representation of the FNIP repeat. At least eight Acanthamoeba polyphaga Mimivirus proteins and over 100 Dictyostelium proteins contain tandem arrays of the IP22 motif and its variants. cmbB also shares structural homology to YopM, from the plague bacterium Yersenia pestis.

Amino Acid Sequence↗

Transcriptional and bioinformatic analysis of the 56.8 kb DNA region amplified in tandem repeats containing the penicillin gene cluster in Penicillium chrysogenum.

High penicillin-producing strains of Penicillium chrysogenum contain 6-14 copies of the three clustered structural biosynthetic genes, pcbAB, pcbC, and penDE [Barredo, J.L., Díez, B., Alvarez, E., Martín, J.F., 1989. Large amplification of a 35-kb DNA fragment carrying two penicillin biosynthetic genes in high penicillin producing strains of Penicillium chrysogenum. Curr. Genet. 16, 453-459; Smith, D.J., Bull, J.H., Edwards, J., Turner, G., 1989. Amplification of the isopenicillin N synthetase gene in a strain of Penicillium chrysogenum producing high levels of penicillin. Mol. Gen. Genet. 216, 492-497.] . The cluster is located in a 56.8 kb DNA region bounded by a conserved TGTAAA/T hexanucleotide that undergoes amplification in tandem repeats [Fierro, F., Barredo, J.L., Díez, B., Gutiérrez, S., Fernández, F.J., Martín, J.F., 1995. The penicillin gene cluster is amplified in tandem repeats linked by conserved hexanucleotide sequences. Proc. Natl. Acad. Sci. USA 92, 6200-6204; Newbert, R.W., Barton, B., Greaves, P., Harper, J., Turner, G., 1997. Analysis of a commercially improved Penicillium chrysogenum strain series: involvement of recombinogenic regions in amplification and deletion of the penicillin biosynthesis gene cluster. J. Ind. Microbiol. Biotechnol. 19, 18-27]. Transcriptional analysis of this amplified region (AR) revealed the presence of at least eight transcripts expressed in penicillin producing conditions. Three of them correspond to the known penicillin biosynthetic genes, pcbAB, pcbC, and penDE. To locate genes related to penicillin precursor formation, or penicillin transport and regulation we have sequenced and analyzed the 56.8 kb amplified region of P. chrysogenum AS-P-78, finding a total of 16 open reading frames. Two of these ORFs have orthologues of known function in the databases. Other ORFs showed similarities to specific domains occurring in different proteins and superfamilies which allowed to infer their probable function. ORF11 encodes a D-amino acid oxidase that might be responsible for the conversion of D-amino acids in the tripeptide L-alpha-aminoadipyl-L-cysteinyl-D-valine or other beta-lactam intermediates to deaminated by-products. ORF12 encodes a predicted protein with similarity to saccharopine dehydrogenases that seems to be related to biosynthesis of the penicillin precursor alpha-aminoadipic acid. A deletion mutant, P. chrysogenum npe10 lacking the entire AR including ORF12, shows a partial requirement of L-lysine for growth. ORF13 encodes a putative protein containing a Zn(II)2-Cys6 fungal-type DNA-binding domain, probably a transcriptional regulator. Although some of the ORFs in the AR may play roles in increasing penicillin production, none of the 13 ORFs other than pcbAB, pcbC, and penDE seem to be strictly indispensable for penicillin biosynthesis. The genes located in the P. chrysogenum AR have been compared with those found in the Aspergillus nidulans 50 kb DNA region adjacent to the penicillin gene cluster, showing no conservation between these two fungi.

Amino Acid Sequence↗