PubMed Health⌕ Search

Biomedical subjects

Jan C Biro

Publications and source records attributed to Jan C Biro.

6 recordsLinked to original sources

A novel intra-molecular protein-protein interaction code based on partial complementary coding of co-locating amino acids.

Proteins are assumed to contain all the information necessary for unambiguous folding and specific interaction with each other. However, ab initio structure prediction is often not successful because the amino acid sequence itself is simply not sufficient to guide between endless folding possibilities. It seems to be logical to try to find the "missing" information in nucleic acids, in the redundant codon. Statistical analyses of approximately 35K amino acid co-locations in 80 different protein structures indicate the existence of a weak intra-molecular protein-protein interaction code. Co-locating amino acids are preferentially coded by codons which are complementary in reverse orientation to each other at the 1st and 3rd codon positions, but not necessarily at the 2nd. This code, called D-1 X 3/RC-3 X 1, limits the number of preferred amino acid pairs from 20 to 10.3+/-0.8 (SEM, n=20) and emphasizes the importance of "strictly" defined amino acids (those having less synonymous codons). The existence of this code does not by any means violate the known physicochemical rules of protein folding or interaction. It is suggested that the biological source of preferential (specific) amino acid co-locations is the partial complementarity of their codons. This special coding of co-locating amino acids is important to better understanding of some fundamental biochemical processes and observations such as: (a) protein folding; (b) specific and high affinity protein-protein interactions; (c) the role of the wobble bases; (d) the significance of the redundant genetic code; (e) the origin of specific protein-protein interactions. Furthermore it might be useful even in protein design.

Amino Acid Sequence↗

Nucleic acid chaperons: a theory of an RNA-assisted protein folding.

BACKGROUND: Proteins are assumed to contain all the information necessary for unambiguous folding (Anfinsen's principle). However, ab initio structure prediction is often not successful because the amino acid sequence itself is not sufficient to guide between endless folding possibilities. It seems to be a logical to try to find the "missing" information in nucleic acids, in the redundant codon base. RESULTS: mRNA energy dot plots and protein residue contact maps were found to be rather similar. The structure of mRNA is also conserved if the protein structure is conserved, even if the sequence similarity is low. These observations led me to suppose that some similarity might exist between nucleic acid and protein folding. I found that amino acid pairs, which are co-located in the protein structure, are preferentially coded by complementary codons. This codon complementarity is not perfect; it is suboptimal where the 1st and 3rd codon residues are complementary to each other in reverse orientation, while the 2nd codon letters may be, but are not necessarily, complementary. CONCLUSION: Partial complementary coding of co-locating amino acids in protein structures suggests that mRNA assists in protein folding and functions not only as a template but even as a chaperon during translation. This function explains the role of wobble bases and answers the mystery of why we have a redundant codon base.

Amino Acids↗

SeqX: a tool to detect, analyze and visualize residue co-locations in protein and nucleic acid structures.

BACKGROUND: The interacting residues of protein and nucleic acid sequences are close to each other - they are co-located. Structure databases (like Protein Data Bank, PDB and Nucleic Acid Data Bank, NDB) contain all information about these co-locations; however it is not an easy task to penetrate this complex information. We developed a JAVA tool, called SeqX for this purpose. RESULTS: SeqX tool is useful to detect, analyze and visualize residue co-locations in protein and nucleic acid structures. The user: a. selects a structure from PDB; b. chooses an atom that is commonly present in every residues of the nucleic acid and/or protein structure(s). c. defines a distance from these atoms (3-15 A). The SeqX tool detects every residue that is located within the defined distances from the defined "backbone" atom(s); provides a DotPlot-like visualization (Residues Contact Map), and calculates the frequency of every possible residue pairs (Residue Contact Table) in the observed structure. It is possible to exclude +/- 1 to 10 neighbor residues in the same polymeric chain from detection, which greatly improves the specificity of detections (up to 60% when tested on dsDNA). Results obtained on protein structures showed highly significant correlations with results obtained from literature (p < 0.0001, n = 210, four different subsets). The co-location frequency of physico-chemically compatible amino acids is significantly higher than is calculated and expected in random protein sequences (p < 0.0001, n = 80). CONCLUSION: The tool is simple and easy to use and provides a quick and reliable visualization and analyses of residue co-locations in protein and nucleic acid structures. AVAILABILITY AND REQUIREMENTS: http://janbiro.com/Downloads.html SeqX, Java J2SE Runtime Environment 5.0 (available from [see Additional file 1] http://www.sun.com) and at least a 1 GHz processor and with a minimum 256 Mb RAM. Source codes are available from the authors.

Amino Acid Sequence↗

Frequent occurrence of recognition site-like sequences in the restriction endonucleases.

BACKGROUND: There are two different theories about the development of the genetic code. Woese suggested that it was developed in connection with the amino acid repertoire, while Crick argued that any connection between codons and amino acids is only the result of an "accident". This question is fundamental to understand the nature of specific protein-nucleic acid interactions. RESULTS: The nature of specific protein-nucleic acid interaction between restriction endonucleases (RE) and their recognition sequences (RS) was studied by bioinformatics methods. It was found that the frequency of 5-6 residue long RS-like oligonucleotides is unexpectedly high in the nucleic acid sequence of the corresponding RE (p < 0.05 and p < 0.001 respectively, n = 7). There is an extensive conservation of these RS-like sequences in RE isoschizomers. A review of the seven available crystallographic studies showed that the amino acids coded by codons that are subsets of recognition sequences were often closely located to the RS itself and they were in many cases directly adjacent to the codon-like triplets in the RS.Fifty-five examples of this codon-amino acid co-localization are found and analyzed, which represents 41.5% of total 132 amino acids which are localized within 8 A distance to the C1' atoms in the DNA. The average distance between the closest atoms in the codons and amino acids is 5.5 +/- 0.2 A (mean +/- S.E.M, n = 55), while the distance between the nitrogen and oxygen atoms of the co-localized molecules is significantly shorter, (3.4 +/- 0.2 A, p < 0.001, n = 15), when positively charged amino acids are involved. This is indicating that an interaction between the nucleic- and amino acids might occur. CONCLUSION: We interpret these results in favor of Woese and suggest that the genetic code is "rational" and there is a stereospecific relationship between the codes and the amino acids.

Amino Acids↗

A novel sequence similarity searching and visualization method based on overlappingly translated nucleic acids: the blastNP.

Sequence data are stored in nucleic acid and protein databases. Searching the nucleic acid databases is very specific but rather insensitive method. Searching protein databases is sensitive but not very specific procedure. It was expected that the combination of these methods might provide an optimal approach. Therefore an alternative method to TblastX has been developed, known as blastNP. Nucleic acids in database and query sequences were translated into overlapping protein-like sequences (overlappingly translated sequences or OTSs) before searching with blastP. Thus, each nucleic acid sequence is represented by a single "protein like" sequence (instead of three hypothetical proteins in different reading frames). The blastNP method is defined as a blastP that is performed on an overlappingly translated nucleic acid database using a similarly converted nucleic acid query. The specificity and sensitivity of blastNP and TblastX is very similar, however blastNP is more sensitive to detect short sequence similarities (less than 50 residues). BlastNP combines the advantages of nucleotide and protein blasts and bypasses many difficulties: (1). it is more sensitive to weak sequence similarities than blastN, (2). codon redundancy is eliminated, (3). the sensitivity to single nucleotide polymorphism, mutation and sequencing errors are reduced, (4). it is insensitive to frame shifts. This novel method was proved to find significant sequence similarities which remained hidden for other methods and is a promising tool for further understanding (and annotating) the function of many old and new sequences.

Amino Acid Sequence↗

The prion paradox: infection or polymerisation?

A weak but significant similarity was found between the prion protein (PrP) and some transcription factors and zinc-finger proteins. A possible interpretation of this similarity is that the PrP is a metal- (copper-) binding transcription factor and might behave like a Zn-finger protein, with the Cu2+ binding to its histidine and serine residues. Copper-binding could create intramolecular bridges in the PrPC (normal, cytoplasmic) molecule, but intermolecular bridges in the PrPSc (scrapie pathogenic) molecule. A molecular model of the Cu2+ -binding monomeric PrPC and the Cu2+-stabilised polymeric PrP Sc is presented here and named the 'cuprion model'. In this model, the PrPC has two idioforms. The stable, normal PrPC-idioform-I contains Cu2+ in -2His-Cu2+-2His- complexes and an intramolecular disulphide bridge. An unstable, transient form, PrPC-idioform-II, contains a -2His-Cu 2+-2Cys- complex, which destabilises the intramolecular disulphide bridge and makes the PrPC molecule highly reactive with other PrPC molecules.

Amino Acid Sequence↗