PubMed Health⌕ Search

Biomedical subjects

Liaofu Luo

Publications and source records attributed to Liaofu Luo.

7 recordsLinked to original sources

Protein structure preference, tRNA copy number, and mRNA stem/loop content.

From statistical analyses of protein sequences for humans and Escherichia coli we found that the messenger RNA segment of m-codons (for m=2 to 6) with average high tRNA copy number (TCN) (larger than approximately 10.5 for humans or approximately 1.95 for E. coli) preferably code for the alpha helix and that with low TCN (smaller than approximately 7.5 for humans or approximately 1.7 for E. coli) preferably code for coil. Between them there is an intermediate region without correlation to structure preference. For the beta strand the preference/ avoidance tendency is not obvious. All strong preference-modes of TCN for protein secondary structures have been deduced. The mutual interaction between two factors--protein secondary structural type and codon TCN--is tested by F distribution. A phenomenological model on the relation between structure preference and translational efficiency or accuracy is proposed. It is pointed out that the structure preference of codons is related to the distribution of mRNA stem/loop content in three TCN regions.

Bacterial Proteins↗

Statistical correlation between protein secondary structure and messenger RNA stem-loop structure.

A new integrated sequence-structure database, called IADE (Integrated ASTRAL-DSSP-EMBL), incorporating matching mRNA sequence, amino acid sequence, and protein secondary structural data, is constructed. It includes 648 protein domains. Based on the IADE database, we studied the relation between RNA stem-loop frequencies and protein secondary structure. It was found that the alpha-helices and beta-strands on proteins tend to be preferably "coded" by mRNA stem region, while the coils on proteins tend to be preferably "coded" by mRNA loop region. These tendencies are more obvious if we observe the structural words (SWs). An SW is defined by a four-amino-acid-fragment that shows the pronounced secondary structural (alpha-helix or beta-strand) propensity. It is demonstrated that the deduced correlation between protein and mRNA structure can hardly be explained as the stochastic fluctuation effect.

Databases as Topic↗

Splice site prediction with quadratic discriminant analysis using diversity measure.

Based on the conservation of nucleotides at splicing sites and the features of base composition and base correlation around these sites we use the method of increment of diversity combined with quadratic discriminant analysis (IDQD) to study the dependence structure of splicing sites and predict the exons/introns and their boundaries for four model genomes: Caenorhabditis elegans, Arabidopsis thaliana, Drosophila melanogaster and human. The comparison of compositional features between two sequences and the comparison of base dependencies at adjacent or non-adjacent positions of two sequences can be integrated automatically in the increment of diversity (ID). Eight feature variables around a potential splice site are defined in terms of ID. They are integrated in a single formal framework given by IDQD. In our calculations 7 (8) base region around the donor (acceptor) sites have been considered in studying the conservation of nucleotides and sequences of 48 bp on either side of splice sites have been used in studying the compositional and base-correlating features. The windows are enlarged to 16 (donor), 29 (acceptor) and 80 bp (either side) to improve the prediction for human splice sites. The prediction capability of the present method is comparable with the leading splice site detector--GeneSplicer.

Algorithms↗

Minimal model for genome evolution and growth.

Textual analysis of typical microbial genomes reveals that they have the statistical characteristics of a DNA sequence of a much shorter length. This peculiar property supports an evolutionary model in which a genome evolves by random mutation but primarily grows by random segmental duplication. That genomes grew mostly by duplication is consistent with the observation that repeat sequences in all genomes are widespread and intragenomic and intergenomic homologous genes are preponderant across all life forms.

DNA, Bacterial↗

Coding rules for amino acids in the genetic code: the genetic code is a minimal code of mutational deterioration.

Coding rules for amino acids in the genetic code are discussed from the point that the genetic code is a minimal code of mutational deterioration. The global mutational deterioration (GMD) function is defined through several parameters describing single base mutations and amino acid distances. The problem of searching for the global minimum of the GMD function is discussed in some detail. From GMD minimization under initial constraints we have succeeded in deducing the standard genetic code.

Amino Acids↗

Sequence-dependent flexibility in promoter sequences.

The non-neighbor interactions between base-pairs were taken into account to calculate the angular parameters (Omega, rho and tau) describing the orientation of successive base-pair planes and the translation parameters (D(y)) along the long axis of base-pair steps for 36 independent tetramers. A statistical mechanical model was proposed to predict the DNA flexibility that is mainly related to the thermal fluctuations at individual base-pair steps. The DNA flexibility can be described by the root-mean-square deviation of the end-to-end distance of DNA helical structure. The present model was then used to investigate the extreme flexible pattern in prokaryotic and eukaryotic promoter sequences. The results demonstrated several extreme flexible regions related to functionally important elements exist both in prokaryotic promoters and in eukaryotic promoters, DNA flexibility and AT content are highly correlated. The probabilities finding flexibility pattern in promoter sequences were also estimated statistically. The biological implications were discussed briefly.

Animals↗

Construction of genetic code from evolutionary stability.

The construction of the genetic code is investigated based on a stability principle. The concept and formulation of mutational deterioration (MD) of the genetic code is proposed. It is proved that the degeneracies of codon multiplets obey the rule to best resist MD. The MD for each ideal multiplet of codons is expressed by four parameters and it takes on a minimum value for real distributions of codons in the multiplet. Then the global mutational deterioration (GMD) of code table is calculated and the minimal code is deduced. The domain-like distribution of hydrophobic and hydrophilic amino acids on the genetic code is explained from the minimization of GMD. It is demonstrated that the standard code is approximately GMD-minimal. By introducing some constraints that are related to the initial condition of the system, we have deduced the standard genetic code from the minimization of GMD. The minimization shows the general trend of evolutionary process to some stable state while the constraints reflect a 'frozen accident.' Many deviant codon assignments are also explained through MD minimization assuming the changeable degrees of degeneracies for some multiplets. So, a possible answer to the question of "Why are synonymous codons and amino acids distributed in the code table just as they are?" is given.

Biological Evolution↗