Bioinformatics in the pre- and post-genomic eras.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to H A Lim.
Explore the source record for details and available documents.
Splice junction shadows (ancient exon-exon junctions) presumably reflect the existence of amino acid primary blocks which were used in the course of evolution for the construction of new proteins. The lengths of such blocks (i.e. regions between splice junctions), as the lengths of corresponding inserted or duplicated ancient exons, should be divisible by three in order to store the preexisting coding frame in the course of evolution. In this paper, we will test the hypothesis of intron-mediated recombination in a model of block molecular evolution (exon shuffling) by revealing corresponding blocks in existing database-contained coding sequences. For this purpose, we use a weight matrix prediction of ancient splice junction shadows in coding regions of the nucleotide sequences in current databases. The usage of splice junction shadows allows us to test the block evolution hypothesis in better detail in comparison with previous methods which were based only on currently existing recent exons. Our result of block length distribution at the nucleotide level shows a clear tendency to be divisible by three. At the protein level, several unexpected favorable block lengths, which are six, nine, 12 and 15 amino acids in length, were observed. Further refinements in our method for revealing splice junction shadows (structural block boundaries) might reveal peptides which probably maintain stable folds in different structures. The latter can in turn be used for protein structure prediction.
We have created an algorithm for compressing a PIR database to assist individual researchers and software developers who utilize sequence database information but may not have huge storage space. The resulting compact databank contains compressed PIR information and an interface written in C which allows fast direct access to the stored information without extensive decompression of corresponding files. The databank files as well as the interface C-file can be used on both PC-compatibles and UNIX-based computers without any modifications. The interface supports all standard PIR Request Network queries (i.e. gets databank SEQ number by entry; for a defined databank SEQ number, gets specified information like: name, organism(s), keyword(s), sequence, sequence features with coordinates, etc.). In contrast with PIR Request Network, our package allows us to call PIR-contained information directly from the C programs, even on a personal computer not on a network. Our PIR-derived databank, SAGITTARIUS PIR, was implemented in the form of separate file sets. Each file set contains database information of independent types (i.e. sequences, entry indexes, organisms, etc.). On a particular computer, the available configuration of the PIR information (and storage space) can be easily changed as needed by the user without affecting retrievals of other types of stored information. Due to an original alignment-based algorithm, in the compression of protein sequences themselves, our package out-performs the well-known ZIP file compressor. For PC-compatibles, a dialogue shell is available which supports all standard PIR Request Network queries plus homology searches, alignments, etc.
A combinatorial sequence space (CSS) model was introduced to represent sequences as a set of overlapping k-tuples of some fixed length which correspond to points in the CSS. The aim was to analyze clusterization of protein sequences in the CSS and to test various hypotheses about the possible evolutionary basis of this clusterization. The authors developed an easy-to-use technique which can reveal and analyze such a clusterization in a multidimensional CSS. Application of the technique led to an unexpectedly high clusterization of points in the CSS corresponding to k-tuples from known proteins. The clusterization could not be inferred from nonuniform amino acid frequencies or be explained by the influence of homologous data. None of the tested possible evolutionary and structural factors could explain the clusterization observed either. It looked as if certain protein sequence variations occurred and were fixed in the early course of evolution. Subsequent evolution (predominantly neutral) allowed only a limited number of changes and permitted new variants which led to preservation of certain k-tuples during the course of evolution. This was consistent with the theory of exon shuffling and protein block structure evolution. Possible applications of sequence space features found were also discussed.
A new algorithm for data bank homology search is proposed. The principal advantages of the new algorithm are: (i) linear computation complexity; (ii) low memory requirements; and (iii) high sensitivity to the presence of local region homology. The algorithm first calculates indicative matrices of k-tuple 'realization' in the query sequence and then searches for an appropriate number of matching k-tuples within a narrow range in database sequences. It does not require k-tuple coordinates tabulation and in-memory placement for database sequences. The algorithm is implemented in a program for execution on PC-compatible computers and tested on PIR and GenBank databases with good results. A few modifications designed to improve the selectivity are also discussed. As an application example, the search for homology of the mouse homeotic protein HOX 3.1 is given.
A well-drawn picture acts as an excellent metaphor for something real, and human vision provides instant, random access to any part of which the picture represents. It is in this sense that pictures can convey information more effectively than words alone. The power of the graphics work-stations available today makes visual presentation of scientific results a reality. A molecular graphics program for investigating protein structures, as well as several sample plots that show the power of the program, are presented.
A theoretical analysis of the reptational motion of DNA in a gel that includes the effects of molecular fluctuations has been used to explain the main features found in experiments involving periodic inversion of the electric field. The resonance-like decrease of the electrophoretic mobility as a function of pulse duration is related to transient "undershoots" in the orientation of the molecule, in agreement with recent experimental data. These features arise from a delicate interplay of internal and center of mass motion of the molecules under pulsed field conditions, and are important for the separation of DNA molecules in the size range 0.2 to 10 million base pairs.