PubMed Health⌕ Search

Biomedical subjects

A E Kister

Publications and source records attributed to A E Kister.

At least 19 recordsLinked to original sources

Strict rules determine arrangements of strands in sandwich proteins.

From a computer analysis of the spatial organization of the secondary structures of beta-sandwich proteins, we find certain sets of consecutive strands that are connected by hydrogen bonds, which we call "strandons." The analysis of the arrangements of strandons in 491 protein structures that come from 69 different superfamilies reveals strict regularities in the arrangements of strandons and the formation of what we call "canonical supermotifs." Six such supermotifs account for approximately 90% of all observed structures. Simple geometric rules are described that dictate the formation of these supermotifs.

Amino Acid Motifs↗

A geometric construction determines all permissible strand arrangements of sandwich proteins.

For a large class of proteins called sandwich-like proteins (SPs), the secondary structures consist of two beta-sheets packed face-to-face, with each beta-sheet consisting typically of three to five beta-strands. An important step in the prediction of the three-dimensional structure of a SP is the prediction of its supersecondary structure, namely the prediction of the arrangement of the beta-strands in the two beta-sheets. Recently, significant progress in this direction was made, where it was shown that 91% of observed SPs form what we here call "canonical motifs." Here, we show that all canonical motifs can be constructed in a simple manner that is based on thermodynamic considerations and uses certain geometric structures. The number of these structures is much smaller than the number of possible strand arrangements. For instance, whereas for SPs consisting of six strands there exist a priori 900 possible strand arrangements, there exist only five geometric structures. Furthermore, the few motifs that are noncanonial can be constructed from canonical motifs by a simple procedure.

Amino Acid Motifs↗

Prediction of the structural motifs of sandwich proteins.

We investigate the supersecondary structure of a large group of proteins, the so-called sandwich proteins. The analysis of a large number of such proteins has led us to propose a set of rules that can be used to predict the possible arrangements of strands in the two beta-sheets forming a given sandwich structure. These rules imply the existence of certain invariant supersecondary substructures common to all sandwich proteins. Furthermore, they dramatically restrict the number of permissible arrangements. For example, whereas for proteins consisting of three strands in each beta-sheet 180 possible strand arrangements exist a priori, our rules imply that only 15 of them are permissible. Five of these predicted arrangements describe all currently known sandwich proteins with six strands.

Amino Acid Sequence↗

The sequence determinants of cadherin molecules.

The sequence and structural analysis of cadherins allow us to find sequence determinants-a few positions in sequences whose residues are characteristic and specific for the structures of a given family. Comparison of the five extracellular domains of classic cadherins showed that they share the same sequence determinants despite only a nonsignificant sequence similarity between the N-terminal domain and other extracellular domains. This allowed us to predict secondary structures and propose three-dimensional structures for these domains that have not been structurally analyzed previously. A new method of assigning a sequence to its proper protein family is suggested: analysis of sequence determinants. The main advantage of this method is that it is not necessary to know all or almost all residues in a sequence as required for other traditional classification tools such as BLAST, FASTA, and HMM. Using the key positions only, that is, residues that serve as the sequence determinants, we found that all members of the classic cadherin family were unequivocally selected from among 80,000 examined proteins. In addition, we proposed a model for the secondary structure of the cytoplasmic domain of cadherins based on the principal relations between sequences and secondary structure multialignments. The patterns of the secondary structure of this domain can serve as the distinguishing characteristics of cadherins.

Algorithms↗

Class-defining characteristics in the mouse heavy chains of variable domains.

Analysis of residue correlation in over 2700 mouse heavy chains of the V(H) domains was carried out on three hierarchical levels. At the 'position' level, statistical analysis revealed 45 positions that conserve similar residues in almost all chains. At the 'fragment' level, the focus of investigation shifted to the study of combinations of amino acids in strands and loops. It was found that no more than 10 patterns were sufficient for describing strands and loops in the chains. At the 'sequence' level, we determined all possible combinations of these patterns and classified the mouse heavy chains. Comparison of the sequences in the eight classes revealed residues at the class-determining positions that were unique to each class. Because a strong correlation of residues was found, one only needs several residues to classify a sequence. It follows that no all residue alignment procedure is necessary to divide sequences into classes. An important corollary of our approach is the possibility of predicting residues in an incomplete sequence from a small sequence fragment. On the basis of our analysis of mouse heavy chains we hypothesize about the presently unknown mouse V(H) germline repertoire.

Amino Acid Sequence↗

Predicting amino acid sequences of the antibody human VH chains from its first several residues.

A new method for classification of Ig sequences is suggested. The defining characteristic of a class is presence of particular residues at several class-determining positions. Sequences within a class follow the same amino acid pattern, i.e., residues at identical positions are, in an overwhelming majority of sequences of that class, identical or chemically related. Thus, once the class of a sequence is determined, one can predict the residue(s) at almost any position in the sequence. In this paper, results of analysis of 1,172 human heavy chains are presented. It was shown that a sequence can be assigned to one of six classes depending on which residues are found at its positions 1, 3, 5, 6, 7, 9, 10, 12, and 13. It is important to note that it is possible to achieve same six-class classification of the human heavy chains on the basis of a different set of positions found not at the beginning but near the end of the sequence (around position 80). For every class, an amino acid pattern of an entire sequence (complementarity determining regions excepting) has been determined. Our approach allowed us to reconstruct the incomplete human heavy chains in which residues at certain positions at the beginning or end of the chain are known. We developed a software tool for analysis, classification, and prediction of residues in sequences of the Ig family.

Algorithms↗

A very limited number of keywords (main patterns) describes all sequences of the human variable heavy (VH) and kappa (Vkappa) domains.

Sequences of the variable heavy (VH) and kappa (Vkappa) domains of Ig structures were divided into 21 fragments that correspond to strands, loops, or parts of these structural units of the variable domains. Amino acid sequences of fragments (termed "words") were collected from the 1,172 human heavy and 668 human kappa chains available in the Kabat database. Statistical analysis of words of 17 fragments was performed (fragments that comprise the complementary determining regions' fragments will not be discussed in this paper). The number of different words (those with different residues in at least one position) ranged, for various fragments, from 11 to 75 in the kappa chains, and from 23 to 189 in the heavy chains. The main result of this study is that very few keywords, or main patterns of words, were necessary to describe over 90% of the sequences (no more than two keywords per fragment in the kappa and no more than five per fragment in the heavy chains). No identical keywords were found for different fragments of the variable domains. Keywords of aligned fragments of the VH and Vkappa domains were different in all but two instances. Thus, knowing the keywords, one can determine whether any given small part of a sequence belongs to a heavy or kappa chain and predict its precise localization in the sequence. In addition, by using all of the keywords obtained through analysis of the Kabat database, it was possible to describe completely the sequences of the human VH and Vkappa germ-line segments.

Computer Simulation↗

The invariant system of coordinates of antibody molecules: prediction of the "standard" C alpha framework of VL and VH domains.

A new approach of comparing protein structures that does not involve the procedure of superposition is suggested. An invariant system of coordinates for immunoglobulin molecules that is based on the geometrical symmetry inherent to the variable domain light-chain (VL)-heavy-chain (VH) complex is described. The coordinates of the Calpha atoms in 22 immunoglobulin structures are calculated in the invariant system of coordinates. We found that 76 identical positions in this Calpha framework are symmetrical about the twofold axis. Comparison of the identical positions in these molecules allows us to select 96 positions in the light chains and 87 positions in the heavy chains whose Calpha atom coordinates are approximately the same. To check whether the average coordinates of Calpha atoms in these positions complies with the stereochemical requirements, we calculated Calpha-Calpha distances. Seventy-three positions of the light chains and 72 positions of the heavy chains satisfy the Calpha-Calpha distance criterion. The Calpha atoms in these positions are used for constructing the "standard" Calpha framework of VL and VH complexes. The average coordinates of Calpha atoms are presented.

Animals↗

Analysis of the relation between the sequence and secondary and three-dimensional structures of immunoglobulin molecules.

Methods of structural and statistical analysis of the relation between the sequence and secondary and three-dimensional structures are developed. About 5000 secondary structures of immunoglobulin molecules from the Kabat data base were predicted. Two statistical analyses of amino acids reveal 47 universal positions in strands and loops. Eight universally conservative positions out of the 47 are singled out because they contain the same amino acid in > 90% of all chains. The remaining 39 positions, which we term universally alternative positions, were divided into five groups: hydrophobic, charged and polar, aromatic, hydrophilic, and Gly-Ala, corresponding to the residues that occupied them in almost all chains. The analysis of residue-residue contacts shows that the 47 universal positions can be distinguished by the number and types of contacts. The calculations of contact maps in the 29 antibody structures revealed that residues in 24 of these 47 positions have contacts only with residues of antiparallel beta-strands in the same beta-sheet and residues in the remaining 23 positions always have far-away contacts with residues from other beta-sheets as well. In addition, residues in 6 of the 47 universal positions are also involved in interactions with residues of the other variable or constant domains.

Amino Acid Sequence↗

Murine common acute lymphoblastic leukemia antigen (CD10 neutral endopeptidase 24.11). Molecular characterization, chromosomal localization, and modeling of the active site.

To further analyze CD10/NEP function in lymphoid and nonlymphoid cells using well characterized murine systems, we isolated the murine CD10/NEP homologue, determined its chromosomal location, and modeled the enzyme's active site. The murine CD10/NEP cDNA predicts a 750-amino acid (aa) type II integral membrane protein with 90% identity to the human CD10 sequence and 100% conservation of critical aa and functional motifs. The latter include the pentapeptide consensus sequence required for zinc binding and catalytic activity, additional aa associated with substrate binding, and the extracellular cysteines that participate in disulfide bonds required for enzymatic activity. Like its human homologue, murine CD10/NEP has multiple alternative 5'-untranslated region sequences. The gene is localized on the proximal half of murine chromosome 3. In Northern analysis, murine CD10/NEP transcripts are abundant in bone marrow stromal cells that support pre-B cell differentiation but are undetectable in representative Abelson transformed pre-B cell lines. The murine CD10/NEP active site was modeled by aligning critical conserved CD10/NEP residues with comparable residues in the active site of thermolysin, a bacterial metalloprotease with similar substrate specificity. The model predicts that the two enzymes have similar clefts that comprise the active site and permit zinc-dependent substrate interactions.

Amino Acid Sequence↗

A kinetic approach to the prediction of RNA secondary structures.

A new approach to the prediction of secondary RNA structures based on the analysis of the kinetics of molecular self-organisation is proposed herein. The Markov process is used to describe structural reconstructions during secondary structure formation. This process is modelled by a Monte-Carlo method. Examples of the calculation by this method of the secondary structures kinetic ensemble are given. Distribution of time-dependent probabilities within the ensembles is obtained. An effective method for search for the equilibrium ensemble is also suggested. This method is based on the construction of a tree of all possible secondary structures of RNA. By ascribing a probability for each structure (according to its free energy) the Boltzmann equilibrium ensemble can be obtained.

Computers↗

[A theoretical analysis of structural restructuring during formation of secondary RNA structures].

An improved method for predicting the RNA secondary structure is proposed. The process of self-organization of structure is considered as a Markov chain. The kinetic of secondary structure is analysed by the Monte-Carlo method. The topological compatibility of helices is discussed. On the base of analysis it follows that the dynamical process of secondary structure formation is so that it is impossible to define a static set of complementary pairs. The method was used for predicting the mRNA secondary structure of a series of recombinant plasmids, containing the cro gene. The observed variation in expression can be explained by secondary structure.

Computer Simulation↗

[Computer programs for the analysis of nucleotide sequences (MALK)].

A system for the computer analysis of nucleic acid and protein sequences ("Helix") is described. Format of the DNA sequences is EMBL--compatible and may be easily commented with the help of convenient menus. "Helix" has also following possibilities: an effective alignment of gele reading data and formation of the final sequence; simple making of recombined molecules "in calcular"; calculations of nucleotide and dinucleotide distribution along the sequence; looking for coding frames; calculations percentage of codons and amino acids in coding frames; searching for direct and inverted repeats; sequences alignment; protein secondary structure prediction; restriction mapping; DNA--protein translation. "Helix" also contain programs for RNA-structure prediction, looking for homologies throughover the EMAL bank, choosing optimal sequence for probes and searching promoters. All the programs are written at FORTRAN-77 and automatically translated into FORTRAN-4. "Helix" require only 64 kbite.

Base Sequence↗