PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Progress of 1D protein structure prediction at last.

Accuracy of predicting protein secondary structure and solvent accessibility from sequence information has been improved significantly by using information contained in multiple sequence alignments as input to a neural network system. For the Asilomar meeting, predictions for 13 proteins were generated automatically using the publicly available prediction method PHD. The results confirm the estimate of 72% three-state prediction accuracy. The fairly accurate predictions of secondary structure segments made the tool useful as a starting point for modeling of higher dimensional aspects of protein structure.

Amino Acid Sequence↗

Purification, characterization, and primary structure of four depressant insect-selective neurotoxin analogs from scorpion (Buthus sindicus) venom.

Four depressant insect-selective neurotoxin analogs (termed Bs-dprIT1 to 4) from the venom of the scorpion Buthus sindicus were purified to homogeneity in a single step using reverse-phase HPLC. The molecular masses of the purified toxins were 6820.9, 6892.4, 6714.7, and 6657.1 Da, respectively, as determined by mass spectrometry. These long-chain neurotoxins were potent against insects with half lethal dose values of 67, 81, 103, and 78 ng/100 mg larva and 138, 160, 163, and 142 ng/100 mg cockroach, respectively, but were not lethal to mice even at the highest applied dose of 10 microg/20 g mouse. When injected into blowfly larvae (Sarcophaga falculata), Bs-dprIT1 to 4 induced classical manifestations of depressant toxins, i.e., a slow depressant flaccid paralysis. The primary structures of Bs-dprIT 1 to 4 revealed high sequence homology (60-75%) with other depressant insect toxins isolated from scorpion venoms. Despite the high sequence conservation, Bs-dprIT1 to 4 showed some remarkable features such as (i) the presence of methionine (Met(6) in Bs-dprIT1 and Met(24) in Bs-dprIT2 to 4) and histidine (His(53) and His(57) in Bs-dprIT1) residues, i.e., amino acid residues that are uncommon to this type of toxin; (ii) the substitution of two highly conserved tryptophan residues (Trp43 --> Ala and Trp53 --> His) in the sequence of Bs-dprIT1; and (iii) the occurrence of more positively charged amino acid residues at the C-terminal end than in other depressant insect toxins. Multiple sequence alignment, sequence analysis, sequence-based structure prediction, and 3D homology modeling studies revealed a protein fold and secondary structural elements similar to those of other scorpion toxins affecting sodium channel activation. The electrostatic potential calculated on the surface of the predicted 3D model of Bs-dprIT1 revealed a significant positive patch in the region of the toxin that is supposed to bind to the sodium channel.

Amino Acid Sequence↗

Amino acid sequence and glycosylation of functional unit RtH2-e from Rapana thomasiana (gastropod) hemocyanin.

The complete amino acid sequence of Rapana thomasiana hemocyanin functional unit RtH2-e was determined by direct sequencing and matrix-assisted laser desorption ionization mass spectrometry of peptides obtained by cleavage with EndoLysC proteinase, chymotrypsin, and trypsin. The single-polypeptide chain of RtH2-e consists of 413 amino acid residues and contains two consensus sequences NXS/T (positions 11-19 and 127-129), potential sites for N-glycosylation. Monosaccharide analysis of RtH2-e revealed a carbohydrate content of about 1.1% and the presence of xylose, fucose, mannose, and N-acetylglucosamine, demonstrating that only N-linked carbohydrate chains of high-mannose type seem to be present. On basis of the monosaccharide composition and MALDI-MS analysis of native and PNGase-F-treated chymotryptic glycopeptide fragment of RtH2-e the oligosaccharide Man(5)GlcNAc(2), attached to Asn(127), is suggested. Multiple sequence alignments with other molluscan hemocyanin e functional units revealed an identity of 63% to the cephalopod Octopus dofleini and of 69% to the gastropod Haliotis tuberculata. The present results are discussed in view of the recently determined X-ray structure of the functional unit g of the O. dofleini hemocyanin.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of two novel fission yeast casein kinase-1 isoforms.

The cDNAs for two casein kinase-1 homologs, hhp1 and hhp2, have been isolated from Schizosaccharomyces pombe and characterized. Their corresponding genes reside on chromosomes II and I, respectively, and encode approximately 42-46 kDa proteins that are related structurally to the HRR25 gene product of budding yeast. On the basis of multiple sequence alignment, the CK1 family appears to consist of three main branches. We predict that the branch containing the hhp genes encodes nuclear kinases involved in the regulation of DNA metabolism.

Amino Acid Sequence↗

Sequence analysis of frog rho-crystallin by cDNA cloning and sequencing: a member of the aldo-keto reductase family.

rho-Crystallin is a major enzyme crystallin present in the lenses of amphibian species with a blocked amino terminus. In order to facilitate the determination of the primary sequence of this taxon-specific crystallin, cDNA mixture was synthesized from the poly(A)+mRNA of bullfrog eye lenses. cDNAs encoding rho-crystallin were then amplified by polymerase chain reaction (PCR) using a new protocol of Rapid Amplification of cDNA Ends (RACE). PCR-amplified product corresponding to rho-crystallin was obtained, which was then subcloned into pUC18 vector and then transformed into E. coli strain JM109. Plasmids purified from the positive clones were prepared for nucleotide sequencing by the automatic fluorescence-based dideoxynucleotide chain-termination method. Sequencing more than 15 clones containing DNA inserts coding for rho-crystallin constructed only one unique and complete full-length reading frame of 975 base pairs covering a deduced protein sequence of 324 amino acids including the universal initiating methionine. It shows 96, 59, 46 and 37 percent sequence similarity to another rho-crystallin from European common frog, bovine prostaglandin-F synthase, human aldose reductase and human aldehyde reductase, respectively, revealing the close relationship between rho-crystallins from related amphibian species and its possible evolutionary relatedness with various aldo-keto reductases. In this study a phylogenetic tree for rho-crystallin and related enzymes is constructed based on multiple-sequence alignment program using a combination of distance matrix and approximate parsimony methods. We have thus established the remote phylogenetic relationship between rho-crystallin and some aldehyde/aldose reductases, which may provide a possible link for the recruitment of this crystallin from detoxification-related enzymes and its physiological role in maintaining a transparent and clear lens.

Alcohol Oxidoreductases↗

A template for generation and comparison of three-dimensional selectin models.

We have complemented multiple sequence alignments of the lectin domains of the selectins with an analysis of structurally invariant regions in X-ray structures of the mannose-binding protein (MBP) and E-selectin. The analysis shows that regions of structural conservation between MBP and E-selectin extend beyond regions of rigorous sequence conservation within the selectin family and suggests that reliable three-dimensional models of selectins from different species can be generated by modification of only a few backbone segments in E-selectin. A model of the L-selectin lectin domain is built and discussed with regard to observed differences in selectin specificity.

Amino Acid Sequence↗

Altered subunit communication in subfamilies of trimeric dUTPases.

The enzyme dUTPase is essential in preventing uracil incorporation into DNA. Design of antagonists against this novel chemotherapeutic target requires identification of species-specific differences in the structure and mechanism of the enzyme. This task is now approached via comparisons of available crystallographic structures of dUTPases from Homo sapiens, Escherichia coli, and retroviruses. The eukaryotic protein uniquely displays polar and charged amino acid residues participating in threefold intersubunit interactions. In bacterial and retroviral dUTPases, threefold interactions are mainly hydrophobic. The residues responsible for this contrast are mapped in multiple sequence alignment to positions differently and characteristically conserved in distinct evolutionary branches. The general feature of this contrast is further strengthened by a second eukaryotic model structure constructed using comparative modeling. The dUTPase cDNA from Drosophila melanogaster was identified, sequenced, and the model structure of the encoded polypeptide displayed a polar hydrogen-bonding network of threefold interactions, identically to the human structure. Results allow clear distinction between two subfamilies of trimeric dUTPases where altered subunit communication may account for a functional difference in the catalytic cycle.

Amino Acid Sequence↗

Mutations affecting the calcium-binding site of myeloperoxidase and lactoperoxidase.

Both myeloperoxidase (MPO) and lactoperoxidase (LPO) contain high affinity bound calcium, which has been suggested to play a structural role. Asp-96 in MPO, a residue next to the histidine distal from the heme prosthetic group, has been assigned to the calcium-binding site of the enzyme by X-ray crystallography. Multiple sequence alignment of known animal peroxidases has revealed that the calcium-binding site is highly conserved. In this study, we replaced Asp-96 in MPO and the counterpart Asp-227 in LPO both with Ala by site-directed mutagenesis. The level of peroxidase activity in insect cells infected with recombinant baculoviruses and their culture supernatants was reduced to virtually zero as a result of these mutations. Immunoblotting revealed that these mutant peroxidases were expressed in the cells but not secreted as effectively as the wild-type enzymes. Our findings suggest that a functional calcium-binding site is essential for the biosynthesis of active animal peroxidases.

Amino Acid Sequence↗

Structural modeling and characterization of a thermostable lipase from Bacillus stearothermophilus P1.

The moderate thermophilic bacterium Bacillus stearothermophilus P1 expresses a thermostable lipase that was active and stable at the high temperature. Based on secondary structure predictions and secondary structure-driven multiple sequence alignment with the homologous lipases of known three-dimensional (3-D) structure, we constructed the 3-D structure model of this enzyme and the model reveals the topological organization of the fold, corroborating our predictions. We hypothesized for this enzyme the alpha/beta-hydrolase fold typical of several lipases and identified Ser-113, Asp-317, and His-358 as the putative members of the catalytic triad that are located close to each other at hydrogen bond distances. In addition, the strongly inhibited enzyme by 10 mM PMSF and 1-hexadecanesulfonyl chloride was indicated that it contains a serine residue which plays a key role in the catalytic mechanism. It was also confirmed by site-directed mutagenesis that mutated Ser-113, Asp-317, and His-358 to Ala and the activity of the mutant enzyme was drastically reduced.

Amino Acid Sequence↗

Characterization of a cDNA encoding the bovine coxsackie and adenovirus receptor.

Non-human adenoviruses such as bovine adenovirus type 3 (BAV-3) that do not replicate in human cells but can infect human cells in culture could provide an attractive alternative to human adenoviral vectors for gene therapy. In addition, a large-animal model for genetic diseases can be very useful for the assessment of the efficacy of adenovector-mediated gene delivery in man. Recombinant human subgroup C adenovectors use the coxsackie and adenovirus receptor (CAR) to enter their target cells. Through RT-PCR and sequencing we determined the complete coding sequence of bovine CAR which serves as the primary adenoviral attachment site on bovine cells. A multiple sequence alignment, involving all the previously identified CAR species (man, mouse, rat, pig, and dog) showed that bovine CAR was most related to porcine CAR (92% nucleotide similarity) and demonstrated a highly conserved adenovirus binding Ig1 domain.

Amino Acid Sequence↗

Cloning and expression of a lombricine kinase from an echiuroid worm: insights into structural correlates of substrate specificity.

Phosphagen kinases constitute a large family of enzymes catalyzing the reversible phosphorylation of guanidino acceptor compounds. These guanidino substrates differ substantially in size and chemical properties. In spite of the appearance of X-ray crystal structures for two members of this family, creatine kinase (CK) and arginine kinase (AK), the structural correlates of substrate specificity remain to be fully elucidated. We have determined the cDNA and deduced amino acid sequences for lombricine (guanidinethylphosphoserine) kinase (LK) from the echiuroid worm Urechis caupo and expressed the cDNA in Escherichia coli. The recombinant protein was purified by affinity chromatography and showed high capacity for phosphorylation of lombricine. Phosphagen kinases consist of a small, N-terminal domain and a much larger domain connected by a linker sequence. A key event in catalysis in CK and AK, and certainly all other phosphagen kinases, is a large conformational change involving involving a rotation of the two domains and the movement of two highly conserved flexible loops (one located in the small domain; the other located in the large domain of these enzymes) which clamp down on the substrates. Multiple sequence alignments of Urechis LK with the only other LK sequence available and CK, AK and glycocyamine kinase sequences, confirm the importance of the small flexible loop located in the N-terminal domain of phosphagen kinases as one component of the structural determinants of guanidine specificity. The role of the other flexible loop in the large domain in terms of substrate specificity remains questionable.

Amino Acid Sequence↗

A method for predicting common structures of homologous RNAs.

We have developed a procedure, composed of a set of computer programs, for predicting common RNA structures of homologous sequences. Given a set of homologous RNAs, these programs perform a multiple sequence alignment, generate a list of possible helical stems that are thermodynamically favored in RNA folding from a selected individual sequence, establish a conserved stem list by inspecting the equivalent base pairings and/or conserved helical stems from the derived alignment of homologous RNAs, and build common RNA secondary structures with the maximum scores (i.e., compensatory base changes and number of base pairs, etc.). The approach is a combination of phylogenetic and thermodynamic methods and has been applied to the prediction of common folding structures of the 5' untranslated regions in a number of positive RNA viruses.

Algorithms↗

Analysis of a cDNA sequence encoding the immunoglobulin heavy chain of the Antarctic teleost Trematomus bernacchii.

A spleen cDNA library was constructed from the Antarctic teleost Trematomus bernacchii and immunoscreened with rabbit IgG specific for T. bernacchii Ig heavy chain. Eleven cDNA clones, varying in size and encoding the entire heavy chain or parts of it, were isolated. Here the complete nucleotide and deduced amino acid sequences of clone 2C2 encoding the secretory IgH chain form are reported. Comparison of the amino acid sequence of the entire constant region of the T. bernacchii Ig heavy chain with those from other teleosts and two holostean fish showed percent identity ranging 53.6-60.6%, with the highest values found for Salmoniformes. The multiple sequence alignment revealed the presence of two remarkable insertions: one at the VH-CH1 boundary and a second one, not found in any other IgM heavy chain, localised at the CH2-CH3 boundary. The latter occurred in the region proposed to act as a 'hinge', and resulted in a CH2-CH3 hinge peptide longer than any other IgM hinge. Differences were also found in the number and position of putative N-glycosylation sites of the compared sequences. It is suggested that the unusual features found in the T. bernacchii Ig heavy chain might contribute to the flexibility of the Ig molecule and help understand more about the adaptation of Ig molecules to the polar sea environment.

Amino Acid Sequence↗

Prediction of protein secondary structure at better than 70% accuracy.

We have trained a two-layered feed-forward neural network on a non-redundant data base of 130 protein chains to predict the secondary structure of water-soluble proteins. A new key aspect is the use of evolutionary information in the form of multiple sequence alignments that are used as input in place of single sequences. The inclusion of protein family information in this form increases the prediction accuracy by six to eight percentage points. A combination of three levels of networks results in an overall three-state accuracy of 70.8% for globular proteins (sustained performance). If four membrane protein chains are included in the evaluation, the overall accuracy drops to 70.2%. The prediction is well balanced between alpha-helix, beta-strand and loop: 65% of the observed strand residues are predicted correctly. The accuracy in predicting the content of three secondary structure types is comparable to that of circular dichroism spectroscopy. The performance accuracy is verified by a sevenfold cross-validation test, and an additional test on 26 recently solved proteins. Of particular practical importance is the definition of a position-specific reliability index. For half of the residues predicted with a high level of reliability the overall accuracy increases to better than 82%. A further strength of the method is the more realistic prediction of segment length. The protein family prediction method is available for testing by academic researchers via an electronic mail server.

Mathematical Computing↗

A structural analysis of phosphate and sulphate binding sites in proteins. Estimation of propensities for binding and conservation of phosphate binding sites.

The high resolution X-ray structures of 38 proteins that bind phosphate containing groups and 36 proteins binding sulphate ions were analysed to characterise the structural features of anion binding sites in proteins. 34 of the 66 phosphates found were in close proximity to the amino terminus of an alpha-helix. 27% of phosphate groups bind to only one amino acid, but there is a wide distribution, with 3% of phosphates binding to seven residues. Similarly, there is a large variability in the number of contacts each phosphate group makes to the protein. This ranges from none (3% of phosphates) to nine (3% of phosphates). The most common number of contacts is two (23% of phosphates). The most commonly found residue at helix-type binding sites is glycine, followed by Arg, Thr, Ser and Lys. At non-helix binding sites, the most commonly found residue is Arg followed by Tyr, His, Lys and Ser. There is no typical phosphate binding site. There are marked differences between propensities for phosphate binding at helix and non-helix type binding sites. Non-helix binding sites show more discrimination between the types of residues involved in binding when compared to the helix set. The propensities for binding of the amino acids reveal the expected trend of positively charged and polar residues being good at binding (although that for lysine is unexpectedly low) with the bulky non-polar residues being poor at binding. Bulky residues are less likely to bind with the amide nitrogen. Sulphate binding sites show similar trends. Analysis of multiple sequence alignments that include phosphate and sulphate binding proteins reveals the degree of conservation at the binding site residues compared to the average conservation of residues in the protein. Phosphate binding site residues are more conserved than sulphate binding sites.

Crystallography, X-Ray↗

Cloning and mRNA expression of human unconventional myosin-IC. A homologue of amoeboid myosins-I with a single IQ motif and an SH3 domain.

The complete deduced amino acid sequence and mRNA expression of human unconventional myosin-IC (HuncM-IC) are described. Sequencing of overlapping cDNA clones reveals a message of 4666 nucleotides with a single open reading frame predicted to encode a 127 kDa protein of 1109 amino acids. HuncM-IC is composed of three discrete regions: a characteristic N-terminal myosin head with predicted actin and ATP-binding sites; a neck domain with an "IQ motif", predicted to bind a single light chain; and a C-terminal tail with a putative membrane-binding site. In addition, the tail contains an src-homology 3 domain. The presence of a single IQ motif and an src-homology 3 domain is reminiscent of "long-tailed" myosins-I from amoeboid organisms, a supposition confirmed by multiple sequence alignment. Northern blot analysis of human tissues shows that HuncM-IC is ubiquitously expressed, with the highest levels in kidney, prostate, colon, liver and ovary. The results show that "amoeboid" myosins-I are not restricted to amoeboid organisms, rather they are expressed in the metazoa as well.

Acanthamoeba↗

Structure-guided analysis reveals nine sequence motifs conserved among DNA amino-methyltransferases, and suggests a catalytic mechanism for these enzymes.

Previous X-ray crystallographic studies have revealed that the catalytic domain of a DNA methyltransferase (Mtase) generating C5-methylcytosine bears a striking structural similarity to that of a Mtase generating N6-methyladenine. Guided by this common structure, we performed a multiple sequence alignment of 42 amino-Mtases (N6-adenine and N4-cytosine). This comparison revealed nine conserved motifs, corresponding to the motifs I to VIII and X previously defined in C5-cytosine Mtases. The amino and C5-cytosine Mtases thus appear to be more closely related than has been appreciated. The amino Mtases could be divided into three groups, based on the sequential order of motifs, and this variation in order may explain why only two motifs were previously recognized in the amino Mtases. The Mtases grouped in this way show several other group-specific properties, including differences in amino acid sequence, molecular mass and DNA sequence specificity. Surprisingly, the N4-cytosine and N6-adenine Mtases do not form separate groups. These results have implications for the catalytic mechanisms, evolution and diversification of this family of enzymes. Furthermore, a comparative analysis of the S-adenosyl-L-methionine and adenine/cytosine binding pockets suggests that, structurally and functionally, they are remarkably similar to one another.

Amino Acid Sequence↗

Thermodynamic prediction of conserved secondary structure: application to the RRE element of HIV, the tRNA-like element of CMV and the mRNA of prion protein.

An algorithm for prediction of conserved secondary structure of single-stranded RNA is presented. For each RNA of a set of homologous RNAs optimal and suboptimal secondary structures are calculated and stored in a base-pair probability matrix. A multiple sequence alignment is performed for the set of RNAs. The resulting gaps are introduced into the individual probability matrices. These homologous probability matrices are summed to give a consensus probability matrix emphasizing the conserved secondary structure elements of the RNA set. Thus the algorithm combines the advantages of thermodynamic structure prediction by energy minimization with the information obtained from phylogenetic alignment of sequences. The algorithm is applied to three examples. The REV-responsive element of HIV, the structure of which is well known from the literature, was chosen to test the algorithm. The second example is the 3' terminal segment of genomic single-stranded RNAs of cucumber mosaic viruses; a structure similar to that of the related brome mosaic virus was expected and was confirmed. The third example is the prion-protein mRNA from different organisms; the structure of this mRNA is not known. By application of the algorithm highly conserved hairpins were found in the prion-protein mRNA.

Algorithms↗