PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Pattern recognition and self-correcting distance geometry calculations applied to myohemerythrin.

A topological list, consisting of segments of regular secondary structures and a list of buried and solvent accessible residues, is automatically predicted from multiple aligned sequences in a protein family. This topological list is translated into geometric constraints for distance geometry calculation in torsion angle space. A new self-correcting distance geometry method detects and eliminates false distance constraints. In an application to the four-helix bundle protein, myohem-erythrin, the right-handed global fold was correctly reproduced with a root-mean-square deviation of 2.6 A, when the topological list was derived from the X-ray structure. A predicted topological list, coupled with constraints from the residues in the active site of myohemerythrin, predicted the correct fold with a root-mean-square deviation of 4 A for backbone atoms.

Amino Acid Sequence↗

Specificity of the cytochrome P-450 interaction with cytochrome b5.

The specificity of the interaction of cytochrome b5 with different forms of cytochrome P-450 was examined. Immunopurification of cytochromes P-450 1A1, 2B1 and 2E1 from rat liver microsomes resulted in co-purification of cytochrome b5 with cytochrome P-450 forms 2B1 and 2E1 but not 1A1. This specificity was evaluated in conjunction with multiple sequence alignment of the three cytochrome P-450s and a molecular model of the cytochrome P-450-cytochrome b5 complex [(1989) Biochemistry 28, 8201-8205]. These analyses suggest two basic residues in the arginine cluster region of P-450, which are present in P-450s 2B1 and 2E1 but are absent in P-450 1A1, as potential binding sites for cytochrome b5.

Amino Acid Sequence↗

Hydrophobic cluster analysis and secondary structure predictions revealed that major and minor structural subunits of K88-related adhesins of Escherichia coli share a common overall fold and differ structurally from other fimbrial subunits.

The structural relatedness of K88-related major and minor subunits was deduced from their sequences by hydrophobic cluster analysis (HCA) and secondary structure predictions produced by the profile neural network prediction program (PHD) on multiple sequence alignments. Although the weak residue identity between major and minor subunits is evidence of a high evolutionary distance, an overall structural similarity was observed In addition, clear amphipathic conformations were conserved in predicted secondary structure. On the basis of this predicted structural similarity, a schematic 2D model of ClpG subunit was developed.

Amino Acid Sequence↗

Identification of additional homologues of subunits VII and VIII of the ubiquinol-cytochrome c oxidoreductase enables definition of consensus sequences.

The Candida utilis QCR7 gene encoding subunit VII of the ubiquinol-cytochrome c oxidoreductase was isolated by functional complementation of the Saccharomyces cerevisiae subunit VII-null mutant. Several other subunit VII homologues as well as homologues for subunit VIII were identified by screening the GenBank database. Some of these homologues for subunit VII could only be identified as such using a consensus sequence that was derived from the multiple sequence alignment. Definition of the consensus should facilitate further analysis of structure/function relationships in this protein.

Amino Acid Sequence↗

Ribosomal protein L22 from Thermus thermophilus: sequencing, overexpression and crystallisation.

The gene for the ribosomal protein L22 from Thermus thermophilus has been sequenced and overexpressed in Escherichia coli. A multiple sequence alignment was carried out for all proteins of the L22 family reported so far. The recombinant protein was purified and crystallized. The crystals belong to the space group P2(1)2(1)2(1), with cell parameters of a = 32.6 A, b = 66.0 A, c = 67.8 A.

Amino Acid Sequence↗

Spectroscopic study of an HIV-1 capsid protein major homology region peptide analog.

The capsid (CA) domain of retroviral Gag proteins possesses one subdomain, the major homology region (MHR), which is conserved among nearly all avian and mammalian retroviruses. While it is known that the mutagenesis of residues in the MHR will impair virus infectivity, the precise structure and function of the MHR is not known. In order to obtain further information on the MHR, we have examined the structure of a synthetic peptide encompassing the MHR of human immunodeficiency virus type I (HIV-1) CA protein. Multiple sequence alignment and secondary structure prediction indicate that the peptide could form 50% alpha-helix and 10% beta-sheet. In addition, circular dichroism studies indicate that, in the presence of 50% trifluoroethanol (TFE), the peptide adopts an alpha-helical structure over half of its length. Further analysis by proton nuclear magnetic resonance spectroscopy suggests that the C-terminal portion of the MHR forms a helix in aqueous solution. Upon the addition of TFE, the position of the helix remains nearly constant, but the magnitude of the changes in H alpha chemical shifts of the residues indicate a more stable helix. These results suggest that a helical C-terminus of retroviral MHRs may be integral to the function of this region.

Amino Acid Sequence↗

Cloning and sequencing of cDNA clones encoding chicken lamins A and B1 and comparison of the primary structures of vertebrate A- and B-type lamins.

Nuclear lamins are intermediate-filament-type proteins forming a fibrillar meshwork underlying the inner nuclear membrane. The existence of multiple isoforms of lamin proteins in vertebrates is believed to reflect functional specializations during cell division and differentiation. Although biochemical criteria may be used to classify many lamin isoforms into A- and B-type subfamilies, the structural features distinguishing the members of these subfamilies remain to be characterized fully. Here, we report the complete primary structures of chicken lamins A and B1, as they are deduced from cloned cDNAs; in the accompanying paper we present the complete sequence of lamin B2, a second avian B-type lamin. Comparisons of the chicken lamin sequences with each other and with those of other lamins allow us to establish structural features that are common to members of both subfamilies. Conversely, multiple sequence alignments make it possible to identify a number of structural motifs that clearly differentiate B-type lamins from A-type lamins. With this information at hand, we attempt to correlate different biochemical properties of A- and B-type lamins with the presence or absence of specific sequence motifs.

Amino Acid Sequence↗

Analysis of insertions/deletions in protein structures.

An analysis of insertions and deletions (indels) occurring in a databank of multiple sequence alignments based on protein tertiary structure is reported. Indels prefer to be short (1 to 5 residues). The average intervening sequence length between them versus the percentage of residue identity in pairwise alignments shows an exponential behaviour, suggesting a stochastic process such that nearly every loop in an ancestral structure is a possible target for indels during evolution. The results also suggest a limit to the average size of indels accommodated by protein structures. The preferred indel conformations are reverse turn and coil as are the preferred conformations at the indel edges (N- and C-terminal sides). Interruptions in helices and strands were observed as very rare events.

Amino Acid Sequence↗

Position-based sequence weights.

Sequence weighting methods have been used to reduce redundancy and emphasize diversity in multiple sequence alignment and searching applications. Each of these methods is based on a notion of distance between a sequence and an ancestral or generalized sequence. We describe a different approach, which bases weights on the diversity observed at each position in the alignment, rather than on a sequence distance measure. These position-based weights make minimal assumptions, are simple to compute, and perform well in comprehensive evaluations.

Amino Acid Sequence↗

Sequence analysis of steroid- and prostaglandin-metabolizing enzymes: application to understanding catalysis.

Amino acid sequence comparisons have revealed that mammalian 11 beta-hydroxysteroid and 17 beta-hydroxysteroid dehydrogenases and bacterial 3 alpha, 20 beta- and 3 beta-hydroxysteroid dehydrogenases are homologs; that is, these enzymes are descended from a common ancestor. These steroid dehydrogenases are also homologous to human 15-hydroxyprostaglandin dehydrogenase and to proteins found in Rhizobia, bacteria that form nitrogen-fixing nodules in the roots of legumes. We constructed a multiple sequence alignment of these proteins, which, when combined with the recently determined tertiary structure of Streptomyces hydrogenans 3 alpha, 20 beta-hydroxysteroid dehydrogenase and a homologous enzyme, rat dihydropteridine reductase, identifies segments and residues that are likely to be structurally important in the functioning of these enzymes especially regarding specificity for NADPH and NADH.

Amino Acid Sequence↗

Toward the unification of sequence and structural data for identification of structural and functional constraints.

The identification and characterization of local residue patterns or conserved segments shared by a set of biopolymers has provided a number of insights in molecular biology. Biopolymer sequences are observations from macro molecules that share common structural or function features. The approach taken here rests on the notion that information may be most efficiently extracted from these observations through the use of a model that faithfully represents macro-molecular characteristics. Accordingly, our efforts are focused on statistical models which attempt to capture central features of protein structure, function, and change. Here the assumptions that underlie two new methods for the analysis of protein sequence data are explicitly delineated. (1) Threading of a sequence through structural motifs seeks to determine if a protein sequence fits a known protein structure. The assumptions delineated here also generally apply to other contact based threading methods that have been recently described. (2) Multiple sequence alignment via the Gibbs sampling algorithm seeks to identify position specific empirical free energy models for residue sites in common motifs and simultaneously the align sequence observations form these motifs.

Algorithms↗

Structural analysis of homologous repeated domains in alpha-actinin and spectrin.

The amino acid sequences of chick and slime mould alpha-actinin each contain four repeats of approximately 122 residues. These repeats are homologous to the 18-22 repeats, each of approximately 106 residues, found in the alpha and beta subunits of spectrin and fodrin, and to the multiple repeats of approximately 110 residues found in the Duchenne muscular dystrophy protein (dystrophin). The repeats correspond to the elongated rod-like portion of these molecules. We present a multiple sequence alignment of 21 repeats from this superfamily (8 alpha-actinin and 13 spectrin/fodrin), based on optimal pairwise alignments, from which a characteristic consensus pattern of amino acid types is deduced. Trp 46 is invariant in all but one repeat, and physicochemical classes of amino acids are conserved at 25 other positions. Secondary structure prediction on both the alpha-actinin and spectrin repeats taken together with the distribution of proline residues in the sequences, strongly suggest that each repeated domain consists of a four-helix structure. Our predictions differ significantly from previous three-helix models based on analyses of fewer sequences. To determine possible interdomain regions, sites of limited proteolysis of the native chick alpha-actinin dimer were determined and located in the amino acid sequence. The majority of these sites were in corresponding positions in different repeats within a segment predicted as a long helix. We propose a model, consistent with the overall dimensions of the rod-like portions of the molecules, in which these long, probably interrupted helices, link adjacent domains.

Actinin↗

Comparative molecular modelling of the Fas-ligand and other members of the TNF family.

A number of proteins with significant similarity to the tumour necrosis factor (TNF) have been identified over the last years. Upon interaction with their cognate receptor (members of the TNF-receptor family), all members of this protein family induce either cell death or proliferation/differentiation of the receptor-bearing cells. One of the last identified members of the TNF family is the apoptosis-inducing ligand of the Fas-receptor, termed Fas-ligand (FasL). Here we report the cloning and sequencing of the mouse cDNA for the FasL. Using knowledge-based protein modelling, we demonstrate that all members of the TNF family form trimeric complexes, and define the residues located at the subunit interfaces. The resulting structurally corrected multiple sequence alignment allows the identification of residues potentially involved in receptor recognition, and should help design mutagenesis experiments for structure-function relationship studies.

Amino Acid Sequence↗

Identification of Zucchini yellow mosaic potyvirus by RT-PCR and analysis of sequence variability.

A reverse transcription-polymerase chain reaction (RT-PCR) method was used to identify Zucchini yellow mosaic virus (ZYMV) in leaves of infected cucurbits. Oligonucleotide primers which annealed to regions in the nuclear inclusion body (NIb) and the coat protein (CP) genes, generated a 300-bp product from ZYMV and also from the closely related watermelon mosaic virus type 2 (WMV-2). However, no product was obtained from papaya ringspot potyvirus which also infects cucurbits. ZYMV and WMV-2 were differentiated using a third primer which was complementary to a sequence in the 3'-untranslated region; a 1186-bp amplified product was obtained for ZYMV only. Nucleotide sequence analysis of the 300-bp fragments of Australian ZYMV and WMV-2 strains revealed 93.7-100% sequence identity between ZYMV strains. Multiple sequence alignments indicated that the nucleotide sequence which codes for the N-terminus of the CP was 74-100% identical for different isolates of ZYMV. The Australian isolate of WMV-2 was 43-46% identical to all isolates of ZYMV and was 84.6% identical to a Florida isolate of WMV-2.

Amino Acid Sequence↗

VISTAS: a package for VIsualizing STructures and sequences of proteins.

VISTAS is a suite of programs for protein sequence and structure analysis. The system allows the simultaneous display, in separate windows, of multiple sequence alignments, of known or model 3D structures, and of 2D graphic representations of sequence and/or alignment properties. The displays are fully integrated, and therefore manipulations in one window can be reflected in each of the others. Beyond its display facilities, VISTAS brings together a number of existing tools under a single, user-friendly umbrella: these include a fully functional interactive color alignment procedure, conserved motif selection, a range of database-scanning routines, and interactive access to the OWL composite sequence database and to the PRINTS protein fingerprint database. Exploration of the sequence database is thus straightforward, and predefined structural motifs from the fingerprint database may be readily visualized. Of particular note is the ability to calculate conservation criteria from sequence alignments and to display the information in a 3D context: this renders VISTAS a powerful tool for aiding mutagenesis studies and for facilitating refinement of molecular models.

Amino Acid Sequence↗

A nested polymerase chain reaction for the detection of Borrelia burgdorferi sensu lato based on a multiple sequence analysis of the hbb gene.

A highly sensitive nested polymerase chain reaction method was designed for the detection of a wide spectrum of strains from Borrelia burgdorferi sensu lato. This technique allows the detection of as little as 3 fg of total genomic DNA extracted and purified from pure cultures of the organism, this amount corresponds to less than 10 organisms. Two sets of primers homologous to conserved spots in the coding region of the hbb gene, encoding a conserved histone-like protein, were constructed. These were based on a multiple sequence alignment of 39 strains representing all the genomic groups described in B. burgdorferi sensu lato.

Base Sequence↗

Genomic sequences of murine gamma B- and gamma C-crystallin-encoding genes: promoter analysis and complete evolutionary pattern of mouse, rat and human gamma-crystallins.

The murine genes, gamma B-cry and gamma C-cry, encoding the gamma B- and gamma C-crystallins, were isolated from a genomic DNA library. The complete nucleotide (nt) sequences of both genes were determined from 661 and 711 bp, respectively, upstream from the first exon to the corresponding polyadenylation sites, comprising more than 2650 and 2890 bp, respectively. The new sequences were compared to the partial cDNA sequences available for the murine gamma B-cry and gamma C-cry, as well as to the corresponding genomic sequences from rat and man, at both the nt and predicted amino acid (aa) sequence levels. In the gamma B-cry promoter region, a canonical CCAAT-box, a TATA-box, putative NF-I and C/EBP sites were detected. An R-repeat is inserted 366 bp upstream from the transcription start point. In contrast, the gamma C-cry promoter does not contain a CCAAT-box, but some other putative binding sites for transcription factors (AP-2, UBP-1, LBP-1) were located by computer analysis. The promoter regions of all six gamma-cry from mouse, rat and human, except human psi gamma F-cry, were analyzed for common sequence elements. A complex sequence element of about 70-80 bp was found in the proximal promoter, which contains a gamma-cry-specific and almost invariant sequence (crygpel) of 14 nt, and ends with the also invariant TATA-box. Within the complex sequence element, a minimum of three further features specific for the gamma A-, gamma B- and gamma D/E/F-cry genes can be defined, at least two of which were recently shown to be functional. In addition to these four sequence elements, a subtype-specific structure of inverted repeats with different-sized spacers can be deduced from the multiple sequence alignment. A phylogenetic analysis based on the promoter region, as well as the complete exon 3 of all gamma-cry from mouse, rat and man, suggests separation of only five gamma-cry subtypes (gamma A-, gamma B-, gamma C-, gamma D- and gamma E/F-cry) prior to species separation.

Amino Acid Sequence↗

Primary structure and biological features of a thermostable nuclease isolated from Staphylococcus hyicus.

The nucH gene, encoding a thermostable nuclease (TNase), was isolated from the cellular DNA of Staphylococcus hyicus strain E80 and sequenced. NucH, the 169-amino-acid (aa) protein encoded by this gene, contains, at its N-terminus, a signal peptide which appears to be cleaved at the same site in S. hyicus and Escherichia coli, yielding a mature protein which is exported extracellularly from S. hyicus, but not from E. coli. The aa sequence of NucH is highly homologous with that of the TNase from S. intermedius strain LRA076, whereas significant similarities are observed with the TNase from S. aureus, as well as with three other bacterial proteins of which only one has been shown to exhibit DNase activity. As seen in a multiple sequence alignment, the invariant residues are mostly located in the regions involved in the biological activity of the S. aureus TNase. The ability of crude cell extracts of E. coli strains carrying nucH to degrade various forms of nucleic acids with or without Ca2+ supplementation was studied. Under our experimental conditions, the enzyme encoded by nucH was active at 37 degrees C on both DNA and RNA, had the potential to act as an endonuclease, and functioned in the presence of Ca2+. Moreover, activity was retained after heating at 100 degrees C, suggesting that the enzyme could undergo reversible unfolding.

Amino Acid Sequence↗