PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Molecular cloning of the cDNA for the catalytic subunit of human DNA polymerase delta.

The cDNA of human DNA polymerase delta was cloned. The cDNA had a length of 3.5 kb and encoded a protein of 1107 amino acid residues with a calculated molecular mass of 124 kDa. Northern blot analysis showed that the cDNA hybridized to a mRNA of 3.4 kb. Monoclonal and polyclonal antibodies to the C-terminal 20 residues specifically immunoblotted the human pol delta catalytic polypeptide. A multiple sequence alignment was constructed. This showed that human pol delta is closely related to yeast pol delta and the herpes virus DNA polymerases. The levels of pol delta message were found to be induced concomitantly with DNA pol delta activity and DNA synthesis in serum restimulated proliferating IMR90 cultured cells. The human pol delta gene was localized to chromosome 19 by Southern blotting of EcoRI digested DNA from a panel of rodent/human cell hybrids.

Amino Acid Sequence

Homology modelling and protein engineering strategy of subtilases, the family of subtilisin-like serine proteinases.

Subtilases are members of the family of subtilisin-like serine proteases. Presently, greater than 50 subtilases are known, greater than 40 of which with their complete amino acid sequences. We have compared these sequences and the available three-dimensional structures (subtilisin BPN', subtilisin Carlsberg, thermitase and proteinase K). The mature enzymes contain up to 1775 residues, with N-terminal catalytic domains ranging from 268 to 511 residues, and signal and/or activation-peptides ranging from 27 to 280 residues. Several members contain C-terminal extensions, relative to the subtilisins, which display additional properties such as sequence repeats, processing sites and membrane anchor segments. Multiple sequence alignment of the N-terminal catalytic domains allows the definition of two main classes of subtilases. A structurally conserved framework of 191 core residues has been defined from a comparison of the four known three-dimensional structures. Eighteen of these core residues are highly conserved, nine of which are glycines. While the alpha-helix and beta-sheet secondary structure elements show considerable sequence homology, this is less so for peptide loops that connect the core secondary structure elements. These loops can vary in length by greater than 150 residues. While the core three-dimensional structure is conserved, insertions and deletions are preferentially confined to surface loops. From the known three-dimensional structures various predictions are made for the other subtilases concerning essential conserved residues, allowable amino acid substitutions, disulphide bonds, Ca(2+)-binding sites, substrate-binding site residues, ionic and aromatic interactions, proteolytically susceptible surface loops, etc. These predictions form a basis for protein engineering of members of the subtilase family, for which no three-dimensional structure is known.

Amino Acid Sequence

Structural and evolutionary relationships among the immunophilins: two ubiquitous families of peptidyl-prolyl cis-trans isomerases.

The immunophilins, protein receptors for the immunosuppressing drugs cyclosporin A and FK506 and related proteins from plants, fungi, and bacteria, have been analyzed structurally and evolutionarily. The cyclosporin A binding proteins (cyclophilins) represent one ubiquitous family of homologous proteins, and the FK506- and rapamycin-binding proteins (FKBPs) constitute a second, unrelated family. Multiple sequence alignments of members of each of these two protein families define the highly conserved residues that are likely to play important structural and functional roles, and mutations in representative members of these two families that abolish or alter function have been evaluated. FKBPs have undergone greater evolutionary divergence than the cyclophilins. Evolutionary trees were constructed using two distinct programs, and these trees establish the structural relationships that allow division of each of these families into subgroups. The results lead to the suggestion that several genes encoding isozymic forms of the FKBPs and possibly also of the cyclophilins existed in prokaryotes before the emergence of eukaryotes on earth and that representatives of these genes were transmitted to both kingdoms to give rise to current subfamilies of these proteins. By contrast, compartmentalization of both classes of immunophilins appears to have arisen independently in prokaryotes and eukaryotes, late in evolutionary history.

Amino Acid Isomerases

Genomic divergence of an HIV-2 from a German AIDS patient probably infected in Mali.

The complete nucleotide sequence of an HIV-2 isolate derived from a German AIDS patient with predominantly neurological symptoms is reported. The HIV-2BEN sequence is highly divergent from those of previously described HIV-2 and SIV strains. Evolutionary tree analysis of eight HIV-2 sequences reveals the existence of three HIV-2 groups. HIV-2BEN belongs to a group with two isolates from Ghana and The Gambia. Based on a comparison of HIV-2BEN with six HIV-2 isolates, SIVsmm and SIVmac, the variability of the structural env and gag proteins is similar within the HIV-2/SIVsmm/mac and HIV-1 groups. In contrast, the regulatory HIV-1 proteins are more highly conserved than those from HIV-2 strains. Multiple sequence alignments reveal that some domains of the envelope and regulatory proteins are well conserved among HIV-1, HIV-2/SIVsmm/mac, SIVagm and SIVmnd. The identification of conserved domains within the external glycoprotein could help to develop broadly active vaccines.

Acquired Immunodeficiency Syndrome

Site-directed mutagenesis of the lipoate acetyltransferase of Escherichia coli.

Remote but significant similarities between the primary and predicted secondary structures of the chloramphenicol acetyltransferases (CAT) and lipoate acyltransferase subunits (LAT, E2) of the 2-oxo acid dehydrogenase complexes, have suggested that both types of enzyme may use similar catalytic mechanisms. Multiple sequence alignments for CAT and LAT have highlighted two conserved motifs that contain the active-site histidine and serine residues of CAT. Site-directed replacement of Ser550 in the E2p subunit (LAT) of the pyruvate dehydrogenase complex of Escherichia coli, deemed to be equivalent to the active-site Ser148 of CAT, supported the CAT-based model of LAT catalysis. The effects of other substitutions were also consistent with the predicted similarity in catalytic mechanism although specific details of active-site geometry may not be conserved.

Acetyltransferases

Identification of a gag protein epitope conserved among all four groups of primate immunodeficiency viruses by using monoclonal antibodies.

Five monoclonal antibodies (MAbs) were raised against the gag proteins of simian immunodeficiency virus (SIV) from African green monkey (SIVagmTYO-7). Two MAbs reacted with the matrix protein p17 and the other three with the core protein p24. Studies on the cross-reactivity of the MAbs revealed that the anti-p24 MAbs detected an epitope shared by the viruses belonging to the human immunodeficiency virus type 2 (HIV-2)/SIVmac group and SIVagmTYO-7 and SIVagmTYO-5. The anti-p17 MAbs recognized an epitope present on all these viruses and on SIVagmTYO-1, HIV-1 and SIVmnd. This finding demonstrates for the first time that the matrix protein, p17 or p18, respectively, of all nine HIV and SIV isolates tested in this study expresses at least one conserved immunogenic epitope recognized serologically. By using synthetic peptides, this epitope was identified at the N terminus of p17. Furthermore, this epitope was analysed by multiple sequence alignments of the peptide with homologous sequences of HIV and SIV p17.

Amino Acid Sequence

Citrate synthase from the thermophilic archaebacterium Thermoplasma acidophilium. Cloning and sequencing of the gene.

The gene encoding the citric acid cycle enzyme, citrate synthase, has been cloned from the thermoacidophilic archaebacterium, Thermoplasma acidophilum. We report the sequencing of this gene and its flanking regions, and the derived amino acid sequence of the enzyme is compared by multiple-sequence alignment analysis with those of citrate synthases from eubacterial and eukaryotic organisms. The similarity is less than 30% between the archaebacterial and non-archaebacterial sequences, although the majority of residues implicated in the catalytic action of the enzyme have been conserved across all three kingdoms. The cloned archaebacterial gene has been expressed in Escherichia coli to produce catalytically active citrate synthase. This is the first reported sequence of citrate synthase from the archaebacteria.

Amino Acid Sequence

Insecticidal properties of a crystal protein gene product isolated from Bacillus thuringiensis subsp. kenyae.

A protoxin gene, localized to a high-molecular-weight plasmid from Bacillus thuringiensis subsp. kenyae, was cloned on a 19-kb BamHI DNA fragment into Escherichia coli. Characterization of the gene revealed it to be a member of the CryIE toxin subclass which has been reported to be as toxic as the CryIC subclass to larvae from Spodoptera exigua in assays with crude E. coli extracts. To directly test the purified recombinant gene product, the gene was subcloned as a 4.8-kb fragment into an expression vector resulting in the overexpression of a 134-kDa protein in the form of phase-bright inclusions in E. coli. Treatment of solubilized inclusion bodies with either trypsin or gut juice from the silkworm Bombyx mori resulted in the appearance of a protease-resistant 65-kDa protein. In force-feeding bioassays, the purified activated protein was highly toxic to larvae of B. mori but not to larvae of Choristoneura fumiferana. In diet bioassays with larvae from S. exigua, the purified protoxin was nontoxic. However, prior activation of the protoxin by tryptic digestion resulted in the appearance of some toxic activity. These results demonstrate that this new subclass of protein toxin may not be useful for the control of Spodoptera species as previously reported. Hierarchical clustering of the nine known lepidopteran-specific CryI toxin subclasses through multiple sequence alignment suggests that the toxins fall into four possible subgroups or clusters.

Animals

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking", Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana (Witch Hazel), contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a helix swap; the other has no disulfide bonds and possesses two tandem domains. To explore the structural evolution of bicycle proteins, we predicted bicycle protein structures with Alphafold2 (AF2). While AF2 did not recover the two experimental structures using existing databases, it succeeded after we provided multiple sequence alignments (MSAs) containing protein sequences encoded in new genome sequences from closely related aphid species. Using this customized approach at scale, we generated 2400 high-confidence predictions for bicycle proteins from seven aphid species. This dataset revealed that bicycle proteins without cysteines are outliers in fold space and appear to have evolved from ancestral proteins with disulfide-bonded saposin-like folds. While all bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance.

AlphaFold predictions

The VPH1 gene encodes a 95-kDa integral membrane polypeptide required for in vivo assembly and activity of the yeast vacuolar H(+)-ATPase.

Yeast vacuolar acidification-defective (vph) mutants were identified using the pH-sensitive fluorescence of 6-carboxyfluorescein diacetate (Preston, R. A., Murphy, R. F., and Jones, E. W. (1989) Proc. Natl. Acad. Sci. U.S.A. 86, 7027-7031). Vacuoles purified from yeast bearing the vph1-1 mutation had no detectable bafilomycin-sensitive ATPase activity or ATP-dependent proton pumping. The peripherally bound nucleotide-binding subunits of the vacuolar H(+)-ATPase (60 and 69 kDa) were no longer associated with vacuolar membranes yet were present in wild type levels in yeast whole cell extracts. The VPH1 gene was cloned by complementation of the vph1-1 mutation and independently cloned by screening a lambda gt11 expression library with antibodies directed against a 95-kDa vacuolar integral membrane protein. Deletion disruption of the VPH1 gene revealed that the VPH1 gene is not essential for viability but is required for vacuolar H(+)-ATPase assembly and vacuolar acidification. VPH1 encodes a predicted polypeptide of 840 amino acid residues (molecular mass 95.6 kDa) and contains six putative membrane-spanning regions. Cell fractionation and immunodetection demonstrate that Vph1p is a vacuolar integral membrane protein that co-purifies with vacuolar H(+)-ATPase activity. Multiple sequence alignments show extensive homology over the entire lengths of the following four polypeptides: Vph1p, the 116-kDa polypeptide of the rat clathrin-coated vesicles/synaptic vesicle proton pump, the predicted polypeptide encoded by the yeast gene STV1 (Similar To VPH1, identified as an open reading frame next to the BUB2 gene), and the TJ6 mouse immune suppressor factor.

Amino Acid Sequence

Computational sequence analysis revisited: new databases, software tools, and the research opportunities they engender.

The increasing quantity and complexity of sequences and structural data for proteins and nucleic acids create both problems and opportunities for biomedical researchers. Fortunately, a new generation of practical computer tools for data analysis and integrated information retrieval is emerging. Recent developments in fast database searching, multiple sequence alignment, and molecular modeling are discussed and windows-based, mouse-driven software for CD-ROM and network information retrieval are described. Each method is illustrated with a practical example pertinent to lipid research. In particular, the connection among cholesteryl ester transfer protein, bactericidal permeability-increasing protein, and lipopolysaccharide-binding proteins is determined; novel repetitive sequence motifs in mammalian farnesyltransferase subunits and related yeast prenyltransferases are derived; biochemical insights from a three-dimensional model of human apolipoprotein D based on two insect lipocalins are discussed; the relationship between apolipoprotein D and gross cystic disease fluid protein from human breast is reviewed; and prospects for modeling apolipoprotein E-related proteins are described. In addition, information on a number of general and special-purpose sequence, motif, and structural databases is included.

Amino Acid Sequence

Multiple DNA and protein sequence alignment on a workstation and a supercomputer.

This paper describes a multiple alignment method using a workstation and supercomputer. The method is based on the alignment of a set of aligned sequences with the new sequence, and uses a recursive procedure of such alignment. The alignment is executed in a reasonable computation time on diverse levels from a workstation to a supercomputer, from the viewpoint of alignment results and computational speed by parallel processing. The application of the algorithm is illustrated by several examples of multiple alignment of 12 amino acid and DNA sequences of HIV (human immunodeficiency virus) env genes. Colour graphic programs on a workstation and parallel processing on a supercomputer are discussed.

Algorithms

The orotidine-5'-monophosphate decarboxylase gene of Myxococcus xanthus. Comparison to the OMP decarboxylase gene family.

The nucleotide sequence of the Myxococcus xanthus orotidine-5'-monophosphate decarboxylase (OMP DCase) gene was determined. The derived protein sequence is not closely related to other prokaryotic OMP DCase sequences; nor is it closely related to any eukaryotic OMP DCase sequences. Progressive multiple alignment of the M. xanthus OMP DCase protein sequence with 19 other OMP DCase sequences revealed four conserved regions present in all 20 sequences. Ten entirely conserved residues were found in these four regions and one region contains a tight cluster of 5 conserved residues, certain of which may be catalytically active residues. A second open reading frame was found upstream of uraA and oriented in the same direction as uraA. A stretch of 21 consecutive pyrimidine (C or T) residues were found in the intercistronic region between the potential ribosome-binding site of uraA and the UGA stop codon of the upstream open reading frame. RNA directly upstream of the pyrimidine run, including the UGA stop codon of the upstream open reading frame, could be folded into a stable hairpin structure resembling Rho-independent terminators of Escherichia coli. Expression of the uraA gene may be regulated by an intercistronic transcription termination mechanism.

Amino Acid Sequence

Evolutionary divergence plots of homologous proteins.

A simple and efficient method is described for analyzing quantitatively multiple protein sequence alignments and finding the most conserved blocks as well as the maxima of divergence within the set of aligned sequences. It consists of calculating the mean distance and the root-mean-square distance in each column of the multiple alignment, averaging the values in a window of defined length and plotting the results as a function of the position of the window. Due attention is paid to the presence of gaps in the columns. Several examples are provided, using the sequences of several cytochromes c, serine proteases, lysozymes and globins. Two distance matrices are compared, namely the matrix derived by Gribskov and Burgess from the Dayhoff matrix, and the Risler Structural Superposition Matrix. In each case, the divergence plots effectively point to the specific residues which are known to be essential for the catalytic activity of the proteins. In addition, the regions of maximum divergence are clearly delineated. Interestingly, they are generally observed in positions immediately flanking the most conserved blocks. The method should therefore be useful for delineating the peptide segments which will be good candidates for site-directed mutagenesis and for visualizing the evolutionary constraints along homologous polypeptide chains.

Amino Acid Sequence

Phylogenetic relationships among megabats, microbats, and primates.

We present 744 nucleotide base positions from the mitochondrial 12S rRNA gene and 236 base positions from the mitochondrial cytochrome oxidase subunit I gene for a microbat, Brachyphylla cavernarum, and a megabat, Pteropus capestratus, in phylogenetic analyses with homologous DNA sequences from Homo sapiens, Mus musculus (house mouse), and Gallus gallus (chicken). We use information on evolutionary rate differences for different types of sequence change to establish phylogenetic character weights, and we consider alternative rRNA alignment strategies in finding that this mtDNA data set clearly supports bat monophyly. This result is found despite variations in outgroup used, gap coding scheme, and order of input for DNA sequences in multiple alignment bouts. These findings are congruent with morphological characters including details of wing structure as well as cladistic analyses of amino acid sequences for three globin genes and indicate that neurological similarities between megabats and primates are due to either retention of primitive characters or to convergent evolution rather than to inheritance from a common ancestor. This finding also indicates a single origin for flight among mammals.

Amino Acid Sequence

Calculating percent identity between protein or DNA sequences with a word processor.

Two macros, to calculate percentage identity between protein or DNA sequences using the Microsoft Word word processor, are described. The user prepares an alignment file of multiple sequences which is used by the macros to calculate number of matches, number of mismatches, total number of compared positions, and the percent identity. The macros are especially useful when alignment of multiple sequences is possible only by eye.

Algorithms

Evolutionary relationship between the TonB-dependent outer membrane transport proteins: nucleotide and amino acid sequences of the Escherichia coli colicin I receptor gene.

The nucleotide sequence of the Escherichia coli colicin I receptor gene (cir) has been determined. The predicted mature protein consists of 599 amino acids and has a molecular weight of 67,169. Several previously noted characteristics of other E. coli outer membrane protein sequences were also identified in the sequence of Cir. These include an overall acidic nature, the absence of long hydrophobic stretches of amino acids, and a lack of predicted alpha-helical secondary structure. Because two classes of outer membrane proteins (the TonB-dependent transport proteins and the porins) share some structural features, protein sequences from both of these groups were aligned pairwise and scored for sequence similarity. Statistical evidence suggested that the porins were not related to the proteins in the TonB-dependent group; however, there was a significant relationship between the proteins in the TonB-dependent group. On the basis of the multiple progressive sequence alignment and the similarity scores derived from it, a tree representing evolutionary distance between five TonB-dependent outer membrane transport proteins was generated.

Amino Acid Sequence