PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence

Sequence analysis of the gene coding for glyceraldehyde-3-phosphate dehydrogenase (gpd) of Podospora anserina: use of homologous regulatory sequences to improve transformation efficiency.

The glyceraldehyde-3-phosphate dehydrogenase (gpd) gene of Podospora anserina has been isolated from a genomic library by heterologous hybridization with the corresponding gene of Curvularia lunata. The coding region consists of 1014 nucleotides and is interrupted by a single intron. The amino-acid sequence encoded by the gpd gene shows a high degree of sequence identity with the corresponding gene products of various fungi. Multiple alignments of all fungal GPD sequences so far available resulted in the construction of a phylogenetic tree. The evolutionary relationships of the various fungi belonging to different taxa will be discussed on the basis of these data. Sequence analysis of 1.9 kbp of the 5' non-coding region revealed the presence of typical fungal promoter elements. Utilizing different parts of the 5' regulatory sequence of the Podospora gpd gene, expression vectors containing a dominant selectable marker gene (hygromycin B phosphotransferase) have been constructed for the transformation of P. anserina protoplasts. The use of these homologous gpd regulatory sequences resulted in a significant increase in transformation efficiencies compared to those obtained with vectors in which the selectable marker gene is under the control of the corresponding heterologous promoter of Aspergillus nidulans.

Amino Acid Sequence

Flexible protein sequence patterns. A sensitive method to detect weak structural similarities.

The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Amino Acid Sequence

Protein database searches for multiple alignments.

Protein database searches frequently can reveal biologically significant sequence relationships useful in understanding structure and function. Weak but meaningful sequence patterns can be obscured, however, by other similarities due only to chance. By searching a database for multiple as opposed to pairwise alignments, distant relationships are much more easily distinguished from background noise. Recent statistical results permit the power of this approach to be analyzed. Given a typical query sequence, an algorithm described here permits the current protein database to be searched for three-sequence alignments in less than 4 min. Such searches have revealed a variety of subtle relationships that pairwise search methods would be unable to detect.

Algorithms

Evaluation of the sequence template method for protein structure prediction. Discrimination of the (beta/alpha)8-barrel fold.

A multiple alignment of five (beta/alpha)8-barrel enzymes has been derived from their structure. The eight beta-strands and eight alpha-helices of the (beta/alpha)8-barrel are correctly aligned and the equivalenced residues in these regions fulfil similar structural roles. Each beta-strand has a central core of usually four residues, two residues contribute side-chains to the barrel core and the other two residues are involved in beta-strand/alpha-helix contacts. However, the fold imposes no constraints on the volumes of the residues at either a local or global level: the volume of the beta-barrel core varies between 1088 A3 in glycolate oxidase and 1571 A3 in taka-amylase. Sequence motifs derived from the multiple alignment were scanned against a database of 124 protein sequences, including 17 (beta/alpha)8-barrel enzymes. The results were evaluated in terms of the discrimination of (beta/alpha)8-barrel sequences and the quality of the alignments obtained. One motif was able to identify the top 12% of high scoring sequences as forming (beta/alpha)8-barrels with 50% accuracy and the bottom 50% of sequences as not being (beta/alpha)8-barrel proteins with 100% accuracy. However, in most instances the alignments were poor. The reasons for this are discussed with reference to the (beta/alpha)8-barrel proteins and the sequence motif method in general.

Alcohol Oxidoreductases

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

ANTHEPROT 2.0: a three-dimensional module fully coupled with protein sequence analysis methods.

ANTHEPROT is a fully interactive graphics program devoted to the analysis of the sequences and structures of proteins. This program, originally developed to facilitate the protein sequence analysis coupled with multiple alignments and predicted secondary structures of proteins, now comprises a powerful 3D module to display and handle macromolecular structures. All the methods that were previously integrated into ANTHEPROT are now directly coupled with a 3D window that provides the user all the classic features of a molecular modeling package. Indeed, it allows real-time rotation and translation of 3D structures with many kinds of models in depth-cueing mode (space filling, backbone, wire models, main chain, and ribbons), selections (atom type, residue type, segments, and chain), color-coding systems (amino acid properties, predicted or observed secondary structures, temperature B factor, and subunits), geometric calculations (Ramachandran plot, distances, and angles), and fitting molecules. Stereo views are possible as well as HPGL standard files. A module specifically devoted to the determination of 3D structures using nuclear magnetic resonance is also available. This major release of our program for IBM rs6000 workstations is available by anonymous ftp to ibcp.fr for academic institutions.

Antigens

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software

Distribution and molecular characterization of integron classes from Escherichia coli and Klebsiella pneumoniae isolates in Sulaymaniyah province of Iraq.

UNLABELLED: The environmental pollution from the misuse of antimicrobial drugs is fueling selection pressure in bacteria, thereby exacerbating the threat to global health. In Iraq, the situation is made worse by the poor implementation of the World Health Organization's Global Antimicrobial Resistance and Use Surveillance System (WHO-GLASS). Consequently, this study aimed to increase surveillance of the spread of antimicrobial resistance in Sulaymaniyah, Iraq. A total of 296 Enterobacteriaceae comprising 147 Klebsiella pneumoniae and 149 Escherichia coli were isolated from humans, poultry, and dairy farms. The isolates were screened using multiplex PCR to assess the prevalence of the clinically important integron integrase (intI) classes and antimicrobial resistance genes (ARGs) of commonly used antibiotics. Remarkably, 81.14% of the isolates carried at least 2 ARGs, 10.47% intI1, and 3.72% intI2. No intI3 was detected. A total of 663 ARGs were identified using multiplex PCR in the two Enterobacteriaceae: beta-lactamase genes were 43%, tetracycline resistance genes 25.20%, sulfonamide resistance gene 16.10%, quinolone resistance gene 10.2%, and aminoglycoside resistance genes 5.7%. K. pneumoniae harbored more integrons and ARGs than E. coli, thus posing a higher antimicrobial resistance threat in this province. This study underscores the importance of implementing more stringent WHO-GLASS and antibiotic stewardship to end the multidrug resistance crisis in Iraq. IMPORTANCE: These data are about the prevalence of integrons and resistance genes, helping to fill a significant gap in global surveillance efforts. Results can be used by global health authorities and the World Health Organization to develop national and international antimicrobial resistance (AMR) control strategies. The study is important because integrons are key genetic platforms that capture and disseminate antibiotic resistance genes among bacteria. In addition, Escherichia coli and Klebsiella spp. are among the top causes of hospital- and community-acquired infections, especially urinary tract infections, bloodstream infections, and pneumonia. Therefore, it will be riskier when these bacteria have a high rate of integrons and resistance genes because it impedes treatments during infection. Another importance of this study is that the study was carried out in Iraq. Iraq, like many low- and middle-income countries, faces challenges with unregulated antibiotic use, leading to high rates of AMR.

Escherichia coli

MATCH-BOX: a fundamentally new algorithm for the simultaneous alignment of several protein sequences.

Original algorithms for simultaneous alignment of protein sequences are presented, including sequence clustering and within- or between-groups multiple alignment. The way of matching similar regions is fundamentally new. Complete matches are formed by segments more similar than expected by random, according to a given probability limit. Any classic or user-defined score matrix can be used to express the similarity between the residues. The algorithm seeks for complete matches common to all the sequences without performing pairwise alignment and regardless of gap weighting. An automatic screening delineates all the similar regions (boxes) that may be defined for a given maximal shift between the sequences. The shift can be large enough to allow the matching of any region of a sequence with any region of another one. It can also be short and used to refine the alignment around anchor points. The algorithm provides the most likely optimal alignment and a comprehensive list of the alignment dilemma. Duality between automatism and interactivity is provided. Depending on the problem complexity, a final alignment is obtained fully automatically or requires some interactive handling to discriminate alternative pathways.

Algorithms

Numerous group I introns with variable distributions in the ribosomal DNA of a lichen fungus.

The length of the small subunit ribosomal DNA (SSU rDNA) differs significantly among individuals from natural populations of the ascomycetous lichen complex Cladonia chlorophaea. The sequence of the 3' region of the SSU rDNA from two individuals, chosen to represent the shortest and longest sequences, revealed multiple insertions within a region that otherwise aligned with a 520-nucleotide sequence of the SSU rDNA in Saccharomyces cerevisiae. The high degree of variability in SSU rDNA size can be accounted for by different numbers of insertions; one individual had two group I introns and the second had five introns, two of which were clearly related to introns at identical positions in the other individual. Yet, introns in different positions, whether within an individual or between individuals, were not similar in sequence. The distribution of introns at three of the positions is consistent with either intron loss or acquisition, and clearly indicates the dynamic variability in this region of the nuclear genome. All seven insertions, which ranged in size from 210 to 228 nucleotides, had the conserved sequence and secondary structural elements of group I introns. The variation in distribution and sequence of group I introns within a short highly conserved region of rDNA presents a unique opportunity for examining the molecular evolution and mobility of group I introns within a systematics framework.

Ascomycota

Secondary structural predictions for the clostridial neurotoxins.

The primary structures of a family of ten clostridial neurotoxins have recently been deduced yet little information is presently available concerning their secondary or tertiary structures. Because the overall similarity percentage of multiply aligned sequences is high, the secondary structures of these metalloendopeptidases are also expected to be conserved. The neural net program, PHD (Rost and Sander, Proc. Natl. Acad. Sci. USA 90:7558-7562, 1993), predicted that the secondary structures of the neurotoxins were indeed conserved in both single and multiple sequence modes of analysis. Predictions for the amounts of helical, extended, and loop states from the single sequence analyses were consistent with previously published data from circular dichroism studies on some of these neurotoxins. In the single analysis mode, only the aligned regions were predicted to show conservation of the three-state structure. In contrast, the multiple sequence analysis predicted that a conserved state (variable loops) also exists in non-aligned regions. Alignments with the primary structure of the prototypic metalloendopeptidase thermolysin showed that about 25% of the residues within this enzyme are similar to those in the neurotoxins. A comparison of thermolysin's known secondary structure with the predictions from this study showed that about 80% of thermolysin's residues could be structurally aligned with those in the neurotoxins. These predictions provide the necessary framework to build a homologous low-resolution tertiary structure of the neurotoxin active site that will be essential in the development of synthetic inhibitors.

Amino Acid Sequence

Repeating sequence homologies in the p36 target protein of retroviral protein kinases and lipocortin, the p37 inhibitor of phospholipase A2.

Although considerable information has emerged on the molecular properties of the p36 target protein its function as well as the possible implications of its tyrosine phosphorylation have remained elusive. Here we show that all sequence segments of p36 published so far can be aligned by homology along the complete sequence of lipocortin, which has been reported recently. This alignment extends beyond multiple Geisow motifs, thought to indicate a sequence principle implicated in Ca2+ and/or lipid binding. While the latter properties are already established for p36 one may expect them also for lipocortin, an inhibitor of phospholipase A2 activity. Certain implications of these results are discussed.

Annexins

Genetic aspects of aromatic amino acid biosynthesis in Lactococcus lactis.

Polymerase chain reaction (PCR) primers designed from a multiple alignment of predicted amino acid sequences from bacterial aroA genes were used to amplify a fragment of Lactococcus lactis DNA. An 8 kb fragment was then cloned from a lambda library and the DNA sequence of a 4.4 kb region determined. This region was found to contain the genes tyrA, aroA, aroK, and pheA, which are involved in aromatic amino acid biosynthesis and folate metabolism. TyrA has been shown to be secreted and AroK also has a signal sequence, suggesting that these proteins have a secondary function, possibly in the transport of amino acids. The aroA gene from L. lactis has been shown to complement an E. coli mutant strain deficient in this gene. The arrangement of genes involved in aromatic amino acid biosynthesis in L. lactis appears to differ from that in other organisms.

3-Phosphoshikimate 1-Carboxyvinyltransferase

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals

Neighboring base composition and transversion/transition bias in a comparison of rice and maize chloroplast noncoding regions.

The correspondence between the transversion/transition ratio and the neighboring base composition in chloroplast DNA is examined. For 18 noncoding regions of the chloroplast genome, alignments between rice (Oryza sativa) and maize (Zea mays) were generated by two different methods. Difficulties of aligning noncoding DNA are discussed, and the alignments are analyzed in a manner that reduces alignment artifacts. Sequence divergence is < 10%, so multiple substitutions at a site are assumed to be rare. Observed substitutions were analyzed with respect to the A+T content of the two immediately flanking bases. It is shown that as this content increases, the proportion of transversions also increases. When both the 5'- and 3'-flanking nucleotides are G or C (A+T content of 0), only 25% of the observed substitutions are transversions. However, when both the 5'- and 3'-flanking nucleotides are A or T (A+T content of 2), 57% of the observed substitutions are transversions. Therefore, the influence of flanking base composition on substitutions, previously reported for a single noncoding region, is a general feature of the chloroplast genome.

Algorithms

Local multiple alignment by consensus matrix.

A new algorithm for aligning several sequences based on the calculation of a consensus matrix and the comparison of all the sequences using this consensus matrix is described. This consensus matrix contains the preference scores of each nucleotide/amino acid and gaps in every position of the alignment. Two modifications of the algorithm corresponding to the evolutionary and functional meanings of the alignment were developed. The first one solves the best-fitting problem without any penalty for end gaps and with an internal gap penalty function independent on the gap length. This algorithm should be used when comparing evolutionary-related proteins for identifying the most conservative residues. The other modification of the algorithm finds the most similar segments in the given sequences. It can be used for finding those parts of the sequences that are responsible for the same biological function. In this case the gap penalty function was chosen to be proportional to the gap length. The result of aligning amino acid sequences of neutral proteases and a compilation of 65 allosteric effectors and substrates of PEP carboxylase are presented.

Algorithms