PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Sequence similarities within the family of dihydrolipoamide acyltransferases and discovery of a previously unidentified fungal enzyme.

A composite protein sequence database was searched for amino acid sequences similar to the C-terminal domain of the dihydrolipoamide acetyltransferase subunit (E2p) of the pyruvate dehydrogenase complex of Escherichia coli. Nine sequences with extensive similarity were found, of which eight were E2 subunits. The other was for a putative mitochondrial ribosomal protein, MRP3, from Neurospora crassa. Alignment of the MRP3 and E2 sequences showed that the similarity extends through the entire MRP3 sequence and that MRP3 is most closely related to the E2p subunit of the pyruvate dehydrogenase complex from Saccharomyces cerevisiae, with 54% identical residues and a further 36% that are conservatively substituted. Other features of the MRP3 gene and protein are also consistent with it being the acyltransferase subunit of a 2-oxo acid dehydrogenase complex. A multiple alignment of 13 E2 sequences indicated that 120 (34%) of 353 equivalenced residues are identical or show some degree of conservation. It also identified residues that are potentially important for the structure, catalytic activity and substrate-specificity of the acyltransferases.

Acetyltransferases

Protein database searches for multiple alignments.

Protein database searches frequently can reveal biologically significant sequence relationships useful in understanding structure and function. Weak but meaningful sequence patterns can be obscured, however, by other similarities due only to chance. By searching a database for multiple as opposed to pairwise alignments, distant relationships are much more easily distinguished from background noise. Recent statistical results permit the power of this approach to be analyzed. Given a typical query sequence, an algorithm described here permits the current protein database to be searched for three-sequence alignments in less than 4 min. Such searches have revealed a variety of subtle relationships that pairwise search methods would be unable to detect.

Algorithms

Evaluation of the sequence template method for protein structure prediction. Discrimination of the (beta/alpha)8-barrel fold.

A multiple alignment of five (beta/alpha)8-barrel enzymes has been derived from their structure. The eight beta-strands and eight alpha-helices of the (beta/alpha)8-barrel are correctly aligned and the equivalenced residues in these regions fulfil similar structural roles. Each beta-strand has a central core of usually four residues, two residues contribute side-chains to the barrel core and the other two residues are involved in beta-strand/alpha-helix contacts. However, the fold imposes no constraints on the volumes of the residues at either a local or global level: the volume of the beta-barrel core varies between 1088 A3 in glycolate oxidase and 1571 A3 in taka-amylase. Sequence motifs derived from the multiple alignment were scanned against a database of 124 protein sequences, including 17 (beta/alpha)8-barrel enzymes. The results were evaluated in terms of the discrimination of (beta/alpha)8-barrel sequences and the quality of the alignments obtained. One motif was able to identify the top 12% of high scoring sequences as forming (beta/alpha)8-barrels with 50% accuracy and the bottom 50% of sequences as not being (beta/alpha)8-barrel proteins with 100% accuracy. However, in most instances the alignments were poor. The reasons for this are discussed with reference to the (beta/alpha)8-barrel proteins and the sequence motif method in general.

Alcohol Oxidoreductases

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software

Distribution and molecular characterization of integron classes from Escherichia coli and Klebsiella pneumoniae isolates in Sulaymaniyah province of Iraq.

UNLABELLED: The environmental pollution from the misuse of antimicrobial drugs is fueling selection pressure in bacteria, thereby exacerbating the threat to global health. In Iraq, the situation is made worse by the poor implementation of the World Health Organization's Global Antimicrobial Resistance and Use Surveillance System (WHO-GLASS). Consequently, this study aimed to increase surveillance of the spread of antimicrobial resistance in Sulaymaniyah, Iraq. A total of 296 Enterobacteriaceae comprising 147 Klebsiella pneumoniae and 149 Escherichia coli were isolated from humans, poultry, and dairy farms. The isolates were screened using multiplex PCR to assess the prevalence of the clinically important integron integrase (intI) classes and antimicrobial resistance genes (ARGs) of commonly used antibiotics. Remarkably, 81.14% of the isolates carried at least 2 ARGs, 10.47% intI1, and 3.72% intI2. No intI3 was detected. A total of 663 ARGs were identified using multiplex PCR in the two Enterobacteriaceae: beta-lactamase genes were 43%, tetracycline resistance genes 25.20%, sulfonamide resistance gene 16.10%, quinolone resistance gene 10.2%, and aminoglycoside resistance genes 5.7%. K. pneumoniae harbored more integrons and ARGs than E. coli, thus posing a higher antimicrobial resistance threat in this province. This study underscores the importance of implementing more stringent WHO-GLASS and antibiotic stewardship to end the multidrug resistance crisis in Iraq. IMPORTANCE: These data are about the prevalence of integrons and resistance genes, helping to fill a significant gap in global surveillance efforts. Results can be used by global health authorities and the World Health Organization to develop national and international antimicrobial resistance (AMR) control strategies. The study is important because integrons are key genetic platforms that capture and disseminate antibiotic resistance genes among bacteria. In addition, Escherichia coli and Klebsiella spp. are among the top causes of hospital- and community-acquired infections, especially urinary tract infections, bloodstream infections, and pneumonia. Therefore, it will be riskier when these bacteria have a high rate of integrons and resistance genes because it impedes treatments during infection. Another importance of this study is that the study was carried out in Iraq. Iraq, like many low- and middle-income countries, faces challenges with unregulated antibiotic use, leading to high rates of AMR.

Escherichia coli

MATCH-BOX: a fundamentally new algorithm for the simultaneous alignment of several protein sequences.

Original algorithms for simultaneous alignment of protein sequences are presented, including sequence clustering and within- or between-groups multiple alignment. The way of matching similar regions is fundamentally new. Complete matches are formed by segments more similar than expected by random, according to a given probability limit. Any classic or user-defined score matrix can be used to express the similarity between the residues. The algorithm seeks for complete matches common to all the sequences without performing pairwise alignment and regardless of gap weighting. An automatic screening delineates all the similar regions (boxes) that may be defined for a given maximal shift between the sequences. The shift can be large enough to allow the matching of any region of a sequence with any region of another one. It can also be short and used to refine the alignment around anchor points. The algorithm provides the most likely optimal alignment and a comprehensive list of the alignment dilemma. Duality between automatism and interactivity is provided. Depending on the problem complexity, a final alignment is obtained fully automatically or requires some interactive handling to discriminate alternative pathways.

Algorithms

Numerous group I introns with variable distributions in the ribosomal DNA of a lichen fungus.

The length of the small subunit ribosomal DNA (SSU rDNA) differs significantly among individuals from natural populations of the ascomycetous lichen complex Cladonia chlorophaea. The sequence of the 3' region of the SSU rDNA from two individuals, chosen to represent the shortest and longest sequences, revealed multiple insertions within a region that otherwise aligned with a 520-nucleotide sequence of the SSU rDNA in Saccharomyces cerevisiae. The high degree of variability in SSU rDNA size can be accounted for by different numbers of insertions; one individual had two group I introns and the second had five introns, two of which were clearly related to introns at identical positions in the other individual. Yet, introns in different positions, whether within an individual or between individuals, were not similar in sequence. The distribution of introns at three of the positions is consistent with either intron loss or acquisition, and clearly indicates the dynamic variability in this region of the nuclear genome. All seven insertions, which ranged in size from 210 to 228 nucleotides, had the conserved sequence and secondary structural elements of group I introns. The variation in distribution and sequence of group I introns within a short highly conserved region of rDNA presents a unique opportunity for examining the molecular evolution and mobility of group I introns within a systematics framework.

Ascomycota

Repeating sequence homologies in the p36 target protein of retroviral protein kinases and lipocortin, the p37 inhibitor of phospholipase A2.

Although considerable information has emerged on the molecular properties of the p36 target protein its function as well as the possible implications of its tyrosine phosphorylation have remained elusive. Here we show that all sequence segments of p36 published so far can be aligned by homology along the complete sequence of lipocortin, which has been reported recently. This alignment extends beyond multiple Geisow motifs, thought to indicate a sequence principle implicated in Ca2+ and/or lipid binding. While the latter properties are already established for p36 one may expect them also for lipocortin, an inhibitor of phospholipase A2 activity. Certain implications of these results are discussed.

Annexins

A data bank merging related protein structures and sequences.

A data collection which merges protein structural and sequence information is described. Structural superpositions amongst proteins with similar main-chain fold were performed or collected from the literature. Sequences taken from the protein primary structure databases were associated with the multiple structural alignments providing they were at least 50% homologous in residue identity to one of the structural sequences and at least 50% of the structural sequence residues were alignable. Such restrictions allow reasonable confidence that the primary sequences share the conformation of the tertiary structural templates, except in the less conserved loop regions. Multiple structural superpositions were collected for 38 familial groups containing a total of 209 tertiary structures; 45 structures had no superposable mates and were used individually. Other information is also provided as main-chain and side-chain conformational angles, secondary structural assignments and the like. Wedding the primary and tertiary structural data resulted in an 8-fold increase of data bank sequence entries over those associated with the known three-dimensional architectures alone.

Amino Acid Sequence

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals

Isolation and characterization of the alkane-inducible NADPH-cytochrome P-450 oxidoreductase gene from Candida tropicalis. Identification of invariant residues within similar amino acid sequences of divergent flavoproteins.

The gene coding for the Candida tropicalis NADPH-cytochrome P-450 oxidoreductase (CPR, NADPH: ferricytochrome oxidoreductase, EC 1.6.2.4) was isolated by immunoscreening of a C. tropicalis lambda gt11 expression library and colony hybridization of a C. tropicalis genomic library. The C. tropicalis CPR gene produces a 2.35-kilobase mRNA transcript, levels of which were shown to be increased 16-fold in cells grown on tetradecane relative to cells grown on glucose as the sole carbon source. A 3-kilobase DNA fragment was sequenced, including 554 and 397 base pairs of 5'- and 3'-noncoding sequence, respectively. A single open reading frame of 2040 base pairs was identified and predicts a 76,683-Da polypeptide of 680 amino acid residues. The deduced C. tropicalis CPR amino acid sequence was compared with each of the CPR sequences reported from other organisms and invariant residues were identified. Multiple pairwise alignments of divergent members of protein families, previously recognized for their sequence similarities in their respective binding domains for FMN, FAD, and NADPH, have allowed identification of a subset of these invariant residues. From these analyses we infer the importance of 25 of the 680 amino acid residues.

Alkanes

Local multiple alignment by consensus matrix.

A new algorithm for aligning several sequences based on the calculation of a consensus matrix and the comparison of all the sequences using this consensus matrix is described. This consensus matrix contains the preference scores of each nucleotide/amino acid and gaps in every position of the alignment. Two modifications of the algorithm corresponding to the evolutionary and functional meanings of the alignment were developed. The first one solves the best-fitting problem without any penalty for end gaps and with an internal gap penalty function independent on the gap length. This algorithm should be used when comparing evolutionary-related proteins for identifying the most conservative residues. The other modification of the algorithm finds the most similar segments in the given sequences. It can be used for finding those parts of the sequences that are responsible for the same biological function. In this case the gap penalty function was chosen to be proportional to the gap length. The result of aligning amino acid sequences of neutral proteases and a compilation of 65 allosteric effectors and substrates of PEP carboxylase are presented.

Algorithms

ADSP--a new package for computational sequence analysis.

A new protein sequence analysis package, ADSP, is described, of which the SOMAP Screen-Oriented Multiple Alignment Procedure forms an integral part. ADSP (Algorithms and Data Structures for Protein sequence analysis) incorporates facilities to generate potent pattern-recognition discriminators and offers four algorithms with which to scan any NBRF format sequence database: the package has been designed, in particular, to interface with the OWL composite sequence database, one of the largest, distributed non-redundant sources of sequence data of its kind. The system incorporates a powerful method for compound feature analysis, which provides the basis for characterizing and predicting the occurrence of complete protein superfamilies and for pinpointing the emergence of related sub-families. Used iteratively, the approach allows diagnostic performance to be rigorously refined and its efficacy to be assessed both qualitatively and quantitatively, and results in the generation of refined structural or functional features suitable for entry into a database: this compilation of characteristic signatures is distinct from, but complementary to, widely used compendia of pattern templates such as PROSITE.

Algorithms

CGEMA and VGAP: a Colour Graphics Editor for Multiple Alignment using a variable GAP penalty. Application to the muscarinic acetylcholine receptor.

Today, more than 40 protein amino acid (AA) sequences of membrane receptors coupled to guanine nucleotide binding proteins (G-proteins) are available. For those working in the field of medicinal chemistry, these sequences present a new type of information that should be taken into consideration. To make maximal use of sequence data it is essential to be able to compare different protein sequences in a similar way to that used for small molecules. A prerequisite, however, is the availability of a processing environment that enables one to handle sequences in an easy way, both by hand and by computer. In order to meet these ends, the package CGEMA (Colour Graphics Editor for Multiple Alignment) was developed in our laboratory. The programme uses a user-definable colour coding for the different AAs. Sequences can be aligned by hand or by computer, using VGAP, and both approaches can be combined. VGAP is a novel in-house written alignment programme with a variable gap penalty that also handles consecutive alignments using one sequence as a probe. In addition, secondary structure prediction tools are available. From the 20 protein sequences, available for the muscarinic acetylcholine receptor, 13 different sequences were selected, covering the subtypes m1 to m5. By comparing the sequences, two major groups are revealed that correspond to those found by considering the transducing system coupled to the various receptor subtypes. Different parts of the protein sequences are identified as characterizing the subtype and binding the ligands, respectively.

Amino Acid Sequence

[Genes of the lipase family: comparison of nucleic and proteinic sequences].

Vertebrates' plasmatic apolipoproteins and a few number of lipases in their metabolism present sequence homologies. They are grouped in genes families. The four exons apolipoproteins gene family includes nine human genes: the divergence rate of their sequences allows to place the first ancestral gene very high in the phylogenetic tree of the evolution. However, a more recent duplication of apolipoprotein C-I gene dating from 40 millions years, may be a phylogenetic marker for the radiation of Monkeys. Pancreatic lipase and isoforms, lipoprotein-lipase and hepatic triacylglycerol-lipase form by their homologies a "superfamily" of genes, which also includes yolk proteins of Dipterians eggs. Sequence homologies of PL, LPL and HL are analysed and compared with multiple alignments of amino-acids and nucleotides on spreadsheets. From these comparisons we may characterize four classes of phylogenetic markers: 1) repetitive DNA sequence (Alu, B1, PRE-1) appeared during Mammals evolution, 2) short insertions or deletions (within N-terminal domain) and a gene conversion in guinea-pig lineage, 3) a progressive reduction of intron number during the lipases evolution, 4) several duplications of genes which have produced the five genes of this superfamily currently known in the human genome.

Amino Acid Sequence

An efficient and reliable method for cloning PCR-amplification products: a survey of point mutations in integrin cDNA.

A highly efficient, non-labor-intensive method for cloning DNA fragments produced by PCR amplification was used to carry out a rapid survey of potential point mutations in integrin alpha 6 cDNA from 17 different cell-type sources. The method includes glass powder purification of the PCR reaction mixture, followed by simultaneous treatment with T4 polynucleotide kinase and DNA polymerase I, and another glass powder purification. Sequences from multiple subclones of each cell type were readily generated, aligned and checked for mismatches. Several commonly used alternative procedures were compared for cloning efficiency and size-fidelity of inserted DNA fragments.

Amino Acid Sequence