PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

ANTHEPROT 2.0: a three-dimensional module fully coupled with protein sequence analysis methods.

ANTHEPROT is a fully interactive graphics program devoted to the analysis of the sequences and structures of proteins. This program, originally developed to facilitate the protein sequence analysis coupled with multiple alignments and predicted secondary structures of proteins, now comprises a powerful 3D module to display and handle macromolecular structures. All the methods that were previously integrated into ANTHEPROT are now directly coupled with a 3D window that provides the user all the classic features of a molecular modeling package. Indeed, it allows real-time rotation and translation of 3D structures with many kinds of models in depth-cueing mode (space filling, backbone, wire models, main chain, and ribbons), selections (atom type, residue type, segments, and chain), color-coding systems (amino acid properties, predicted or observed secondary structures, temperature B factor, and subunits), geometric calculations (Ramachandran plot, distances, and angles), and fitting molecules. Stereo views are possible as well as HPGL standard files. A module specifically devoted to the determination of 3D structures using nuclear magnetic resonance is also available. This major release of our program for IBM rs6000 workstations is available by anonymous ftp to ibcp.fr for academic institutions.

Antigens

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software

Distribution and molecular characterization of integron classes from Escherichia coli and Klebsiella pneumoniae isolates in Sulaymaniyah province of Iraq.

UNLABELLED: The environmental pollution from the misuse of antimicrobial drugs is fueling selection pressure in bacteria, thereby exacerbating the threat to global health. In Iraq, the situation is made worse by the poor implementation of the World Health Organization's Global Antimicrobial Resistance and Use Surveillance System (WHO-GLASS). Consequently, this study aimed to increase surveillance of the spread of antimicrobial resistance in Sulaymaniyah, Iraq. A total of 296 Enterobacteriaceae comprising 147 Klebsiella pneumoniae and 149 Escherichia coli were isolated from humans, poultry, and dairy farms. The isolates were screened using multiplex PCR to assess the prevalence of the clinically important integron integrase (intI) classes and antimicrobial resistance genes (ARGs) of commonly used antibiotics. Remarkably, 81.14% of the isolates carried at least 2 ARGs, 10.47% intI1, and 3.72% intI2. No intI3 was detected. A total of 663 ARGs were identified using multiplex PCR in the two Enterobacteriaceae: beta-lactamase genes were 43%, tetracycline resistance genes 25.20%, sulfonamide resistance gene 16.10%, quinolone resistance gene 10.2%, and aminoglycoside resistance genes 5.7%. K. pneumoniae harbored more integrons and ARGs than E. coli, thus posing a higher antimicrobial resistance threat in this province. This study underscores the importance of implementing more stringent WHO-GLASS and antibiotic stewardship to end the multidrug resistance crisis in Iraq. IMPORTANCE: These data are about the prevalence of integrons and resistance genes, helping to fill a significant gap in global surveillance efforts. Results can be used by global health authorities and the World Health Organization to develop national and international antimicrobial resistance (AMR) control strategies. The study is important because integrons are key genetic platforms that capture and disseminate antibiotic resistance genes among bacteria. In addition, Escherichia coli and Klebsiella spp. are among the top causes of hospital- and community-acquired infections, especially urinary tract infections, bloodstream infections, and pneumonia. Therefore, it will be riskier when these bacteria have a high rate of integrons and resistance genes because it impedes treatments during infection. Another importance of this study is that the study was carried out in Iraq. Iraq, like many low- and middle-income countries, faces challenges with unregulated antibiotic use, leading to high rates of AMR.

Escherichia coli

Positive and negative regulatory elements of the rabbit embryonic epsilon-globin gene revealed by an improved multiple alignment program and functional analysis.

The epsilon-globin genes of mammals are expressed in early embryos, but are silenced during fetal and adult erythropoiesis. As a guide to defining the regulatory elements involved in this developmental switch, we have searched the sequences of epsilon-globin genes from different mammals for highly conserved segments. The search was facilitated by the development of a new program, called yama, to generate a multiple alignment of very long sequences using an improved scoring scheme. This allowed us to generate a multiple alignment of sequences from a more divergent group than previously analyzed, as demonstrated here for representatives of four mammalian orders. In parallel experiments, we constructed a series of deletion mutations in the 5' flank of the rabbit epsilon-globin gene and tested their effect on an epsilon-globin-luciferase hybrid reporter gene. These results show that 121 bp of 5' flank, containing CACC, CCAAT and ATA motifs, is sufficient for expression in erythroid K562 cells. Both positive and negative cis-acting control sequences are located between 218 and 394 bp 5' to the cap site, in a region previously proposed to be a silencer. The positive regulatory sequence contains conserved binding sites for the nuclear protein YY1 adjacent to another highly conserved sequence. The negative element contains a conserved sequence followed by a purine-rich segment. This analysis maps the upstream control sequences more precisely and points to a very complex regulatory scheme for this gene.

Animals

Sequence complexity of the S receptor kinase gene family in Brassica.

A genomic library from an S29/S29 self-incompatible genotype of Brassica oleracea was screened with a probe carrying part of the catalytic domain of a Brassica S-receptor kinase (SRK)-like gene. Six positive phage clones with varying hybridisation intensities (K1 to K6) were purified and characterised. A 650-700 bp region corresponding to the probe was excised from each clone and sequenced. DNA and predicted protein sequence comparisons based on a multiple alignment identified K5 as a pseudogene, whereas the others could encode functional proteins. K3 was found to have lost an intron from its genomic sequence. The six genes display different degrees of sequence similarity and form two distinct clusters in a dendrogram. The 98% similarity between K4 and K6, which extends across intron sequences, suggests that these might be very recently diverged alleles or daughters of a duplication. In addition, K2 showed a comparably high similarity to the probe. Clones K1, K3 and K5 cross-hybridised with an SLG29 cDNA probe, indicating the presence of upstream receptor domains homologous to the Brassica SLG gene. This suggests that the previously reported S sequence complexity may be ascribed to a large receptor kinase gene family.

Base Sequence

On global sequence alignment.

We present a dynamic programming algorithm for computing a best global alignment of two sequences. The proposed algorithm is robust in identifying any of several global relationships between two sequences. The algorithm delivers a best alignment of two sequences in linear space and quadratic time. We also describe a multiple alignment algorithm based on the pairwise algorithm. Both algorithms have been implemented as portable C programs. Experimental results indicate that for a commonly used set of gap penalties, the new programs produce more satisfactory alignments on sequences of various lengths than some existing pairwise and multiple programs based on the dynamic programming algorithm of Needleman and Wunsch.

Algorithms

Fowlpox virus encodes a protein related to human deoxycytidine kinase: further evidence for independent acquisition of genes for enzymes of nucleotide metabolism by different viruses.

It is demonstrated that fowlpox virus (FPV) protein FP26 located in the HindIII D fragment of the genome is related to the human deoxycytidine kinase (dCK) and probably possesses the same enzymatic activity. A homologous protein is not encoded by vaccinia virus. A multiple alignment of the amino acid sequences of the human and FPV dCKs, the thymidine kinases (TK) of herpesviruses, and cellular and vaccinia virus thymidylate kinases (ThyK) was generated and the conserved motifs, at least two of which are implicated in ATP binding, were characterized. An apparent duplication of ATP-binding motif B in the dCKs was revealed, leading to the reassignment of one of the catalytic residues. Phylogenetic analysis based on the multiple alignment suggested that the putative dCK of FPV probably has diverged from the common ancestor with the human dCK at a later stage of evolution than the herpesvirus TKs, with the ThyKs being peripheral members of the family. These results are compatible with hypothesis that genes for enzymes of nucleotide metabolism could be acquired independently by different DNA viruses (Koonin, E.V. and Senkevich, T.G., Virus Genes 6:187-196, 1992).

Amino Acid Sequence

MATCH-BOX: a fundamentally new algorithm for the simultaneous alignment of several protein sequences.

Original algorithms for simultaneous alignment of protein sequences are presented, including sequence clustering and within- or between-groups multiple alignment. The way of matching similar regions is fundamentally new. Complete matches are formed by segments more similar than expected by random, according to a given probability limit. Any classic or user-defined score matrix can be used to express the similarity between the residues. The algorithm seeks for complete matches common to all the sequences without performing pairwise alignment and regardless of gap weighting. An automatic screening delineates all the similar regions (boxes) that may be defined for a given maximal shift between the sequences. The shift can be large enough to allow the matching of any region of a sequence with any region of another one. It can also be short and used to refine the alignment around anchor points. The algorithm provides the most likely optimal alignment and a comprehensive list of the alignment dilemma. Duality between automatism and interactivity is provided. Depending on the problem complexity, a final alignment is obtained fully automatically or requires some interactive handling to discriminate alternative pathways.

Algorithms

Numerous group I introns with variable distributions in the ribosomal DNA of a lichen fungus.

The length of the small subunit ribosomal DNA (SSU rDNA) differs significantly among individuals from natural populations of the ascomycetous lichen complex Cladonia chlorophaea. The sequence of the 3' region of the SSU rDNA from two individuals, chosen to represent the shortest and longest sequences, revealed multiple insertions within a region that otherwise aligned with a 520-nucleotide sequence of the SSU rDNA in Saccharomyces cerevisiae. The high degree of variability in SSU rDNA size can be accounted for by different numbers of insertions; one individual had two group I introns and the second had five introns, two of which were clearly related to introns at identical positions in the other individual. Yet, introns in different positions, whether within an individual or between individuals, were not similar in sequence. The distribution of introns at three of the positions is consistent with either intron loss or acquisition, and clearly indicates the dynamic variability in this region of the nuclear genome. All seven insertions, which ranged in size from 210 to 228 nucleotides, had the conserved sequence and secondary structural elements of group I introns. The variation in distribution and sequence of group I introns within a short highly conserved region of rDNA presents a unique opportunity for examining the molecular evolution and mobility of group I introns within a systematics framework.

Ascomycota

Secondary structural predictions for the clostridial neurotoxins.

The primary structures of a family of ten clostridial neurotoxins have recently been deduced yet little information is presently available concerning their secondary or tertiary structures. Because the overall similarity percentage of multiply aligned sequences is high, the secondary structures of these metalloendopeptidases are also expected to be conserved. The neural net program, PHD (Rost and Sander, Proc. Natl. Acad. Sci. USA 90:7558-7562, 1993), predicted that the secondary structures of the neurotoxins were indeed conserved in both single and multiple sequence modes of analysis. Predictions for the amounts of helical, extended, and loop states from the single sequence analyses were consistent with previously published data from circular dichroism studies on some of these neurotoxins. In the single analysis mode, only the aligned regions were predicted to show conservation of the three-state structure. In contrast, the multiple sequence analysis predicted that a conserved state (variable loops) also exists in non-aligned regions. Alignments with the primary structure of the prototypic metalloendopeptidase thermolysin showed that about 25% of the residues within this enzyme are similar to those in the neurotoxins. A comparison of thermolysin's known secondary structure with the predictions from this study showed that about 80% of thermolysin's residues could be structurally aligned with those in the neurotoxins. These predictions provide the necessary framework to build a homologous low-resolution tertiary structure of the neurotoxin active site that will be essential in the development of synthetic inhibitors.

Amino Acid Sequence

Genetic organization of the streptokinase region of the Streptococcus equisimilis H46A chromosome.

The complete nucleotide sequences of four genes and one open reading frame (ORF1) adjacent to the streptokinase gene, skc, from Streptococcus equisimilis H46A were determined. These genes are encoded on the opposite DNA strand to skc and are arranged as follows: dexB-abc-lrp-skc-ORF1-rel. The dexB gene, coding for an alpha-glucosidase (M(r) 61,733), and abc, encoding an ABC transporter (M(r) 42,080), are similar to the dexB and msmK genes, respectively, from the multiple sugar metabolism operon of S. mutans. The lrp gene specifies a leucine-rich protein (M(r) 32,302) that has a leucine-zipper motif at its C-terminus. The function of the Lrp protein is not known but appeared to be detrimental when overexpressed in Escherichia coli. Although lrp appears not to be an essential gene, as judged by plasmid insertion mutagenesis, it is conserved in all streptococcal strains carrying a streptokinase gene. The rel gene showed significant homology to the E. coli relA and spoT genes involved in the stringent response to amino acid deprivation. Multiple alignment of the amino acid sequences of Rel (M(r) 83,913), RelA and SpoT revealed 59.4% homology of the primary structures. Northern hybridization analyses of the genes in the skc region showed skc to be transcribed most abundantly. In addition to transcripts for skc, monocistronic mRNAs were detected for all three genes divergently transcribed from skc. Although there was also some read-through transcription from lrp into abc, and from abc into dexB, the transcription pattern suggests a high degree of transcriptional and functional independence not only of skc but also abc and dexB. Prominent structural features in intergenic regions included a static DNA bending locus located upstream and a putative bidirectional transcription terminator downstream of skc.

Amino Acid Sequence

Repeating sequence homologies in the p36 target protein of retroviral protein kinases and lipocortin, the p37 inhibitor of phospholipase A2.

Although considerable information has emerged on the molecular properties of the p36 target protein its function as well as the possible implications of its tyrosine phosphorylation have remained elusive. Here we show that all sequence segments of p36 published so far can be aligned by homology along the complete sequence of lipocortin, which has been reported recently. This alignment extends beyond multiple Geisow motifs, thought to indicate a sequence principle implicated in Ca2+ and/or lipid binding. While the latter properties are already established for p36 one may expect them also for lipocortin, an inhibitor of phospholipase A2 activity. Certain implications of these results are discussed.

Annexins

Genetic aspects of aromatic amino acid biosynthesis in Lactococcus lactis.

Polymerase chain reaction (PCR) primers designed from a multiple alignment of predicted amino acid sequences from bacterial aroA genes were used to amplify a fragment of Lactococcus lactis DNA. An 8 kb fragment was then cloned from a lambda library and the DNA sequence of a 4.4 kb region determined. This region was found to contain the genes tyrA, aroA, aroK, and pheA, which are involved in aromatic amino acid biosynthesis and folate metabolism. TyrA has been shown to be secreted and AroK also has a signal sequence, suggesting that these proteins have a secondary function, possibly in the transport of amino acids. The aroA gene from L. lactis has been shown to complement an E. coli mutant strain deficient in this gene. The arrangement of genes involved in aromatic amino acid biosynthesis in L. lactis appears to differ from that in other organisms.

3-Phosphoshikimate 1-Carboxyvinyltransferase

A data bank merging related protein structures and sequences.

A data collection which merges protein structural and sequence information is described. Structural superpositions amongst proteins with similar main-chain fold were performed or collected from the literature. Sequences taken from the protein primary structure databases were associated with the multiple structural alignments providing they were at least 50% homologous in residue identity to one of the structural sequences and at least 50% of the structural sequence residues were alignable. Such restrictions allow reasonable confidence that the primary sequences share the conformation of the tertiary structural templates, except in the less conserved loop regions. Multiple structural superpositions were collected for 38 familial groups containing a total of 209 tertiary structures; 45 structures had no superposable mates and were used individually. Other information is also provided as main-chain and side-chain conformational angles, secondary structural assignments and the like. Wedding the primary and tertiary structural data resulted in an 8-fold increase of data bank sequence entries over those associated with the known three-dimensional architectures alone.

Amino Acid Sequence

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals

Neighboring base composition and transversion/transition bias in a comparison of rice and maize chloroplast noncoding regions.

The correspondence between the transversion/transition ratio and the neighboring base composition in chloroplast DNA is examined. For 18 noncoding regions of the chloroplast genome, alignments between rice (Oryza sativa) and maize (Zea mays) were generated by two different methods. Difficulties of aligning noncoding DNA are discussed, and the alignments are analyzed in a manner that reduces alignment artifacts. Sequence divergence is < 10%, so multiple substitutions at a site are assumed to be rare. Observed substitutions were analyzed with respect to the A+T content of the two immediately flanking bases. It is shown that as this content increases, the proportion of transversions also increases. When both the 5'- and 3'-flanking nucleotides are G or C (A+T content of 0), only 25% of the observed substitutions are transversions. However, when both the 5'- and 3'-flanking nucleotides are A or T (A+T content of 2), 57% of the observed substitutions are transversions. Therefore, the influence of flanking base composition on substitutions, previously reported for a single noncoding region, is a general feature of the chloroplast genome.

Algorithms

Local multiple alignment by consensus matrix.

A new algorithm for aligning several sequences based on the calculation of a consensus matrix and the comparison of all the sequences using this consensus matrix is described. This consensus matrix contains the preference scores of each nucleotide/amino acid and gaps in every position of the alignment. Two modifications of the algorithm corresponding to the evolutionary and functional meanings of the alignment were developed. The first one solves the best-fitting problem without any penalty for end gaps and with an internal gap penalty function independent on the gap length. This algorithm should be used when comparing evolutionary-related proteins for identifying the most conservative residues. The other modification of the algorithm finds the most similar segments in the given sequences. It can be used for finding those parts of the sequences that are responsible for the same biological function. In this case the gap penalty function was chosen to be proportional to the gap length. The result of aligning amino acid sequences of neutral proteases and a compilation of 65 allosteric effectors and substrates of PEP carboxylase are presented.

Algorithms