PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Comparison of side chain interactions performed by structurally equivalent residues in homologous protein structures.

The present work describes the computer program Hom-Bond, which allows to identify and compare intra-molecular interactions performed by side chain polar atoms as observed in a family of homologous protein structures with known and conserved 3-D conformation. For this purpose, the side chain to side chain and the side chain to main chain hydrogen bonds, the disulfide and the salt bridges are identified in each considered protein structure. Subsequently, the side chain interactions are displayed according to the multiple sequence alignment. The presented approach allows to easily identify bonds which are conserved in homologous proteins and to analyse rearrangements of the network of side chain interactions that characterize each protein structure.

Amino Acid Sequence

Sisyphus and prediction of protein structure.

The problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and 1D prediction have reopened the field. Mean-force-potentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue contacts (2D) can be detected by analysis of correlated mutations, albeit with low accuracy. Secondary structure, solvent accessibility and transmembrane helices (1D) can be predicted with significantly improved accuracy using multiple sequence alignments. Some of these new prediction methods have proven accurate and reliable enough to be useful in genome analysis, and in experimental structure determination. Moreover, the new generation of theoretical methods is increasingly influencing experiments in molecular biology.

Computers

Efficient discovery of conserved patterns using a pattern graph.

MOTIVATION: We have previously reported an algorithm for discovering patterns conserved in sets of related unaligned protein sequences. The algorithm was implemented in a program called Pratt. Pratt allows the user to define a class of patterns (e.g. the degree of ambiguity allowed and the length and number of gaps), and is then guaranteed to find the conserved patterns in this class scoring highest according to a defined fitness measure. In many cases, this version of Pratt was very efficient, but in other cases it was too time consuming to be applied. Hence, a more efficient algorithm was needed. RESULTS: In this paper, we describe a new and improved searching strategy that has two main advantages over the old strategy. First, it allows for easier integration with programs for multiple sequence alignment and data base search. Secondly, it makes it possible to use branch-and-bound search, and heuristics, to speed up the search. The new search strategy has been implemented in a new version of the Pratt program.

Algorithms

TOPAL: recombination detection in DNA and protein sequences.

UNLABELLED: TOPAL scans a multiple sequence alignment for evidence of recombinant sequences, prior to phylogenetic analysis. AVAILABILITY: The TOPAL package may be accessed at http://www.bioss.sari.ac.uk/grainne, and by anonymous ftp at ftp.bioss. sari.ac.uk in the directory pub/phylogeny/topal. CONTACT: grainne@bioss.sari.ac.uk

Computational Biology

ENVIRON: a software package to compare protein three-dimensional structures with homologous sequences using local structural motifs.

This work presents a method to compare local clusters of interacting residues as observed in a known three-dimensional protein structure with corresponding clusters inferred from homologous protein sequences, assuming conserved protein folding. For this purpose the local environment of a selected residue in a known protein structure is defined as the ensemble of amino acids in contact with it in the folded state. Using a multiple sequence alignment to identify corresponding residues in homologous proteins, a detailed comparison can be performed between the local environment of a selected amino acid in the template protein structure and the expected local environments at the sets of equivalent residues, derived from the aligned protein sequences. The comparison makes it possible to detect conserved local features such as hydrogen bonding or complementarity in residue substitution. A global measure of environmental similarity is also defined, to search for conserved amino acid clusters subject to functional or structural constraints. The proposed approach is useful for investigating protein function as well as for site-directed mutagenesis experiments, where appropriate amino acid substitutions can be suggested by observing naturally occurring protein variants.

Algorithms

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Inference of Cytochrome P450 Evolutionary History Using Structural and Physicochemical Metrics.

Cytochrome P450s are a superfamily of heme-binding monooxygenases involved with the detoxification of intrinsic and extrinsic toxins. They are near ubiquitous within biological domains and are found in all domains. Members of families within the superfamily are defined based on amino acid identity thresholds, with thresholds as low as 40% in some families. Relationships among Cytochrome P450 families have proven elusive due to sub-Twilight Zone interfamily identities (<30%) that result in poor multiple sequence alignment quality and thus low levels of support for downstream phylogenetic reconstructions. Despite the low identities, Cytochrome P450 structures are remarkably well conserved both within and among families. In such cases, structural phylogenetics has the potential to unveil elusive relationships because the selectively favored physicochemical properties giving rise to the structure and function of the proteins persist despite sequence-level divergence. Recently, in two separate publications, we demonstrated that by utilizing physicochemical vectors, dynamic time warping, and hierarchical clustering (PCDTW), large swaths of protein domain families and betacoronavirus receptor-binding domain clades were congruent with validated functional/structural relationships. These were important findings because anomalous sequence alignment-based maximum likelihood phylogenetic findings, which were not congruent with the known functional relationships, were resolved. That also validated the use of physicochemical vectors in making inferences about structural/functional homology. Additionally, it illuminated that the same methods might be applied to other protein families with relationships that are difficult to resolve from sequence data alone. Herein, we used Molecular Weight and Hydrophobicity Physicochemical Dynamic Time Warping (MWHP PCDTW) along with structural and sequence alignment-based phylogenetic methodologies to analyze all of the Cytochrome P450s found both in the high-fidelity Structural Classificaction of Proteins (SCOP) database and the reviewed sequences with both experimentally resolved and de novo predicted structures in the Protein Data Bank and the AlphaFold (AF) Protein Structure Database, respectively. We compared the resulting phylogenetic topologies and found that in some cases, structure-based methods may be less able to resolve random/convergent similarity than physicochemical and sequence-based methodologies. This finding agrees with previous findings that demonstrate the usefulness of physicochemical properties in resolving both random structural similarity and potentially convergent relationships.

Cytochrome P-450 Enzyme System

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Analysis of genetic diversity in cytotoxin-producing and non-cytotoxin-producing Helicobacter pylori strains.

Analysis of 32 Helicobacter pylori strains indicated a strong association between the presence of the cagA gene and a specific type of vacA allele found predominantly in cytotoxin-producing strains (P < .001). To determine whether tox+/CagA+ and tox-/CagA- strains constituted two separate noncombining lineages, sequences of the H. pylori ureC gene, cysS homologue, and the intergenic region between cysS and vacA were determined for multiple strains. The mean levels of nucleotide identity in the three regions were 96.7% +/- 0.5%, 95.0% +/- 1.0%, and 89.0% +/- 2.9%, respectively. Multiple sequence alignments and dendrograms based on these three regions failed to identify two clonal populations of organisms for which cagA and vacA genotypes were markers. The presence of a 63- to 64-bp insertion in the cysS-vacA intergenic region was unrelated to the vacA genotype of the strains. These data suggest that recombination between Helicobacter genomes may occur in vivo.

Base Sequence

Molecular cloning of the cDNA for the catalytic subunit of human DNA polymerase delta.

The cDNA of human DNA polymerase delta was cloned. The cDNA had a length of 3.5 kb and encoded a protein of 1107 amino acid residues with a calculated molecular mass of 124 kDa. Northern blot analysis showed that the cDNA hybridized to a mRNA of 3.4 kb. Monoclonal and polyclonal antibodies to the C-terminal 20 residues specifically immunoblotted the human pol delta catalytic polypeptide. A multiple sequence alignment was constructed. This showed that human pol delta is closely related to yeast pol delta and the herpes virus DNA polymerases. The levels of pol delta message were found to be induced concomitantly with DNA pol delta activity and DNA synthesis in serum restimulated proliferating IMR90 cultured cells. The human pol delta gene was localized to chromosome 19 by Southern blotting of EcoRI digested DNA from a panel of rodent/human cell hybrids.

Amino Acid Sequence

Molecular evolution of SRP cycle components: functional implications.

Signal recognition particle (SRP) is a cytoplasmic ribonucleoprotein that targets a subset of nascent presecretory proteins to the endoplasmic reticulum membrane. We have considered the SRP cycle from the perspective of molecular evolution, using recently determined sequences of genes or cDNAs encoding homologs of SRP (7SL) RNA, the Srp54 protein (Srp54p), and the alpha subunit of the SRP receptor (SR alpha) from a broad spectrum of organisms, together with the remaining five polypeptides of mammalian SRP. Our analysis provides insight into the significance of structural variation in SRP RNA and identifies novel conserved motifs in protein components of this pathway. The lack of congruence between an established phylogenetic tree and size variation in 7SL homologs implies the occurrence of several independent events that eliminated more than half the sequence content of this RNA during bacterial evolution. The apparently non-essential structures are domain I, a tRNA-like element that is constant in archaea, varies in size among eucaryotes, and is generally missing in bacteria, and domain III, a tightly base-paired hairpin that is present in all eucaryotic and archeal SRP RNAs but is invariably absent in bacteria. Based on both structural and functional considerations, we propose that the conserved core of SRP consists minimally of the 54 kDa signal sequence-binding protein complexed with the loosely base-paired domain IV helix of SRP RNA, and is also likely to contain a homolog of the Srp68 protein. Comparative sequence analysis of the methionine-rich M domains from a diverse array of Srp54p homologs reveals an extended region of amino acid identity that resembles a recently identified RNA recognition motif. Multiple sequence alignment of the G domains of Srp54p and SR alpha homologs indicates that these two polypeptides exhibit significant similarity even outside the four GTPase consensus motifs, including a block of nine contiguous amino acids in a location analogous to the binding site of the guanine nucleotide dissociation stimulator (GDS) for E. coli EF-Tu. The conservation of this sequence, in combination with the results of earlier genetic and biochemical studies of the SRP cycle, leads us to hypothesize that a component of the Srp68/72p heterodimer serves as the GDS for both Srp54p and SR alpha. Using an iterative alignment procedure, we demonstrate similarity between Srp68p and sequence motifs conserved among GDS proteins for small Ras-related GTPases. The conservation of SRP cycle components in organisms from all three major branches of the phylogenetic tree suggests that this pathway for protein export is of ancient evolutionary origin.

Amino Acid Sequence

The Genome Sequence DataBase version 1.0 (GSDB): from low pass sequences to complete genomes.

The Genome Sequence DataBase (GSDB) has completed its conversion to an improved relational database. The new database, GSDB 1.0, is fully operational and publicly available. Data contributions, including both original sequence submissions and community annotation, are being accomplished through the use of a graphical client-server interface tool, the GSDB Annotator, and via GIO (GSDB Input/Output) files. Data retrieval services are being provided through a new Web Query Tool and direct SQL. All methods of data contribution and data retrieval fully support the new data types that have been incorporated into GSDB, including discontiguous sequences, multiple sequence alignments, and community annotation.

Animals

Comparative sequence analysis of ribonucleases HII, III, II PH and D.

Escherichia coli ribonucleases (RNases) HII, III, II, PH and D have been used to characterise new and known viral, bacterial, archaeal and eucaryotic sequences similar to these endo- (HII and III) and exoribonucleases (II, PH and D). Statistical models, hidden Markov models (HMMs), were created for the RNase HII, III, II and PH and D families as well as a double-stranded RNA binding domain present in RNase III. Results suggest that the RNase D family, which includes Werner syndrome protein and the 100 kDa antigenic component of the human polymyositis scleroderma (PMSCL) autoantigen, is a 3'-->5' exoribonuclease structurally and functionally related to the 3'-->5' exodeoxyribonuclease domain of DNA polymerases. Polynucleotide phosphorylases and the RNase PH family, which includes the 75 kDa PMSCL autoantigen, possess a common domain suggesting similar structures and mechanisms of action for these 3'-->5' phosphorolytic enzymes. Examination of HMM-generated multiple sequences alignments for each family suggest amino acids that may be important for their structure, substrate binding and/or catalysis.

Amino Acid Sequence

The proofreading domain of Escherichia coli DNA polymerase I and other DNA and/or RNA exonuclease domains.

Prior sequence analysis studies have suggested that bacterial ribonuclease (RNase) Ds comprise a complete domain that is found also in Homo sapiens polymyositis-scleroderma overlap syndrome 100 kDa autoantigen and Werner syndrome protein. This RNase D 3'-->5' exoribonuclease domain was predicted to have a structure and mechanism of action similar to the 3'-->5' exodeoxyibonuclease (proofreading) domain of DNA polymerases. Here, hidden Markov model (HMM) and phylogenetic studies have been used to identify and characterise other sequences that may possess this exonuclease domain. Results indicate that it is also present in the RNase T family; Borrelia burgdorferi P93 protein, an immunodominant antigen in Lyme disease; bacteriophage T4 dexA and Escherichia coli exonuclease I, processive 3'-->5' exodeoxyribonucleases that degrade single-stranded DNA; Bacillus subtilis dinG, a probable helicase involved in DNA repair and possibly replication, and peptide synthase 1; Saccharomyces cerevisiae Pab1p-dependent poly(A) nuclease PAN2 subunit, required for shortening mRNA poly(A) tails; Caenorhabditis elegans and Mus musculus CAF1, transcription factor CCR4-associated factor 1; Xenopus laevis XPMC2, prevention of mitotic catastrophe in fission yeast; Drosophila melanogaster egalitarian, oocyte specification and axis determination, and exuperantia, establishment of oocyte polarity; H.sapiens HEM45, expressed in tumour cell lines and uterus and regulated by oestrogen; and 31 open reading frames including one in Methanococcus jannaschii . Examination of a multiple sequence alignment and two three-dimensional structures of proofreading domains has allowed definition of the core sequence, structural and functional elements of this exonuclease domain.

Amino Acid Sequence

GPCRDB: an information system for G protein-coupled receptors.

The GPCRDB is a G protein-coupled receptor (GPCR) database system aimed at the collection and dissemination of GPCR related data. It holds sequences, mutant data and ligand binding constants as primary (experimental) data. Computationally derived data such as multiple sequence alignments, three dimensional models, phylogenetic trees and two dimensional visualization tools are added to enhance the database's usefulness. The GPCRDB is an EU sponsored project aimed at building a generic molecular class specific database capable of dealing with highly heterogeneous data. GPCRs were chosen as test molecules because of their enormous importance for medical sciences and due to the availability of so much highly heterogeneous data. The GPCRDB is available via the WWW at http://www.gpcr.org/7tm

Computer Communication Networks

Histone Sequence Database: new histone fold family members.

Searches of the major public protein databases with core and linker chicken and human histone sequences have resulted in the compilation of an annotated set of histone protein sequences. In addition, new database searches with two distinct motif search algorithms have identified several members of the histone fold family, including human DRAP1 and yeast CSE4. Database resources include information on conflicts between similar sequence entries in different source databases, multiple sequence alignments, links to the Entrez integrated information retrieval system, structures for histone and histone fold proteins, and the ability to visualize structural data through Cn3D. The database currently contains >1000 protein sequences, which are searchable by protein type, accession number, organism name, or any other free text appearing in the definition line of the entry. All sequences and alignments in this database are available through the World Wide Web at http://www.nhgri.nih. gov/DIR/GTB/HISTONES or http://www.ncbi.nlm.nih. gov/Baxevani/HISTONES

Amino Acid Sequence

Construction of a 3D model of cytochrome P450 2B4.

A three-dimensional structural model of rabbit phenobarbital-inducible cytochrome P450 2B4 (LM2) was constructed by homology modeling techniques previously developed for building and evaluating a 3D model of the cytochrome P450choP isozyme. Four templates with known crystal structures including cytochrome P450cam, terp, BM-3 and eryF were used in multiple sequence alignments and construction of the cytochrome P450 2B4 coordinates. The model was evaluated for its overall quality using available protein analysis programs and found to be satisfactory. The model structure was stable at room temperature during a 140 ps unconstrained full protein molecular dynamics simulation. A putative substrate access channel and binding site were identified. Two different substrates, benzphetamine and androstenedione, that are metabolized by cytochrome P450 2B4 with pronounced product specificity were docked into the putative binding site. Two orientations were found for each substrate that could lead to the observed preferred products. Using a geometric fit method three regions on the surface of the model cytochrome P450 structure were identified as possible sites for interaction with cytochrome b5, a redox partner of P450 2B4. Residues that may interact with the substrates and with cytochrome b5 have been identified and mutagenesis studies are currently in progress.

Amino Acid Sequence

Interactions underlying subunit association in cholinesterases.

Cholinesterases occur in a family of molecular forms, both as homo-oligomers of catalytic subunits, which can be either soluble, amphiphilic or lipid-anchored to the membrane; and hetero-oligomers of catalytic subunits and structural subunits. The structural subunits afford a method for precise localization of cholinesterases for specific function. A number of mutagenesis studies suggest that the C-terminal region of one alternatively spliced form of cholinesterase is involved in association of catalytic subunits into tetramers and in the association of these tetramers with structural subunits, however, there is currently no structural information about this region. In addition, none of the mutagenesis studies have clearly defined the residues important in these interactions. Here, multiple sequence alignment, structure prediction techniques and analysis of three-dimensional structural data are combined with a re-examination of mutagenesis and biochemical data. Three-dimensional models for the C-terminal region and for soluble tetrameric cholinesterase are proposed, and a set of rules governing subunit association are formulated. The simple model for association of catalytic and structural subunits presented is consistent with data for all known cholinesterases from species as divergent as nematode and man.

Amino Acid Sequence