PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Rat liver carboxylesterase: cDNA cloning, sequencing, and evidence for a multigene family.

A cDNA clone was isolated from a rat liver lambda gt11 expression library by screening with polyclonal antibodies raised against a rat liver microsomal carboxylesterase. This clone of 1.8 kb contained an open reading frame encoding a mature protein of 531 amino acids with a predicted molecular weight of 58,084. The 5' portion of the clone coded for 9 amino acids of a putative signal peptide. The 3' end of the clone included an untranslated region and a poly (A) tail. Carboxylesterase active site regions, five potential N-linked glycosylation sites, and 2 postulated cystine disulfide bridges were found in the cDNA-deduced amino acid sequence. Sequences obtained from tryptic peptides and the NH2-terminus of the purified native carboxylesterase were aligned with the deduced amino acid sequence, and the overall identity was 84%. Southern blot analysis suggested the presence of multiple genes. Thus it is concluded that we have cloned a rat liver carboxylesterase, and that this enzyme is a member of a multigene family.

Amino Acid Sequence

Role of light chain variable region in myeloma with light chain deposition disease: evidence from an experimental model.

Light chain deposition disease (LCDD) results from a propensity of some human monoclonal L chains to form tissue deposits. We designed an experimental model for in vivo expression of human kappa L chain sequences in mice and compared a somatically mutated LCDD chain with a closely related control kappa chain, both encoded by the unique V kappa IV gene. Mice secreting the LCDD chain but not those producing the control chain showed deposits with a distribution similar to that observed in patients. These data show that discrete changes in V region sequences can play a major role in tissue deposition of human L chains.

Amino Acid Sequence

Mitochondrial Impostors: Prevalence and Impacts of NUMTs on Genetic and Evolutionary Studies in Carnivora.

Nuclear mitochondrial pseudogenes are mitochondria-derived DNA sequences integrated into the nuclear genome, which can introduce errors in species identification, phylogenetic inference, and population genetics. Although nuclear mitochondrial pseudogene contamination has been reported in some Carnivora species, a systematic investigation into the prevalence and impacts of nuclear mitochondrial pseudogenes across an order is still lacking. In this study, 22,102 mitochondrial DNA sequences of 80 Carnivora species from 14 families and 54 genera were retrieved from the public National Center for Biotechnology Information database and further analyzed. Using alignment-based methods, 158 problematic sequences/sequence groups were identified and categorized into four types: nuclear mitochondrial pseudogenes, species misidentification or mislabeling, sequence errors, and anomalous sites. Among families, Felidae exhibited the highest rate of nuclear mitochondrial pseudogene contamination, particularly in species of the genus Panthera. In contrast, no nuclear mitochondrial pseudogene contamination was detected in members of Ursidae and Ailuridae. Phylogenetic analysis revealed multiple independent origins of nuclear mitochondrial pseudogene, with some tracing back to the common ancestor of Carnivora. To mitigate nuclear mitochondrial pseudogene-related errors, rigorous sequence verification strategies, such as sequence alignment and phylogenetic validation, should be implemented. In conclusion, our findings highlight the necessity of nuclear mitochondrial pseudogene awareness in genetic and evolutionary studies of Carnivora and other taxa.

Animals

High-accuracy SNV calling for bacterial isolates using deep learning with AccuSNV.

Accurate detection of mutations within bacterial species is critical for fundamental studies of microbial evolution, reconstruction of transmission events, and identification of antimicrobial resistance mutations. Although many tools have been developed to identify single-nucleotide variants (SNVs) from whole-genome sequencing, they often suffer from high false-positive rates owing to the complexity of bacterial genomes and the need for different filtering cutoffs across sample types and sequencing depths. As data sets increase in size, the manual filtering required for high accuracy presents a significant obstacle. Here, we present AccuSNV, a novel deep learning-based tool for high-precision and automated bacterial SNV calling. Unlike traditional methods that process one sample at a time, AccuSNV leverages a convolutional neural network (CNN) that integrates alignment information across multiple samples, enhancing precision through learned across-sample patterns. We evaluate AccuSNV against seven popular SNV-calling tools using simulated data from six bacterial species with varied sequencing depths, numbers of isolates, mutations, and divergence levels. To further validate its real-world utility, we test AccuSNV on multiple curated bacterial data sets containing reported SNVs. In both simulated and real-world scenarios, AccuSNV consistently achieves the best performance. Moreover, AccuSNV provides comprehensive user-friendly downstream analysis modules and outputs, including mutation annotation information, phylogenetic inference, d N/d S calculations, and optional manual filtering. Together with the automated deep learning-based calling, these features make AccuSNV broadly accessible to users with different levels of computational expertise.

Deep Learning

A comparison of several similarity indices used in the classification of protein sequences: a multivariate analysis.

The present work describes an attempt to identify reliable criteria which could be used as distance indices between protein sequences. Seven different criteria have been tested: i and ii) the scores of the alignments as given by the BESTFIT and the FASTA programs; iii) the ratio parameter, i.e. the BESTFIT score divided by the length of the aligned peptides; iv and v) the statistical significance (Z-scores) of the scores calculated by BESTFIT and FASTA, as obtained by comparison with shuffled sequences; vi) the Z-scores provided by the program RELATE which performs a segment-by-segment comparison of 2 sequences, and vii) an original distance index calculated by the program DOCMA from all the pairwise dotplots between the sequences. These 7 criteria have been tested against the aminoacid sequences of 39 globins and those of the 20 aminoacyl-tRNA synthetases from E. coli. The distances between the sequences were analyzed by the multivariate analysis techniques. The results show that the distances calculated from the scores of the pairwise alignments are not adequately sensitive. The Z-score from RELATE is not selective enough and too demanding in computer time. Three criteria gave a classification consistent with the known similarities between the sequences in the sets, namely the Z-scores from BESTFIT and FASTA and the multiple dotplot comparison distance index from DOCMA.

Algorithms

[Pattern recognition in the computer analysis of nucleotide sequences].

We have used an algorithm from the pattern recognition theory "generalized portrait" to find a distinguishing vector for Escherichia coli promoters. We have made an attempt to solve closely linked problems for choosing significant signs of that signal, multiple alignment and for calculation of the recognition vector (matrix). The promoters with known strength have been ranged with this vector. The analysis of the occurrence of predicted promoters has been carried out. The promoters search program for IBM-compatible computers is available from the authors.

Base Sequence

Ulysses transposable element of Drosophila shows high structural similarities to functional domains of retroviruses.

We have determined the DNA structure of the Ulysses transposable element of Drosophila virilis and found that this transposon is 10,653 bp and is flanked by two unusually large direct repeats 2136 bp long. Ulysses shows the characteristic organization of LTR-containing retrotransposons, with matrix and capsid protein domains encoded in the first open reading frame. In addition, Ulysses contains protease, reverse transcriptase, RNase H and integrase domains encoded in the second open reading frame. Ulysses lacks a third open reading frame present in some retrotransposons that could encode an env-like protein. A dendrogram analysis based on multiple alignments of the protease, reverse transcriptase, RNase H, integrase and tRNA primer binding site of all known Drosophila LTR-containing retrotransposon sequences establishes a phylogenetic relationship of Ulysses to other retrotransposons and suggests that Ulysses belongs to a new family of this type of elements.

Amino Acid Sequence

Immunoglobulin gene sequence analysis to further assess B-cell origin of multiple myeloma.

To further characterize the B-cell origin of multiple myeloma, our laboratory performed immunoglobulin gene sequence analyses of four cases of myeloma (three immunoglobulin A and one immunoglobulin G). Three tumors expressed VH3 genes and one expressed a VH1 gene, while the light chains included two V lambda and one V kappa III; one light chain was not isolated. The closest homology to published germ line genes ranged from 91 to 97%. In two cases, the expressed VH genes were compared with the putative germ line precursor VH genes isolated from autologous granulocyte DNA and appeared to have mutated randomly from the germ line gene. By sequencing multiple clonal isolates from each tumor sample, we found no evidence for ongoing mutation in three cases; in one case, however, clonotypic heterogeneity was evident. The analysis of DH- and JH-region genes revealed (i) limited or absent N nucleotide insertions (two of four cases), (ii) the presence of a DH-JH junction resulting from sequence overlap between the DH and JH genes (one of four cases), (iii) the absence of somatic mutations (two of four cases), and (iv) restricted JH gene usage of a JH6 polymorphism (three of four cases). These analyses of DH and JH genes suggest that multiple myeloma, similar to what has been proposed for chronic lymphocytic leukemia, may derive from B cells which have rearranged during fetal development rather than during adult life.

Amino Acid Sequence

Alignment-free integration of single-nucleus ATAC-seq across species with sPYce.

Changes in gene regulation largely contribute to differences in cellular identities and phenotypes between species. Single-nucleus assays for transposase-accessible chromatin with sequencing (snATAC-seq) are an efficient strategy to identify putative gene regulatory elements and provide new insight into evolutionary divergence of regulatory programmes. However, no dedicated framework exists to integrate and compare snATAC-seq data across species, while methods designed for single-cell gene expression data have serious limitations. Here we present sPYce, a cross-species snATAC-seq integration method that relies on sequence composition similarities through k-mer histograms of regulatory regions, removing the need for genome alignments to anchor data from different species. sPYce can embed datasets from multiple species into the same mathematical space and permits further downstream analysis steps. We benchmarked sPYce against existing approaches on two publicly available datasets spanning more than 160 myr of evolution, showing that it successfully uncovers conserved cellular programmes while preserving biologically relevant species-specific differences. By comparing cerebellar development in mice and opossums, sPYce identifies regulatory divergence in granule cell differentiation programmes, particularly driven by nuclear factor 1. As an easy-to-use, alignment-free cross-species snATAC-seq integration approach, sPYce opens new perspectives to compare gene regulatory evolution across species.

Animals

Frequency and distribution of DNA uptake signal sequences in the Haemophilus influenzae Rd genome.

The naturally transformable, Gram-negative bacterium Haemophilus influenzae Rd preferentially takes up DNA of its own species by recognizing a 9-base pair sequence, 5'-AAGTGCGGT, carried in multiple copies in its chromosome. With the availability of the complete genome sequence, 1465 copies of the 9-base pair uptake site have been identified. Alignment of these sites unexpectedly reveals an extended consensus region of 29 base pairs containing the core 9-base pair region and two downstream 6-base pair A/T-rich regions, each spaced about one helix turn apart. Seventeen percent of the sites are in inverted repeat pairs, many of which are located downstream to gene termini and are capable of forming stem-loop structures in messenger RNA that might function as signals for transcription termination.

Base Composition

MAFin: motif detection in multiple alignment files.

MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.

Software

Evolutionary relationships among aminotransferases. Tyrosine aminotransferase, histidinol-phosphate aminotransferase, and aspartate aminotransferase are homologous proteins.

A data base was compiled containing the amino acid sequences of 12 aspartate aminotransferases and 11 other aminotransferases. A comparison of these sequences by a standard alignment method confirmed the previously reported homology of all aspartate aminotransferases and Escherichia coli tyrosine aminotransferase. However, no significant similarity between these proteins and any of the other aminotransferases was detected. A more rigorous analysis, focusing on short sequence segments rather than the total polypeptide chain, revealed that rat tyrosine aminotransferase and Saccharomyces cerevisiae and Escherichia coli histidinol-phosphate aminotransferase share several homologous sequence segments with aspartate aminotransferases. For comparison of the complete sequences, a multiple sequence editor was developed to display the whole set of amino acid sequences in parallel on a single work-sheet. The editor allows gaps in individual sequences or a set of sequences to be introduced and thus facilitates their parallel analysis and alignment. Several clusters of invariant residues at corresponding positions in the amino acid sequences became evident, clearly establishing that the cytosolic and the mitochondrial isoenzyme of vertebrate aspartate aminotransferase, E. coli aspartate aminotransferase, rat and E. coli tyrosine aminotransferase, and S. cerevisiae and E. coli histidinol-phosphate aminotransferase are homologous proteins. Only 12 amino acid residues out of a total of about 400 proved to be invariant in all sequences compared; they are either involved in the binding of pyridoxal 5'-phosphate and the substrate, or appear to be essential for the conformation of the enzymes.

Amino Acid Sequence

Heterogeneity of T-cell receptor alpha-chain complementarity-determining region 3 in myelin basic protein-specific T cells increases with severity of multiple sclerosis.

The pathogenesis of multiple sclerosis (MS) is thought to involve a T-cell-mediated autoimmune process. Experimental allergic encephalomyelitis (EAE), an animal model resembling MS, can be induced by immunization with myelin antigens such as myelin basic protein. The T-cell antigen receptor (TCR) usage in EAE is highly restricted in some strains of animals and experimental treatments targeting the TCR have been successful in EAE. Examination of the TCR beta-chain variable-region (V beta) usage of MBP-specific T-cell lines in MS patients has produced conflicting results. Our previous studies of TCR alpha-chain variable-region usage in monozygotic twins demonstrated a general skewing of the TCR repertoire in individuals with MS. This skewing became apparent only after stimulation with antigens; in peripheral blood lymphocyte preparations from individuals with MS V alpha 8-bearing T cells were preferentially selected by stimulation with myelin basic protein. In the present study we examined complementarity-determining region 3 of those V alpha 8-positive TCRs. Marked sequence heterogeneity was found in all individuals with severe MS. In contrast, restricted areas of complementarity-determining region 3 were found in healthy control individuals and in individuals with a mild form of MS. Sequences from tetanus toxoid-specific V alpha 8-positive T cells generated from the same individuals were relatively homogeneous within individuals regardless of disease activity and were distinct from the sequences of complementarity-determining region 3 in myelin basic protein-stimulated lines. These findings suggest that disease severity may be associated with increased heterogeneity of myelin antigen-specific T cells and could reflect an impaired ability of the immune system to down-regulate these anti-self responses.

Adult

Unique T-cell receptor junctional sequences found in multiple sclerosis and T-cells mediating experimental allergic encephalomyelitis.

We have used two approaches to isolate TCR sequences that are unique to patients with multiple sclerosis. One strategy was to sequence TCR gene rearrangements directly from MS lesions. The second strategy utilized T-cell clones with a selectable mutation that are found only in MS patients. The selection of T-cell clones with mutations in the hypoxanthine guanine phosphoribosyltransferase (hprt) gene was used to isolate T-cells reactive to myelin basic protein (MBP) in patients with multiple sclerosis (MS). These T-cell clones are activated in vivo, and are not found in healthy individuals. The third complementarity determining regions (CDR3) of the T-cell receptor (TCR) alpha and beta chains are the putative contact sites for peptide fragments of MBP bound in the groove of the HLA molecule. The TCR V gene usage and CDR3s of these MBP-reactive hprt- T-cell clones are homologous to TCRs from other T-cells relevant to MS, including T-cells causing experimental allergic encephalomyelitis (EAE) and T-cells found in brain lesions and in the cerebrospinal fluid (CSF) of MS patients. In vivo activated MBP-reactive T-cells in MS patients may be critical in the pathogenesis of MS.

Amino Acid Sequence

Detection of Caenorhabditis transposon homologs in diverse organisms.

Although transposons that move via DNA intermediates are common in bacteria, invertebrates, and plants, none have been clearly documented in vertebrates and certain other classes of organisms. One such family of transposons includes invertebrate elements related to Caenorhabditis elegans Tc1. Blocks of aligned protein segments derived from this family were used to search a nucleotide sequence databank. Among the relatives detected were known bacterial insertion elements, revealing the ancient origin of the family. Furthermore, a Tc1-like homolog was detected in a catfish, raising the possibility that this valuable tool of C. elegans genetics can be used with vertebrate genomes. This study illustrates the use of multiple protein blocks for detection and evaluation of distant relationships.

Amino Acid Sequence

Recognition of distantly related protein sequences using conserved motifs and neural networks.

A sensitive technique for protein sequence motif recognition based on neural networks has been developed. It involves three major steps. (1) At each appropriate alignment position of a set of N matched sequences, a set of N aligned oligopeptides is specified with preselected window length. N neural nets are subsequently and successively trained on N-1 amino acid spans after eliminating each ith oligopeptide. A test for recognition of each of the ith spans is performed. The average neural net recognition over N such trials is used as a measure of conservation for the particular windowed region of the multiple alignment. This process is repeated for all possible spans of given length in the multiple alignment. (2) The M most conserved regions are regarded as motifs and the oligopeptides within each are used to train intensively M individual neural networks. (3) The M networks are then applied in a search for related primary structures in a databank of known protein sequences. The oligopeptide spans in the database sequence with strongest neural net output for each of the M networks are saved and then scored according to the output signals and the proper combination that follows the expected N- to C-terminal sequence order. The motifs from the database with highest similarity scores can then be used to retrain the M neural nets, which can be subsequently utilized for further searches in the databank, thus providing even greater sensitivity to recognize distant familial proteins. This technique was successfully applied to the integrase, DNA-polymerase and immunoglobulin families.

Aldehyde Dehydrogenase

A sequence property approach to searching protein databases.

Currently available sequence alignment programs are generally not capable of detecting functional and structural homologs in the twilight zone of sequence similarity, i.e. when the sequence identity falls below about 25%. Here we attempt to detect such weak similarities using an approach based on a notion of protein sequence similarity radically different from that used in sequential alignment. The approach defines protein sequence dissimilarity (or distance) as a weighted sum of differences of compositional properties such as singlet and doublet amino acid composition, molecular weight, isoelectric point (protein property search or PropSearch). With PropSearch, either single sequences can be used for a database query, or multiple sequences can be merged into an "average" sequence reflecting the average composition of a protein family. First, we show that members of structural protein families have a low mutual PropSearch distance when the weights are optimized to discriminate maximally between structural families. Second, we demonstrate the results of database searches using the PropSearch method. Such searches are very rapid when scanning a preprocessed database and do not require alignments. In cases in which conventional alignment tools fail to detect similarities, PropSearch can be used to generate hypotheses about possible structural or functional relationships between a new sequence and sequences in the database.

Algorithms

Structure-function analysis of human IL-6: identification of two distinct regions that are important for receptor binding.

Interleukin-6 (IL-6) is a multifunctional cytokine that plays an important role in host defense. It has been predicted that IL-6 may fold as a 4 alpha-helix bundle structure with up-up-down-down topology. Despite a high degree of sequence similarity (42%) the human and mouse IL-6 polypeptides display distinct species-specific activities. Although human IL-6 (hIL-6) is active in both human and mouse cell assays, mouse IL-6 (mIL-6) is not active on human cells. Previously, we demonstrated that the 5 C-terminal residues of mIL-6 are important for activity, conformation, and stability (Ward LD et al., 1993, Protein Sci 2:1472-1481). To further probe the structure-function relationship of this cytokine, we have constructed several human/mouse IL-6 hybrid molecules. Restriction endonuclease sites were introduced and used to ligate the human and mouse sequences at junction points situated at Leu-62 (Lys-65 in mIL-6) in the putative connecting loop AB between helices A and B, at Arg-113 (Val-117 in mIL-6) at the N-terminal end of helix C, at Lys-150 (Asp-152 in mIL-6) in the connecting loop CD between helices C and D, and at Leu-178 (Thr-180 in mIL-6) in helix D. Hybrid molecules consisting of various combinations of these fragments were constructed, expressed, and purified to homogeneity. The conformational integrity of the IL-6 hybrids was assessed by far-UV CD. Analysis of their biological activity in a human bioassay (using the HepG2 cell line), a mouse bioassay (using the 7TD1 cell line), and receptor binding properties indicates that at least 2 regions of hIL-6, residues 178-184 in helix D and residues 63-113 in the region incorporating part of the putative connecting loop AB through to the beginning of helix C, are critical for efficient binding to the human IL-6 receptor. For human IL-6, it would appear that interactions between residues Ala-180, Leu-181, and Met-184 and residues in the N-terminal region may be critical for maintaining the structure of the molecule; replacement of these residues with the corresponding 3 residues in mouse IL-6 correlated with a significant loss of alpha-helical content and a 200-fold reduction in activity in the mouse bioassay. A homology model of mIL-6 based on the X-ray structure of human granulocyte colony-stimulating factor is presented.

Amino Acid Sequence