PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Identification of a functionally conserved surface region of rat cytochromes P450IA.

A region of rat cytochrome P450IA1 at residues 294-301 (Gln-Asp-Arg-Arg-Leu-Asp-Glu-Asn), equivalent to a proinhibitory region of cytochrome P450IA2, was identified by sequence alignment. Anti-peptide antibodies were successfully raised when the peptide was coupled through either its N- or its C-terminus to carrier protein, but no antibodies were produced against the so-called multiple peptide antigen, which consisted of eight copies of the peptide attached through its C-terminus to a synthetic base. Both of the anti-peptide antibodies bound specifically to cytochrome P450IA1 in the rat, as shown by e.l.i.s.a. and immunoblotting. They inhibited microsomal aryl hydrocarbon hydroxylase activity and the mutagenic activation of 2-acetylaminofluorene (these reactions are catalysed by cytochrome P450IA1), but not high-affinity phenacetin O-de-ethylation activity, which is catalysed by cytochrome P450IA2. However, there was differences in the properties of the two antisera in their binding to cytochromes P450IA1 in species other than the rat, their relative binding to the multiple peptide antigen, the yield of antibody following affinity purification using peptide coupled through its N-terminus to CNBr-activated Sepharose, and the binding of the purified preparations to N- and C-terminal-coupled peptide conjugates. These observations indicated that the antibodies were directed to the region of the peptide opposite to the end which was coupled to the carrier protein. Nevertheless, both of the antibody preparations bound equally well to the target cytochrome P450, thus indicating that, in the native protein, the whole of the peptide region is exposed on the surface of cytochrome P450IA1 and is available for binding by the antibodies. The role of this region appears to be the same in both cytochromes P450IA1 and P450IA2, despite the difference in its primary structure in the two cytochromes P450.

Amino Acid Sequence

Prokaryotic and eukaryotic pyridoxal-dependent decarboxylases are homologous.

A database search has revealed significant and extensive sequence similarities among prokaryotic and eukaryotic pyridoxal phosphate (PLP)-dependent decarboxylases, including Drosophila glutamic acid decarboxylase (GAD) and bacterial histidine decarboxylase (HDC). Based on these findings, the sequences of seven PLP-dependent decarboxylases from five different organisms have been aligned to derive a consensus sequence for this family of enzymes. In addition, quantitative methods have been employed to calculate the relative evolutionary distances between pairs of the decarboxylases comprising this family. The multiple sequence analysis together with the quantitative results strongly suggest an ancient and common origin for all PLP-dependent decarboxylases. This analysis also indicates that prokaryotic and eukaryotic HDC activities evolved independently. Finally, a sensitive search algorithm (PROFILE) was unable to detect additional members of this decarboxylase family in protein sequence databases.

Amino Acid Sequence

The product of unr, the highly conserved gene upstream of N-ras, contains multiple repeats similar to the cold-shock domain (CSD), a putative DNA-binding motif.

We show that the open reading frame transcribed from the unr gene (immediately upstream of N-ras) in mammals consists of multiple repeats similar to the cold-shock domain (CSD), a putative DNA-binding motif found in prokaryotic cold-shock proteins, and eukaryotic DNA-binding proteins. Alignment of the CSD sequences of unr with those from other proteins reveals a core of similarity for which a consistent secondary structure prediction can be derived. This prediction suggests that the CSD consists primarily of beta-sheet, in contrast to most known eukaryotic DNA-binding proteins. Sequence analysis of the 3' end of the guinea pig unr gene shows that the core of one CSD repeat is encoded in a single exon, consistent with the modular assembly of the gene from ancestral CSD-coding units.

Amino Acid Sequence

Identification of proteolytic cleavage sites in the conversion of profilaggrin to filaggrin in mammalian epidermis.

Profilaggrin consists of multiple filaggrin domains joined by linker segments which are removed during proteolytic conversion to filaggrin. Analysis of tryptic peptides of filaggrin defined a 26-residue linker segment when aligned on the amino acid sequence of one repeat unit of mouse profilaggrin deduced from a cDNA sequence (Rothnagel, J. A., Mehrel, T., Idler, W. W., Roop, D. R., and Steinert, P. M. (1987) J. Biol. Chem. 262, 15643-15648). Two types of linker segments were distinguished by their different susceptibility to thermolysin and by the presence of a Phe-Tyr-Pro-Val sequence in only one type. These data led to a model of profilaggrin in which the two types of linker segments alternate along the length of profilaggrin. This model provides a structural basis for the two stages of proteolytic processing seen in vivo. In the first stage intermediates accumulate which have several filaggrin domains still joined by linker segments lacking Phe-Tyr-Pro-Val. In the second stage, the other linker segments are cleaved and mature filaggrin domains are released. Proteolytic activity with specificity consistent with first stage cleavage was partially purified from rat epidermis. Chymostatin inhibited both the in vitro enzymatic activity and the processing of profilaggrin in a cultured rat keratinocyte cell line. The products formed in vitro were 3-5 kDa larger than intermediates produced in vivo, suggesting that the linker segments are cleaved at one end only. This implies the existence of a third protease which completes the removal of the linker segments.

Amino Acid Sequence

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence

Is the bacteriophage lambda lysozyme an evolutionary link or a hybrid between the C and V-type lysozymes? Homology analysis and detection of the catalytic amino acid residues.

The relationship between the bacteriophage lambda lysozyme (lambda L) and the C and V-type lysozymes has been investigated by sequence alignment, secondary structure prediction and pattern recognition methods. The alignment of the amino terminal part of lambda L with that of V-type lysozymes suggests that Glu19 is a residue essential for catalysis. Its mutation to Gln leads to a completely inactive enzyme. In the alignment of the sequence of lambda L with those of the C-type lysozymes a strongly homologous fragment of about 30 amino acid residues is detected. Taking into consideration this observation and the published structural alignments between C and V-type lysozymes, a repetition of the beta-sheet motif in lambda L is proposed. The multiple alignment draws the attention to a possible catalytic role for Asp34 that would be positioned in the middle of the second strand of the beta-sheet as in the C-type lysozymes. This role is confirmed by mutagenesis. The implications of these observations in terms of the evolutionary relationship between lambda L and the other lysozymes is discussed.

Amino Acid Sequence

An end-to-end computational framework for "Record-seq" transcriptional recording data.

MOTIVATION: Record-seq captures cumulative transcriptional activity over time in engineered Escherichia coli by integrating cellular RNA-derived spacer sequences into clustered regularly interspaced short palindromic repeats (CRISPR) arrays, which are read out by sequencing. Unlike the approximately uniform transcript sampling of RNA-seq, Record-seq records biological signal as spacers sampled by the CRISPR spacer acquisition machinery. Consequently, standard RNA-seq analysis strategies are not directly applicable, limiting sensitivity and interpretability. Our previous pipeline addressed these challenges only partially, retained inherited RNA-seq assumptions, and had limited algorithmic efficiency. RESULTS: Here, we present an end-to-end computational framework for Record-seq data. To address the primary computational bottleneck of spacer sequence extraction, we implemented a wavefront alignment approach for efficient quasi-local pattern matching, achieving an approximately 30-fold speedup. We introduce transcription unit-based feature counting as an alternative to gene-body quantification to better represent prokaryotic transcription and increase statistical power by capturing signal from untranslated regions, which are spacer acquisition hotspots. For downstream analyses, we incorporate multiple normalization strategies and a nonparametric differential expression testing framework designed for sparse datasets. Further, we analyze spacer acquisition patterns and train sequence-based neural models that predict acquisition propensity from genomic sequence and annotations, providing a framework for assessing whether acquisition rules generalize as Record-seq is extended to new microbial hosts. AVAILABILITY AND IMPLEMENTATION: The primary analysis workflow, the recoRdseq package, acquisition modeling repository, and relevant data are all linked at https://github.com/plattlab/Record-seq-Framework. Acquisition models and training data are on Zenodo at https://doi.org/10.5281/zenodo.18891434.

Escherichia coli

The Drosophila ras oncogenes: structure and nucleotide sequence.

Three Drosophila genes homologous to the Ha-ras probe were isolated and mapped to positions 85D, 64B, and 62B on chromosome 3. Two of these genes (termed Dras 1 and Dras 2) were sequenced. In the case of Dras 1, which contains multiple introns, a cDNA clone was isolated and sequenced. In the case of Dras2, the nucleotide sequence fo the genomic clone was determined. Each gene codes for a protein with a predicted molecular weight of 21.6 kd. Alignment of the amino acid sequence of Dras 1 with the vertebrate Ha-ras protein shows that at the amino terminus and central portion (residues 1-121 and 137-164) the two proteins are remarkably similar, and have an overall homology of 75%. The Dras 2 gene lacks significant homology to the vertebrate counterpart at the extreme amino terminus and is homologous only between positions 28-120 and 139-161 (overall homology of 50%). This result suggests that the N terminus of p21 forms a distinct regulatory or functional domain. At the carboxy terminus, the major region of variability among the vertebrate ras proteins, the two Drosophila sequences also display considerable variability. However, both appear to be more similar to exon 4B of the Ki-ras gene.

Amino Acid Sequence

Evolutionary relationships of avian Eimeria species among other Apicomplexan protozoa: monophyly of the apicomplexa is supported.

Direct, reverse transcriptase-mediated, partial sequencing of the small-subunit (16S-like) ribosomal RNA (srRNA) of Eimeria tenella and E. acervulina was performed. Sequences were aligned by eye with six previously published, partial or complete srRNA sequences of apicomplexan protists (Plasmodium berghei, Theileria annulata, Cryptosporidium sp., Toxoplasma gondii, Sarcocystis muris, and S. gigantea). Six eukaryotic protists (a slime mold, a yeast, two dinoflagellates, and two ciliates) acted as an outgroup for a parsimony-based phylogenetic analysis (PAUP Ver. 3.0). The 188 phylogenetically informative sites (i.e., those positions that neither were unvaried nor had only autapomorphic substitutions) supported a single tree topology 481 steps in length with a consistency index of 0.65 in which the monophyly of the Apicomplexa was supported. The two Eimeria species and S. muris, S. gigantea, and T. gondii formed a pair of monophyletic groups that were sister groups. The two Sarcocystis species were not hypothesized to be sister taxa. The genera Plasmodium and Cryptosporidium were hypothesized to form the sister group to these five coccidia and T. annulata. A priori data-editing techniques that deleted "variable" positions prior to analysis failed to recognize the monophyly of the Apicomplexa when the same parsimony-based tree-building algorithm was used. Inability of the outgroup taxa to root the well-supported ingroup tree (Apicomplexa) at a unique site when these taxa were used individually for this purpose reinforces the need for an appropriate, multiple-taxon outgroup in such analyses.

Animals

European elk papillomavirus: characterization of the genome, induction of tumors in animals, and transformation in vitro.

The European elk papillomavirus (EEPV) genome was cloned in the BamHI cleavage site of the pBR322 vector. The cloned genome was used for construction of a physical map, employing restriction endonucleases BamHI, BglII, HindIII, PvuII, SacI, and XhoI. The sequence homology between the EEPV and bovine papillomavirus type 1 genomes was elucidated by performing hybridizations in different concentrations of formamide. Sequence homology could only be revealed under less stringent conditions, i.e., Tm - 43 degrees C. Nucleotide sequence information was also collected from the regions which lie adjacent to the three HindIII sites that are present in the EEPV genome. The results made it possible to align the EEPV and bovine papillomavirus type 1 genomes. Transformation by EEPV was demonstrated with the C127 mouse cell line, and fibrosarcomas were induced in young hamsters after subcutaneous injection. The transformed cells and the tumors contain multiple, nonintegrated copies of the EEPV genome. Virus particles could not be detected either in tumors or in transformed cells.

Animals

resLens: genomic language models to enhance antibiotic resistance gene detection.

The rise of antibiotic resistance necessitates advanced tools to detect and analyze antibiotic resistance genes (ARGs). We present resLens, a family of genomic language models that leverage latent genomic representations to enhance ARG detection and analysis. Unlike alignment-based methods constrained by reference databases, resLens fine-tunes a pre-trained DNA language model on curated ARG datasets, achieving competitive or superior performance in classifying resistance genes across multiple evaluation scenarios, including when ARGs exhibit sequences and mechanisms of resistance dissimilar to those in reference datasets.

Journal Article

Primary structure determination of two cytochromes c2: close similarity to functionally unrelated mitochondrial cytochrome C.

The amino-acid sequences of the cytochromes c2 from the photosynthetic non-sulfur purple bacteria Rhodomicrobium vannielii and Rhodopseudomonas viridis have been determined. Only a single residue deletion (at position 11 in horse cytochrome c) is necessary to align the sequences with those of mitochondrial cytochromes c. The overall sequence similarity between these cytochromes c2 and mitochondrial cytochromes c is closer than that between mitochondrial cytochromes c and the other cytochromes c2 of known sequence, and in the latter multiple insertions and deletions must be postulated before a match can be obtained. Nevertheless, these two cytochromes c2 show no better reactivity with the mitochondrial cytochrome c oxidase than do the less well-matched cytochromes c2. The bearing of these findings on possible evolutionary relationship between mitochondria and prokaryotes is discussed.

Amino Acid Sequence

Topography of simian virus 40 A protein-DNA complexes: arrangement of protein bound to the origin of replication.

DNA binding regions I, II, and III at the origin of replication have different arrangements of A protein (T antigen) recognition pentanucleotides. The A protein also protects each region from DNase in distinctly different patterns. Footprint and fragment assays led to the following conclusions: (i) in some cases a single recognition pentanucleotide is sufficient to direct the binding and accurate alignment of A protein on DNA; (ii) the A protein binds within isolated region I or II in a sequential process leading to multiple overlapping areas of DNase protection within each region; and (iii) the 23-base pair span of recognition sequences in region II allows binding and protection of a longer length of DNA than the 23-base pair span in region I. We propose a model of protein binding that addresses the problem of variations in the arrangement of pentanucleotides in regions I and II and explains the observed DNase protection patterns. The central feature of the model requires each protomer of A protein to bind to a pentanucleotide in a unique direction. The resulting orientation of protein would protect more DNA at the 5' end of the 5'-GAGGC-3' recognition sequence than at the 3' end. The arrangement of multiple protomers at the origin of simian virus 40 replication is discussed.

Antigens, Viral

Actin as the generator of tension during muscle contraction.

We propose that the key structural feature in the conversion of chemical free energy into mechanical work by actomyosin is a myosin-induced change in the length of the actin filament. As reported earlier, there is evidence that helical actin filaments can untwist into ribbons having an increased intersubunit repeat. Regular patterns of actomyosin interactions arise when ribbons are aligned with myosin thick filaments, because the repeat distance of the myosin lattice (429 A) is an integral multiple of the subunit repeat in the ribbon (35.7 A). This commensurability property of the actomyosin lattice leads to a simple mechanism for controlling the sequence of events in chemical-mechanical transduction. A role for tropomyosin in transmitting the forces developed by actomyosin is proposed. In this paper, we describe how these transduction principles provide the basis for a theory of muscle contraction.

Actins

The amino acid sequence of the light chain of human high-molecular-mass kininogen.

The complete amino acid sequence of the light chain of human high-molecular-mass kininogen has been determined. The peptide chain contains 255 amino acid residues. The half-cystine, which forms the disulfide bridge to the heavy chain, was identified in position 225. Nine carbohydrate attachment sites were found. All carbohydrate side chains are O-glycosidically linked. Alignment of the present sequence with the bovine kininogen light chain sequence shows a high degree of homology, except for an extension of 22 amino acids within the histidine-rich part of the sequence. The histidine-rich region may have arisen by gene multiplication during evolution.

Amino Acid Sequence

Duplication of type IV collagen COOH-terminal repeats and species-specific expression of alpha 1(IV) and alpha 2(IV) collagen genes.

DNA sequencing of a 1.7-kilobase cloned cDNA allowed determination of the complete COOH-terminal noncollagenous region (NC1) of the human alpha 2(IV) collagen chain. This 227-residue domain is composed of two equal-sized sequential "repeats" as had been observed for the corresponding 229-residue domain of the alpha 1(IV) chain. Alignment of the alpha 2(IV) repeats with each other and with those in alpha 1(IV) suggests that the type IV NC1 regions evolved via an intra- to intergenic duplication. In addition, smaller internal units having a consensus sequence GYSSCLFLWYLAF are present multiple times in the 110-115-amino acid halves. Consonant with the predominance of Tyr, Leu, and Phe in the heptads, hydrophilicity profiles of the homologous alpha 1(IV) and alpha 2(IV) regions revealed extended hydrophobic stretches which are similar to those in the noncollagenous COOH terminus of the chicken alpha 1(X) chain (Ninomiya, Y., Gordon, M., van der Rest, M., Schmid, T., Linsenmayer, T., and Olsen, B. R. (1986) J. Biol. Chem. 261, 5041-5050). Isolation of an alpha 2(IV) clone also enabled us to investigate if the alpha 2(IV) and alpha 1(IV) collagen mRNAs were coordinately transcribed and if one or both were consistently associated with either type I (alpha 1 and alpha 2), III, or V (alpha 2) transcripts as a function of cell type. Northern blot hybridization of collagen cDNA probes to poly(A) RNA extracted from human and bovine cell cultures showed that only the alpha 1(III) and alpha 2(V) genes were expressed in all cells examined. Unexpectedly, neither type IV mRNAs were found in bovine endothelial, smooth muscle, or fibroblast cells, whereas both type IV species were present in human fibroblasts and to a much greater extent in human umbilical vein endothelial cells.

Amino Acid Sequence

Detection of weak sequence homology of proteins for tertiary structure prediction.

Multiple measures of similarity were employed to detect weak homologies among protein sequences (e.g., below 30% residue identity). A set of thresholds was empirically determined, by using sample proteins of known structure, so as to select only correct pairs of sequences; correct or incorrect alignment of sequences was judged by direct comparison of corresponding conformations. The empirical criterion thus set up is applicable to the prediction of a protein structure when the structure of the other protein in the pair is known. We searched all the combinations between 84 proteins of known structure and 4610 proteins stored in a sequence database, and found about 4000 pairs of sequences which satisfied the criterion. However, after excluding such pairs of proteins that belong to the same family or superfamily, the number of pairs remaining was reduced to only 19. The reliability of these data for structural prediction is discussed.

Amino Acid Sequence

RNA-catalysed synthesis of complementary-strand RNA.

The Tetrahymena ribozyme can splice together multiple oligonucleotides aligned on a template strand to yield a fully complementary product strand. This reaction demonstrates the feasibility of RNA-catalysed RNA replications.

Animals