PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

resLens: genomic language models to enhance antibiotic resistance gene detection.

The rise of antibiotic resistance necessitates advanced tools to detect and analyze antibiotic resistance genes (ARGs). We present resLens, a family of genomic language models that leverage latent genomic representations to enhance ARG detection and analysis. Unlike alignment-based methods constrained by reference databases, resLens fine-tunes a pre-trained DNA language model on curated ARG datasets, achieving competitive or superior performance in classifying resistance genes across multiple evaluation scenarios, including when ARGs exhibit sequences and mechanisms of resistance dissimilar to those in reference datasets.

Journal Article

Primary structure determination of two cytochromes c2: close similarity to functionally unrelated mitochondrial cytochrome C.

The amino-acid sequences of the cytochromes c2 from the photosynthetic non-sulfur purple bacteria Rhodomicrobium vannielii and Rhodopseudomonas viridis have been determined. Only a single residue deletion (at position 11 in horse cytochrome c) is necessary to align the sequences with those of mitochondrial cytochromes c. The overall sequence similarity between these cytochromes c2 and mitochondrial cytochromes c is closer than that between mitochondrial cytochromes c and the other cytochromes c2 of known sequence, and in the latter multiple insertions and deletions must be postulated before a match can be obtained. Nevertheless, these two cytochromes c2 show no better reactivity with the mitochondrial cytochrome c oxidase than do the less well-matched cytochromes c2. The bearing of these findings on possible evolutionary relationship between mitochondria and prokaryotes is discussed.

Amino Acid Sequence

Topography of simian virus 40 A protein-DNA complexes: arrangement of protein bound to the origin of replication.

DNA binding regions I, II, and III at the origin of replication have different arrangements of A protein (T antigen) recognition pentanucleotides. The A protein also protects each region from DNase in distinctly different patterns. Footprint and fragment assays led to the following conclusions: (i) in some cases a single recognition pentanucleotide is sufficient to direct the binding and accurate alignment of A protein on DNA; (ii) the A protein binds within isolated region I or II in a sequential process leading to multiple overlapping areas of DNase protection within each region; and (iii) the 23-base pair span of recognition sequences in region II allows binding and protection of a longer length of DNA than the 23-base pair span in region I. We propose a model of protein binding that addresses the problem of variations in the arrangement of pentanucleotides in regions I and II and explains the observed DNase protection patterns. The central feature of the model requires each protomer of A protein to bind to a pentanucleotide in a unique direction. The resulting orientation of protein would protect more DNA at the 5' end of the 5'-GAGGC-3' recognition sequence than at the 3' end. The arrangement of multiple protomers at the origin of simian virus 40 replication is discussed.

Antigens, Viral

Actin as the generator of tension during muscle contraction.

We propose that the key structural feature in the conversion of chemical free energy into mechanical work by actomyosin is a myosin-induced change in the length of the actin filament. As reported earlier, there is evidence that helical actin filaments can untwist into ribbons having an increased intersubunit repeat. Regular patterns of actomyosin interactions arise when ribbons are aligned with myosin thick filaments, because the repeat distance of the myosin lattice (429 A) is an integral multiple of the subunit repeat in the ribbon (35.7 A). This commensurability property of the actomyosin lattice leads to a simple mechanism for controlling the sequence of events in chemical-mechanical transduction. A role for tropomyosin in transmitting the forces developed by actomyosin is proposed. In this paper, we describe how these transduction principles provide the basis for a theory of muscle contraction.

Actins

The amino acid sequence of the light chain of human high-molecular-mass kininogen.

The complete amino acid sequence of the light chain of human high-molecular-mass kininogen has been determined. The peptide chain contains 255 amino acid residues. The half-cystine, which forms the disulfide bridge to the heavy chain, was identified in position 225. Nine carbohydrate attachment sites were found. All carbohydrate side chains are O-glycosidically linked. Alignment of the present sequence with the bovine kininogen light chain sequence shows a high degree of homology, except for an extension of 22 amino acids within the histidine-rich part of the sequence. The histidine-rich region may have arisen by gene multiplication during evolution.

Amino Acid Sequence

Duplication of type IV collagen COOH-terminal repeats and species-specific expression of alpha 1(IV) and alpha 2(IV) collagen genes.

DNA sequencing of a 1.7-kilobase cloned cDNA allowed determination of the complete COOH-terminal noncollagenous region (NC1) of the human alpha 2(IV) collagen chain. This 227-residue domain is composed of two equal-sized sequential "repeats" as had been observed for the corresponding 229-residue domain of the alpha 1(IV) chain. Alignment of the alpha 2(IV) repeats with each other and with those in alpha 1(IV) suggests that the type IV NC1 regions evolved via an intra- to intergenic duplication. In addition, smaller internal units having a consensus sequence GYSSCLFLWYLAF are present multiple times in the 110-115-amino acid halves. Consonant with the predominance of Tyr, Leu, and Phe in the heptads, hydrophilicity profiles of the homologous alpha 1(IV) and alpha 2(IV) regions revealed extended hydrophobic stretches which are similar to those in the noncollagenous COOH terminus of the chicken alpha 1(X) chain (Ninomiya, Y., Gordon, M., van der Rest, M., Schmid, T., Linsenmayer, T., and Olsen, B. R. (1986) J. Biol. Chem. 261, 5041-5050). Isolation of an alpha 2(IV) clone also enabled us to investigate if the alpha 2(IV) and alpha 1(IV) collagen mRNAs were coordinately transcribed and if one or both were consistently associated with either type I (alpha 1 and alpha 2), III, or V (alpha 2) transcripts as a function of cell type. Northern blot hybridization of collagen cDNA probes to poly(A) RNA extracted from human and bovine cell cultures showed that only the alpha 1(III) and alpha 2(V) genes were expressed in all cells examined. Unexpectedly, neither type IV mRNAs were found in bovine endothelial, smooth muscle, or fibroblast cells, whereas both type IV species were present in human fibroblasts and to a much greater extent in human umbilical vein endothelial cells.

Amino Acid Sequence

Detection of weak sequence homology of proteins for tertiary structure prediction.

Multiple measures of similarity were employed to detect weak homologies among protein sequences (e.g., below 30% residue identity). A set of thresholds was empirically determined, by using sample proteins of known structure, so as to select only correct pairs of sequences; correct or incorrect alignment of sequences was judged by direct comparison of corresponding conformations. The empirical criterion thus set up is applicable to the prediction of a protein structure when the structure of the other protein in the pair is known. We searched all the combinations between 84 proteins of known structure and 4610 proteins stored in a sequence database, and found about 4000 pairs of sequences which satisfied the criterion. However, after excluding such pairs of proteins that belong to the same family or superfamily, the number of pairs remaining was reduced to only 19. The reliability of these data for structural prediction is discussed.

Amino Acid Sequence

RNA-catalysed synthesis of complementary-strand RNA.

The Tetrahymena ribozyme can splice together multiple oligonucleotides aligned on a template strand to yield a fully complementary product strand. This reaction demonstrates the feasibility of RNA-catalysed RNA replications.

Animals

A new algorithm for best subsequence alignments with application to tRNA-rRNA comparisons.

The algorithm of Smith & Waterman for identification of maximally similar subsequences is extended to allow identification of all non-intersecting similar subsequences with similarity score at or above some preset level. The resulting alignments are found in order of score, with the highest scoring alignment first. In the case of single gaps or multiple gaps weighted linear with gap length, the algorithm is extremely efficient, taking very little time beyond that of the initial calculation of the matrix. The algorithm is applied to comparisons of tRNA-rRNA sequences from Escherichia coli. A statistical analysis is important for proper evaluation of the results, which differ substantially from the results of an earlier analysis of the same sequences by Bloch and colleagues.

Algorithms

Basic local alignment search tool.

A new approach to rapid sequence comparison, basic local alignment search tool (BLAST), directly approximates alignments that optimize a measure of local similarity, the maximal segment pair (MSP) score. Recent mathematical results on the stochastic properties of MSP scores allow an analysis of the performance of this method as well as the statistical significance of alignments it generates. The basic algorithm is simple and robust; it can be implemented in a number of ways and applied in a variety of contexts including straightforward DNA and protein sequence database searches, motif searches, gene identification searches, and in the analysis of multiple regions of similarity in long DNA sequences. In addition to its flexibility and tractability to mathematical analysis, BLAST is an order of magnitude faster than existing sequence comparison tools of comparable sensitivity.

Algorithms

Extensive multiplicity of the miscellaneous type of neurotoxins from the venom of the cobra Naja naja naja and structural characterization of major components.

A multiplicity of miscellaneous type neurotoxins were detected in the venom of the cobra Naja naja naja by use of reverse-phase HPLC and FPLC. The primary structures of major forms were determined, giving 4 novel structures. All four contain 62-65 residues, with 10 half-cystine residues and resemble the miscellaneous type of toxins from other Naja species. Differences within the species are extensive, exchanges occur at 27 positions, giving only 58% residue identity between all forms. However, the differences are largely limited to 3 regions corresponding to structurally important loops where two functional residues participating in receptor binding are exchanged. The four miscellaneous neurotoxins now characterized, together with the minor components of the miscellaneous type, the minimally four neurotoxins reported before, and other related toxins, indicate the existence of an extensive toxin gene multiplicity.

Amino Acid Sequence

Characteristics of short-chain alcohol dehydrogenases and related enzymes.

Different short-chain dehydrogenases are distantly related, constituting a protein family now known from at least 20 separate enzymes characterized, but with extensive differences, especially in the C-terminal third of their sequences. Many of the first known members were prokaryotic, but recent additions include mammalian enzymes from placenta, liver and other tissues, including 15-hydroxyprostaglandin, 17 beta-hydroxysteroid and 11 beta-hydroxysteroid dehydrogenases. In addition, species variants, isozyme-like multiplicities and mutants have been reported for several of the structures. Alignments of the different enzymes reveal large homologous parts, with clustered similarities indicating regions of special functional/structural importance. Several of these derive from relationships within a common type of coenzyme-binding domain, but central-chain patterns of similarity go beyond this domain. Total residue identities between enzyme pairs are typically around 25%, but single forms deviate more or less (14-58%). Only six of the 250-odd residues are strictly conserved and seven more are conserved in all but single cases. Over one third of the conserved residues are glycine, showing the importance of conformational and spatial restrictions. Secondary structure predictions, residue distributions and hydrophilicity profiles outline a common, N-terminal coenzyme-binding domain similar to that of other dehydrogenases, and a C-terminal domain with unique segments and presumably individual functions in each case. Strictly conserved residues of possible functional interest are limited, essentially only three polar residues. Asp64, Tyr152 and Lys156 (in the numbering of Drosophila alcohol dehydrogenase), but no histidine or cysteine residue like in the completely different, classical medium-chain alcohol dehydrogenase family. Asp64 is in the suggested coenzyme-binding domain, whereas Tyr152 and Lys156 are close to the center of the protein chain, at a putative inter-domain, active-site segment. Consequently, the overall comparisons suggest the possibility of related mechanisms and domain properties for different members of the short-chain family.

Alcohol Dehydrogenase

Terminal deoxynucleotidyltransferase. Alignment of alpha- and beta-subunits of the core enzyme along the primary translation product.

Terminal deoxynucleotidyltransferase exists in multiple Mr forms, all apparently generated from a single polypeptide of 62kDa. On isolation and purification, the smallest catalytically active protein of this enzyme consists of two subunits, alpha (12kDa) and beta (30kDa). Recently a complementary-DNA nucleotide sequence has been reported for a portion of the enzyme from human lymphoblast. We have pinpointed the locations of the alpha- and beta-subunits within the elucidated nucleotide sequence. From these data, the portions of the nucleotide sequence coding for the catalytically important area of the transferase can be estimated. Here the amino acid sequence of a number of tryptic peptides from calf alpha- and beta-subunits is presented. Because of the striking homology between the amino acid sequence of the calf enzyme and that predicted for human lymphoblast enzyme, it is possible for us to conclude that the alpha-subunit was generated from the C-terminus of the precursor protein and the beta-subunit was non-overlapping and proximal.

Amino Acid Sequence

The role of basement membranes in vascular development.

Endothelial cells produce and bind to multiple basement membrane components. Fibronectin and interstitial collagens seem to promote migration and proliferation, whereas basement membrane collagen and laminin stimulate attachment and differentiation. Human umbilical vein endothelial cells will rapidly form capillary-like structures when plated on a reconstituted basement membrane gel. This morphological differentiation involves the alignment of the cells followed by their close association with one another and the formation of a central lumen. Using antibodies to basement membrane components, we find that the formation of these vessels is a complex process involving multiple interactions with several matrix components. Synthetic peptides to active sequences in laminin have demonstrated that at least two sites in laminin participate in tube formation. An RGD-containing site on the A chain appears to mediate cell to matrix adhesion, and synthetic RGD-containing peptides block cell to matrix adhesion during tube formation. A YIGSR-containing site on the B1 chain appears to mediate cell to cell adhesion and promote tube formation because synthetic peptides block the strong cell interactions involved in tube formation. Our data with laminin peptides show that for at least one protein, multiple sites are recognized. Such data would also suggest that several cellular receptors are involved in a concerted process in laminin-induced differentiation of endothelial cells. We conclude that vessel formation is a complex, multistep process. Identification of active sites that block this process may have potential use in blocking angiogenesis in diseases such as diabetic retinopathy and Kaposi's sarcoma.

Amino Acid Sequence

Illuminating the mystery of thylacine extinction: a role for relaxed selection and gene loss.

Gene loss shapes lineage-specific traits but is often overlooked in species survival. In this study, we investigate the role of ancestral gene loss using the extinction icon-thylacine (Thylacinus cynocephalus). While studies of neutral genetic variation indicate a population decline before extinction, the impact of thylacine-specific ancestral gene losses remains unexplored. The availability of a chromosomal-level genome of the extinct thylacine offers a unique opportunity for such comparative studies. Here, we leverage palaeogenomic data to compare gene presence/absence patterns between the Tasmanian devil and thylacine. We discovered ancestral (between 13-1 Ma) loss of SAMD9L, HSD17B13, CUZD1 and VWA7 due to multiple gene-inactivating mutations, corroborated by short-read sequencing. The timing of gene loss mirrors the thylacine's shift towards hypercarnivory and increased body size. Notably, the loss of SAMD9 correlates with a carnivorous diet. Our genome-wide analysis reveals olfactory receptor loss and relaxed selection, aligning with reduced olfactory lobes in the thylacine, indicating olfaction is not its primary hunting sense. By integrating palaeogenomic data with comparative genomics, our study reveals ancestral gene losses and their impact on species survival and resilience to environmental changes. Our approach can be extended to other extinct and endangered species, helping to identify genetic factors for conservation efforts.

Animals

Fast-growing Bacillus sensu lato rhizosphere populations are constrained by antagonistic Pseudomonadota, Actinomycetota and other Bacillus sensu lato.

Copiotrophic Bacillus and related taxa grow rapidly and are commonly isolated from soil. Despite their growth rate, Bacillus sensu lato (BSL) constitute less than one percent of soil bacterial communities, and the nutrient-enriched rhizosphere contains even fewer. Amendment of bulk soil with synthetic root exudate did not lead to increase in Bacillus culturable counts. We hypothesized that BSL populations in soil enriched with growth-supporting carbon are suppressed by various soil microbes. A screen using B. pseudomycoides as tester strain yielded 124 growth inhibiting isolates, aligning by 16S rRNA genes to 3 Alphaproteobacteria, 6 Betaproteobacteria, 5 Gammaproteobacteria, 3 Streptomyces, and 19 Bacillaceae. Most antagonists also suppressed four other BSL, and over 70% of the BSL isolates suppressed each other. The 11 sequenced BSL genomes encoded between 2 and 10 antibiotic biosynthetic gene clusters. Incubation of multiple isolates in artificial soil microcosms resulted in population growth restraint through a high percentage of endospores formed. This indicated that growth suppression by antagonists was due primarily to induction of sporulation. These results support our hypothesis that Bacillus populations in soil enriched with growth-supporting carbon are suppressed by various soil microbes.

Bacillus

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

Automated assembly of protein blocks for database searching.

A system is described for finding and assembling the most highly conserved regions of related proteins for database searching. First, an automated version of Smith's algorithm for finding motifs is used for sensitive detection of multiple local alignments. Next, the local alignments are converted to blocks and the best set of non-overlapping blocks is determined. When the automated system was applied successively to all 437 groups of related proteins in the PROSITE catalog, 1764 blocks resulted; these could be used for very sensitive searches of sequence databases. Each block was calibrated by searching the SWISS-PROT database to obtain a measure of the chance distribution of matches, and the calibrated blocks were concatenated into a database that could itself be searched. Examples are provided in which distant relationships are detected either using a set of blocks to search a sequence database or using sequences to search the database of blocks. The practical use of the blocks database is demonstrated by detecting previously unknown relationships between oxidoreductases and by evaluating a proposed relationship between HIV Vif protein and thiol proteases.

Algorithms