PubMed HealthSearch

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Nested co-expression network analysis identifies compact gene clusters in a black box.

MOTIVATION: Digital analysis of biological systems requires methods capable of identifying both broad and nested gene modules reflecting complex biological processes. Existing transcriptomic methods often miss compact gene sets corresponding to subprocesses in specialized cell types, limiting insights into functional heterogeneity. RESULTS: We present Nested-WGCNA, a two-stage unsupervised network analysis algorithm designed to identify coarse-grained and fine-grained gene modules. Applied to bulk RNA-Seq data, Nested-WGCNA reveals stable modules reproducible across datasets. When validated against scRNA-Seq data, these modules correspond to both major and minor immune cell subtypes. Application to immunotherapy response datasets uncovers predictive and prognostic biomarkers, highlighting its utility in treatment stratification and biomarker discovery. AVAILABILITY: The NestedWGCNA source code and analysis pipeline are available on GitHub (https://github.com/ilyada/NestedWGCNA) and archived on Zenodo (https://doi.org/10.5281/zenodo.18959244).

Algorithms

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean = 0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans

Total synthesis and further characterization of the gamma-carboxyglutamate-containing "sleeper" peptide from Conus geographus venom.

The total synthesis of the Gla-containing "sleeper" peptide (Gly-Glu-Gla-Gla-Leu-Gln-Gla-Asn-Gln-Gla-Leu-Ile-Arg-Gla-Lys-Ser-Asn-NH2 ) from Conus geographus is described. A new strategy for the synthesis of acid-sensitive peptide amides was developed, which allowed complete deprotection and cleavage of the L-gamma-carboxyglutamate-containing peptide from the 2,4-dimethoxybenzhydrylamine resin. Synthetic sleeper peptide, after preparative high-performance liquid chromatography (HPLC) purification, was shown to be identical with the native peptide by all criteria (coelution experiments of HPLC, sequence analysis, and biological activity). In addition, a developmental switch in the behavioral symptoms induced by the peptide after intracerebral administration in mice was documented. At low doses of the peptide (4-30 pmol/g), a sleeplike state was induced in mice under 2 weeks old; in contrast, older mice became markedly hyperactive. It is proposed that, in the presence of Ca2+, the sleeper peptide assumes an alpha-helical configuration in which all the gamma-carboxyglutamate residues are located on the same side of the alpha-helix.

1-Carboxyglutamic Acid

Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite.

Nucleotide recoding RNA sequencing methods (NR-seq; TimeLapse-seq, SLAM-seq, TUC-seq, etc.) are powerful approaches for assaying transcript population dynamics. In addition, these methods have been extended to probe a host of regulated steps in the RNA life cycle. Current bioinformatic tools significantly constrain analyses of NR-seq data. To address this limitation, we developed EZbakR (https://github.com/isaacvock/EZbakR), an R package to facilitate a more comprehensive set of NR-seq analyses, and fastq2EZbakR (https://github.com/isaacvock/fastq2EZbakR), a Snakemake pipeline for flexible preprocessing of NR-seq datasets, collectively referred to as the EZbakR suite. Together, these tools generalize many aspects of the NR-seq analysis workflow. The fastq2EZbakR pipeline can assign reads to a diverse set of genomic features (e.g., genes, exons, splice junctions), and EZbakR can perform analyses on any combination of these features. EZbakR extends standard NR-seq mutational modeling to support multi-label analyses (e.g., s4U and s6G dual labeling), and implements an improved hierarchical model to better account for transcript-to-transcript variance in metabolic label incorporation. EZbakR also generalizes dynamical systems modeling of NR-seq data to support analyses of premature mRNA processing and flow between subcellular compartments. Finally, EZbakR implements flexible and well-powered comparative analyses of all estimated parameters via design matrix-specified generalized linear modeling. The EZbakR suite will thus allow researchers to make full, effective use of NR-seq data.

Software

Isolation and characterization of TGF-beta 2 and TGF-beta 5 from medium conditioned by Xenopus XTC cells.

TGF-beta 2 and -beta 5 have been purified from medium conditioned by Xenopus cultured cells (XTC) and identified based on their N-terminal amino acid sequence analysis and biological activity. When applied in high concentrations, Xenopus TGF-beta 2, like porcine TGF-beta 2, induces expression of mesodermal markers from cultured Xenopus ectodermal explants, whereas TGF-beta 5 is inactive in this assay. However, the TGF-beta 's could be separated from the major mesoderm-inducing activity present in XTC medium. Xenopus TGF-beta 2 and -beta 5 are approximately equivalent to TGF-beta 1 in their abilities to inhibit the growth of mink lung CCL-64 cells, induce anchorage-independent growth of rat NRK cells, inhibit the proliferation and antibody secretion of human B-lymphocytes, and stimulate chemotaxis of human monocytes. These data establish the functional activity of TGF-beta 5 and suggest that more complex multicellular systems, in contrast to most isolated cells, discriminate between the different TGF-beta s.

Amino Acid Sequence

Activation of an N-ras gene in acute myeloblastic leukemia through somatic mutation in the first exon.

A transforming N-ras gene has been cloned from acute myeloblastic leukemia bone marrow cells, in parallel with the N-ras gene derived from fibroblasts of the same patient. N-ras derived from fibroblasts lacked focus-forming activity in NIH/3T3 cells, indicating that gene activation in the leukemia cells must have occurred by a somatic event. Construction of chimeric molecules between the transforming and the normal N-ras genes and subsequent biological and sequence analysis of these constructs revealed that the transforming gene was altered by a point mutation changing amino acid 12 of the N-ras protein from glycine to aspartic acid.

Alleles

vcfgl: a flexible genotype likelihood simulator for VCF/BCF files.

MOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl.

Software

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing

Fast and flexible minimizer digestion with digest.

SUMMARY: Minimizer digestion is an increasingly common component of bioinformatics tools, including tools for de Bruijn graph assembly and sequence classification. We describe a new open source tool and library to facilitate efficient digestion of genomic sequences. It can produce digests based on the related ideas of minimizers, modimizers or syncmers. Digest uses efficient data structures, scales well to many threads, and produces digests with expected spacings between digested elements. AVAILABILITY AND IMPLEMENTATION: Digest is implemented in C++17 with a Python API, and is available open-source at https://github.com/VeryAmazed/digest. The python library is available on Bioconda. Rust bindings are available as a public crate at https://crates.io/crates/digest-rs.

Software

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning

GraphyloVar: predicting the impact of non-coding variants using a multi-species sequence model.

MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on &#x223c;149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818.

Phylogeny

Novel transfer RNAs that are active in Escherichia coli.

Many of the mammalian mitochondrial tRNAs contain significant nucleotide deletions in the dihydrouridine (D) stem or T psi C stem, so that they cannot fold into the canonical cloverleaf structure. This suggests that alternative forms and shapes are possible for a mitochondrial tRNA that functions in the specialized translational apparatus of the mammalian mitochondria. The question of whether significant structural alterations may be accommodated by a bacterial protein synthesis machinery, such as in Escherichia coli, is unanswered. In this work, all but ten positions in the gene for the 76-nucleotide coding sequence of an E. coli amber suppressor tRNA were permuted and screened for biological activity in vivo. Sequence analysis of a collection of biologically active variants established that many have unusual structures that include base-pair mismatches in helical stems, substitutions of normally conserved bases, and deletions. Independent mutations were obtained that weaken base pairs or tertiary interactions that normally stabilize the coaxial stacking of the D and anticodon stems, suggesting that the translational apparatus can accommodate considerable flexibility in this part of the molecule. The results demonstrate the capacity of the bacterial protein synthetic apparatus to accommodate altered tRNA structures that are not represented by any naturally occurring tRNAs.

Base Sequence

The chemistry and biology of thymosin. II. Amino acid sequence analysis of thymosin alpha1 and polypeptide beta1.

The amino acid sequences of two polypeptide components of thymosin Fraction 5 termed thymosin alpha1 and polypeptide beta1 have been established. The sequences were determined by automatic Edman degradation of the intact molecules as well as by manual sequence analysis of the enzymatic cleavage products. Thymosin alpha1, an immunologically active polypeptide, is highly acidic with an isoelectric point of 4.2. This molecule is composed of 28 amino acid residues with acetylserine as the NH2 terminus. A chemically synthesized molecule of thymosin alpha1 has been found to be as active as the natural molecule in our bioassay systems. Polypeptide beta1 is a molecule consisting of 74 amino acid residues and has an isoelectric point of 6.7. This peptide is not biologically active in our assay systems, suggesting that it is not involved in thymic hormone action. The sequence of beta1 was found to be identical with ubiquitin and a portion of protein A24, a nuclear chromosomal protein. The relationships among these proteins are discussed.

Amino Acid Sequence

Predicting and comparing transcription start sites in single cell populations.

The advent of 5' single-cell RNA sequencing (scRNA-seq) technologies offers unique opportunities to identify and analyze transcription start sites (TSSs) at a single-cell resolution. These technologies have the potential to uncover the complexities of transcription initiation and alternative TSS usage across different cell types and conditions. Despite the emergence of computational methods designed to analyze 5' RNA sequencing data, current methods often lack comparative evaluations in single-cell contexts and are predominantly tailored for paired-end data, neglecting the potential of single-end data. This study introduces scTSS, a computational pipeline developed to bridge this gap by accommodating both paired-end and single-end 5' scRNA-seq data. scTSS enables joint analysis of multiple single-cell samples, starting with TSS cluster prediction and quantification, followed by differential TSS usage analysis. It employs a Binomial generalized linear mixed model to accurately and efficiently detect differential TSS usage. We demonstrate the utility of scTSS through its application in analyzing transcriptional initiation from single-cell data of two distinct diseases. The results illustrate scTSS's ability to discern alternative TSS usage between different cell types or biological conditions and to identify cell subpopulations characterized by unique TSS-level expression profiles.

Transcription Initiation Site

EPIPDLF: a pretrained deep learning framework for predicting enhancer-promoter interactions.

MOTIVATION: Enhancers and promoters, as regulatory DNA elements, play pivotal roles in gene expression, homeostasis, and disease development across various biological processes. With advancing research, it has been uncovered that distal enhancers may engage with nearby promoters to modulate the expression of target genes. This discovery holds significant implications for deepening our comprehension of various biological mechanisms. In recent years, numerous high-throughput wet-lab techniques have been created to detect possible interactions between enhancers and promoters. However, these experimental methods are often time-intensive and costly. RESULTS: To tackle this issue, we have created an innovative deep learning approach, EPIPDLF, which utilizes advanced deep learning techniques to predict EPIs based solely on genomic sequences in an interpretable manner. Comparative evaluations across six benchmark datasets demonstrate that EPIPDLF consistently exhibits superior performance in EPI prediction. Additionally, by incorporating interpretable analysis mechanisms, our model enables the elucidation of learned features, aiding in the identification and biological analysis of important sequences. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/xzc196/EPIPDLF.

Deep Learning

v-maf, a viral oncogene that encodes a "leucine zipper" motif.

We have molecularly cloned the provirus of the avian musculoaponeurotic fibrosarcoma virus AS42. Nucleotide sequence analysis of a biologically active clone of AS42 showed that this virus encodes a viral oncogene, maf. The deduced amino acid sequence of the v-maf gene product contains a "leucine zipper" motif similar to that found in a number of DNA binding proteins, including the gene products of the fos, jun, and myc oncogenes. However, unlike these oncogenes, the cellular maf gene was not transcriptionally activated by growth stimulation of cultured cells.

Amino Acid Sequence

A computer method for finding common base paired helices in aligned sequences: application to the analysis of random sequences.

We describe a new computer program that identifies conserved secondary structures in aligned nucleotide sequences of related single-stranded RNAs. The program employs a series of hash tables to identify and sort common base paired helices that are located in identical positions in more than one sequence. The program gives information on the total number of base paired helices that are conserved between related sequences and provides detailed information about common helices that have a minimum of one or more compensating base changes. The program is useful in the analysis of large biological sequences. We have used it to examine the number and type of complementary segments (potential base paired helices) that can be found in common among related random sequences similar in base composition to 16S rRNA from Escherichia coli. Two types of random sequences were analyzed. One set consisted of sequences that were independent but they had the same mononucleotide composition as the 16S rRNA. The second set contained sequences that were 80% similar to one another. Different results were obtained in the analysis of these two types of random sequences. When 5 sequences that were 80% similar to one another were analyzed, significant numbers of potential helices with two or more independent base changes were observed. When 5 independent sequences were analyzed, no potential helices were found in common. The results of the analyses with random sequences were compared with the number and type of helices found in the phylogenetic model of the secondary structure of 16S ribosomal RNA. Many more helices are conserved among the ribosomal sequences than are found in common among similar random sequences. In addition, conserved helices in the 16S rRNAs are, on the average, longer than the complementary segments that are found in comparable random sequences. The significance of these results and their application in the analysis of long non-ribosomal nucleotide sequences is discussed.

Base Composition