PubMed HealthSearch

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

CryoSCAPE: Scalable immune profiling using cryopreserved whole blood for multi-omic single cell and functional assays.

BACKGROUND: The field of single cell technologies has rapidly advanced our comprehension of the human immune system, offering unprecedented insights into cellular heterogeneity and immune function. While cryopreserved peripheral blood mononuclear cell (PBMC) samples enable deep characterization of immune cells, challenges in clinical isolation and preservation limit their application in underserved communities with limited access to research facilities. We present CryoSCAPE (Cryopreservation for Scalable Cellular And Proteomic Exploration), a scalable method for immune studies of human PBMC with multi-omic single cell assays using direct cryopreservation of whole blood. RESULTS: Comparative analyses of matched human PBMC from cryopreserved whole blood and density gradient isolation demonstrate the efficacy of this methodology in capturing cell proportions and molecular features. The method was then optimized and verified for high sample throughput using fixed single cell RNA sequencing and liquid handling automation with a single batch of 60 cryopreserved whole blood samples. Additionally, cryopreserved whole blood was demonstrated to be compatible with functional assays, enabling this sample preservation method for clinical research. CONCLUSIONS: The CryoSCAPE method, optimized for scalability and cost-effectiveness, allows for high-throughput single cell RNA sequencing and functional assays while minimizing sample handling challenges. Utilization of this method in the clinic has the potential to democratize access to single-cell assays and enhance our understanding of immune function across diverse populations.

Humans

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

Multichannel genomic recording of biological information with ENGRAM.

Molecular recording is an emerging paradigm for measuring biology over time. Enhancer-mediated genomic recording of activity in multiplex (ENGRAM) is a recently described synthetic biology circuit architecture that converts the transient activity of cis-regulatory elements (CREs) into stable genomic records that can be retrospectively recovered via DNA sequencing. Here we provide a step-by-step protocol for conducting ENGRAM experiments and analyzing the resulting data. We also describe key design considerations for ENGRAM recorders, summarize the strengths and limitations of ENGRAM, and highlight applications, including multiplex signal recording and high-throughput CRE screening. In contrast to other systems for DNA-based recording in mammalian systems, ENGRAM relies on prime editing-mediated insertions to record the activity of a given CRE, such that it is inherently multiplexable-for example, four-base-pair insertions can represent the activities of up to 256 distinct CREs. A further contrast lies with ENGRAM's compatibility with DNA Typewriter, which facilitates the capture of signal order. For users with basic skills in molecular biology, mammalian cell culture and DNA sequencing analysis, ENGRAM experiments can typically be completed within 5-6 weeks.

Genomics

Eclosion hormone of the silkworm Bombyx mori. Expression in Escherichia coli and location of disulfide bonds.

A gene encoding eclosion hormone (EH) from the silkworm, Bombyx mori was chemically synthesized, inserted into a secretion vector and expressed in Escherichia coli, leading to the production of biologically active EH. Sequence analysis of cystine-containing peptides in a thermolysin digest of this EH established the locations of 3 disulfide bonds in the molecule. Evidence was also obtained that the 6 residues at the NH2-terminal are dispensable but 4 residues at the COOH-terminal play an important role in EH activity.

Amino Acid Sequence

Analyzing and comparing nucleic acid sequences by hybridization to arrays of oligonucleotides: evaluation using experimental models.

An efficient method was developed for making complete sets of oligonucleotides of defined length, covalently attached to the surface of a glass plate, by synthesizing them in situ. A device carrying all octapurine sequences was used to explore factors affecting molecular hybridization of the tethered oligonucleotides, to develop computer-aided methods for analyzing the data, and to test the feasibility of using the method for sequence analysis. Further development is needed before the method can be used routinely, but our work shows that it has a number of potential advantages over gel-based methods: it should be easy to automate; the quality of the sequence results can be evaluated statistically; it provides a powerful way of comparing related sequences and detecting mutation; it can be applied to both DNA and RNA; and specific motifs can be incorporated into all sequences of the array to focus analysis on sequences of biological interest.

Base Sequence

Expression and characterization of recombinant TGF-beta 2 proteins produced in mammalian cells.

Recombinant DNA plasmids coding for transforming growth factor beta 2 (TGF-beta 2) precursor and a hybrid TGF-beta 1(NH2)/beta 2(COOH) molecule consisting of the amino-terminal precursor portion of transforming growth factor-beta 1 (TGF-beta 1) linked in phase to the carboxyl terminus of mature TGF-beta 2 were constructed and transfected into COS cells. Both plasmids directed the synthesis of active TGF-beta 2 which was secreted into the supernatants of transfected cells. The TGF-beta 2 was secreted in a latent form, as an acidification step was required to demonstrate optimal biological activity. Using site-specific anti-peptide antibodies, we show that precursor and mature forms of TGF-beta 2 are produced. A stable Chinese hamster ovary (CHO) cell line expressing the hybrid TGF-beta 1(NH2)/beta 2(COOH) protein was isolated. This cell line secreted both precursor and mature forms of TGF-beta 1(NH2)/beta 2(COOH); acidification was required to demonstrate biological activity. Protein sequence analysis of recombinant TGF-beta 2 produced by this CHO clone demonstrated that correct proteolytic cleavage had occurred, suggesting that the processing signals contained within the TGF-beta 1 amino portion can function in producing mature TGF-beta 2. Receptor binding studies showed that TGF-beta 2 specifically bound predominantly to type III receptors on the surface of human palatal mesenchymal cells. The availability of active TGF-beta 2 should aid in determining its potential therapeutic use as a growth modulator.

Animals

A new family of powerful multivariate statistical sequence analysis techniques.

A novel multivariate statistical approach is presented for extracting and exploiting intrinsic information present in our ever-growing sequence data banks. The information extraction from the sequences avoids the pitfalls of intersequence alignment by analyzing secondary invariant functions derived from the sequences in the data bank rather than the sequences themselves. Such typical invariant function is a 20 x 20 histogram of occurrences of amino acid pairs in a given sequence or fragment thereof. To illustrate the potential of the approach an analysis of 10,000 protein sequences from the National Biomedical Research Foundation Protein Identification Resource is presented, whose analysis already reveals great biological detail. For example, zeta-hemoglobin is found to lie close to amphibian and fish chi-hemoglobin which, in turn, is an important clue to the physiological function of this mammalian early embryonic hemoglobin. The multivariate statistical framework presented unifies such apparently unrelated issues as phylogenetic comparisons between a set of sequences and distance matrices between the constituents of the biological sequences. The Multivariate Statistical Sequence Analysis (MSSA) principles can be used for a wide spectrum of sequence analysis problems such as: assignment of family memberships to new sequences, validation of new incoming sequences to be entered into the database, prediction of structure from sequence, discrimination of coding from non-coding DNA regions, and automatic generation of an atlas of protein or DNA sequences. The MSSA techniques represent a self-contained approach to learning continuously and automatically from the growing stream of new sequences. The MSSA approach is particularly likely to play a significant role in major sequencing efforts such as the human genome project.

Amino Acid Sequence

Protocol to predict gene expression from transcriptomic data using PREDICT.

Linking DNA sequence variation to context-specific transcriptional programs is a critical challenge in regulatory genomics, especially for non-model organisms. Here, we present PREDICT, a modular Python package for discovering cis-regulatory elements and transcription factor binding motifs. We describe steps to identify enriched k-mers from differentially expressed genes, map them to known motifs, quantify their impact on gene expression, and visualize motif co-occurrences. PREDICT provides a robust, k-mer-based approach to uncover regulatory logic in diverse genomic systems. For complete details on the use and execution of this protocol, please refer to Yen et al. and Liu et al.1,2.

Gene Expression Profiling

Nucleotide sequence of the type A staphylococcal enterotoxin gene.

We determined the nucleotide sequence of the gene encoding staphylococcal enterotoxin A (entA). The gene, composed of 771 base pairs, encodes an enterotoxin A precursor of 257 amino acid residues. A 24-residue N-terminal hydrophobic leader sequence is apparently processed, yielding the mature form of staphylococcal enterotoxin A (Mr, 27,100). Mature enterotoxin A has 82, 72, 74, and 34 amino acid residues in common with staphylococcal enterotoxins B and C1, type A streptococcal exotoxin, and toxic shock syndrome toxin 1, respectively. This level of homology was determined to be significant based on the results of computer analysis and biological considerations. DNA sequence homology between the entA gene and genes encoding other types of staphylococcal enterotoxins was examined by DNA-DNA hybridization analysis with probes derived from the entA gene. A 624-base-pair DNA probe that represented an internal fragment of the entA gene hybridized well to DNA isolated from EntE+ strains and some EntA+ strains. In contrast, a 17-base oligonucleotide probe that encoded a peptide conserved among staphylococcal enterotoxins A, B, and C1 hybridized well to DNA isolated from EntA+, EntB+, EntC1+, and EntD+ strains. These hybridization results indicate that considerable sequence divergence has occurred within this family of exotoxins.

Amino Acid Sequence

Human and murine interleukin 1 possess sequence and structural similarities.

The molecular cloning and sequence analysis for human and murine interleukin 1 precursor have recently been described. Comparison of the amino acid sequences resulting from these data can be used to aid in the identification of conserved regions essential to biological activity. We report results which confirm the relationship between these two molecules and suggest that specific regions may be essential for activity. Amino terminal sequence analysis of a 19,000 Mr biologically active IL-1 isolated from stimulated human monocytes reveals a sequence which is in good agreement with that inferred from the human cDNA and, furthermore, locates the processed amino terminus at a site similar to that described for the murine sequence.

Amino Acid Sequence

Characterizing the regulatory logic of transcriptional control at the DNA sequence level by ensembles of thermodynamic models.

MOTIVATION: Understanding how the genome encodes the regulatory logic of transcription is a main challenge of the post-genomic era, and can be overcome with the aid of customized computational tools. RESULTS: We report an automated framework for analyzing an ensemble of fits to data of a thermodynamics-based sequence-level model for transcriptional regulation. The fits are clustered accordingly with their intrinsic regulatory logic. A multiscale analysis enables visualization of quantitative features resulting from the deconvolution of the regulatory profile provided by multiple transcription factors interacting with the locus of a gene. Quantitative experimental data on reporters driven by the whole locus of the even-skipped gene in the blastoderm of Drosophila embryos was used for validating our approach. A few clusters of highly active DNA binding sites within the enhancers collectively modulate even-skipped gene transcription. Analysis of variable enhancers' length shows the importance of bound protein-protein interactions for transcriptional regulation. The interplay between activation and quenching enables function conservation of enhancers despite length variations. AVAILABILITY AND IMPLEMENTATION: The transcription factor level data used for performing the reported study is accessible in the input files in Zenodo and GitHub as well the full code. Additional data from formerly FlyEx database will be available under request.

Thermodynamics

A scale-independent signal processing method for sequence analysis.

In this paper, we present methods to detect and localize patterns in biologically related protein sequences (family). The patterns common to the sequences of the family are detected by using Fourier analysis. No previous scales (codes) are needed, they are actually produced as a result of the analysis procedure, together with the frequencies of the Fourier decompositions. Characteristic features of the family are thus expressed as (code-frequency) pairs. Various tools are proposed in order to localize the patterns, to compare the codes, and to evaluate the proximity of an arbitrary sequence to the investigated family. The general strategy is illustrated on a family composed of calcium-binding proteins.

Amino Acid Sequence

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics

Receptor mechanisms. Structure and molecular biology of transmitter receptors.

Details of receptor structure and function that were unavailable as recently as two years ago are now readily obtainable through the application of molecular biological techniques. Cloning and sequence analysis of neurotransmitter receptor genes have provided information on the primary structure of these proteins, revealing the relationship between pharmacologically diverse families of receptors. Knowledge of the primary structure of receptors has allowed for prediction of secondary structure and the construction of three-dimensional models. Permanent expression of cloned neurotransmitter receptor genes in cultured cells is providing unlimited sources of pure receptor, which allows for pharmacological and biochemical studies on single receptor subtypes. The use of site-directed mutagenesis to elucidate the relationship between protein structure and function has provided considerable information on the role of certain conserved amino acids in receptor function and has suggested possible molecular mechanisms of signal transduction across membranes. The article will review some of these recent developments in the area of neurotransmitter receptors and point out the utility of molecular biology in these endeavors.

Amino Acid Sequence

Whole metagenome sequencing: not deep enough for complete microbial function recovery.

BACKGROUND: Whole metagenome shotgun sequencing (WMS) is widely used to profile microbial function. However, technical variability in sequencing and analysis often obscures true biological patterns. Large-scale studies are particularly susceptible to batch effects, such as differences in sequencing depth and platform and annotation strategies, as well as sample-to-flow-cell assignments. However, the relative effects of these factors on functional inference in such studies have yet to be systematically evaluated. We analyzed oral-rinse WMS data from 671 Nigerian youths aged 9-18, sequenced on two Illumina platforms. Microbial molecular functionality encoded in these data was annotated using the mi-faser/Fusion pipeline, to capture the broad functional repertoire, and HUMAnN 3/EC numbers pipeline to characterize curated enzymatic activities. We then quantified how technical factors and batch effects shaped the recovery of microbial functionality. RESULTS: Three findings of our work were most salient. First, we observed that the choice of annotation strategy traded off between breadth and specificity of functional coverage. Second, we found that low-prevalence functions were disproportionately lost at shallow sequencing depths, indicating that in, e.g., case-control studies with few representatives of the minor class, sequencing depth could critically impact study resolution. Finally, using our newly developed model relating sequencing depth to functional recovery, we demonstrated that increasing sequencing depth does not directly or proportionally improve functional recall. That is, at as little as 10% of this study's sequencing depth, 30% of the estimated complete microbiome functional repertoire was detectable. However, even at the full depth used in this study, we were only able to recover an estimated 60% of that complete functional repertoire. We further showed that despite biomes differences in functional diversity and host contamination levels (e.g., soil, fecal), incomplete functional recovery at commonly used sequencing depths was consistently observed. CONCLUSIONS: Together, these findings and our depth-to-function mapping framework provide practical guidelines for the design and interpretation of WMS studies. Coordinating sequencing depth planning with annotation strategy, experimental design, and rigorous batch control is thus essential for robust detection of microbial functions and for ensuring reproducible microbiome insights. Video Abstract.

Humans

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~ 28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA

Synthesis and characterization of the Kunitz protease-inhibitor domain of the beta-amyloid precursor protein.

To understand the pathological process by which amyloid is deposited in Alzheimer's disease, it is important to characterize the proteolytic processing events of the beta-amyloid precursor protein (beta-APP) from which the amyloid-forming fragment is excised. A potentially important component in beta-APP processing is the 57-amino acid (aa) Kunitz serine protease inhibitor (KPI) located within the extracellular domain of both the 751- and 770-aa isoforms of beta-APP. We have synthesized DNA encoding the 57-aa KPI domain as a necessary step in identifying the role of the protease inhibitor in beta-APP processing and amyloid formation. A bacterial secretion system directed by the alkaline phosphatase signal peptide of Escherichia coli linked to a synthetic gene encoding KPI was used to produce soluble, extracellular recombinant KPI (reKPI) protein. The reKPI protein was purified to homogeneity from bacterial supernatants and was biochemically and biologically characterized. Complete aa sequence analysis confirmed the fidelity of the reKPI, and fast-atom bombardment mass-spectral analysis was used to document that reKPI was of the predicted Mr. The reKPI is as active on a molar basis as the inhibitor-containing beta-APP when assayed for inhibition of trypsin activity. Together these data suggest that reKPI protein is properly folded and lacking in modified aa. Hence, this reKPI will be an important reagent in gaining a better understanding of the role of the KPI domain in beta-APP function and metabolism, as well as in the proteolytic events involved in beta-amyloid formation.

Alzheimer Disease

Non-reciprocal coevolution in a fungus-gardening ant.

Symbioses are often characterized by nonrandom associations between hosts and symbionts. Hosts may obtain symbionts horizontally from the environment or vertically from a parent or sometimes use both methods. Macroevolutionary examinations of fungus-gardening ants and their fungi have shown either a 1:1 coevolution model or a 'diffuse' model between ant host and fungal symbionts. However, some of these conclusions may have been based on using relatively conservative molecular markers, which could obscure cryptic variation. The use of whole genome approaches potentially offer more power in elucidating coevolutionary history. In this study, we examined patterns of coevolution in a single species (Trachymyrmex septentrionalis) using genomic and experimental approaches. We tested whether ant-fungal specificity patterns reflected either 1:1 or diffuse models of coevolution. While we report significant co-phylogenetic signal among intraspecific ant host and fungal symbiont lineages, we found evidence of 1:1 coevolution in some lineages and diffuse in others. These conclusions were supported by the results of experiments where newly mated T. septentrionalis queens were forced to grow novel fungi that suggested that not all fungi are equivalent symbionts and would require specialized hosts. Thus, within a single ant species, there is a mixed support for both models.

Animals