PubMed HealthSearch

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Non-reciprocal coevolution in a fungus-gardening ant.

Symbioses are often characterized by nonrandom associations between hosts and symbionts. Hosts may obtain symbionts horizontally from the environment or vertically from a parent or sometimes use both methods. Macroevolutionary examinations of fungus-gardening ants and their fungi have shown either a 1:1 coevolution model or a 'diffuse' model between ant host and fungal symbionts. However, some of these conclusions may have been based on using relatively conservative molecular markers, which could obscure cryptic variation. The use of whole genome approaches potentially offer more power in elucidating coevolutionary history. In this study, we examined patterns of coevolution in a single species (Trachymyrmex septentrionalis) using genomic and experimental approaches. We tested whether ant-fungal specificity patterns reflected either 1:1 or diffuse models of coevolution. While we report significant co-phylogenetic signal among intraspecific ant host and fungal symbiont lineages, we found evidence of 1:1 coevolution in some lineages and diffuse in others. These conclusions were supported by the results of experiments where newly mated T. septentrionalis queens were forced to grow novel fungi that suggested that not all fungi are equivalent symbionts and would require specialized hosts. Thus, within a single ant species, there is a mixed support for both models.

Animals

Construction and biological analysis of deletion mutants of Fujinami sarcoma virus: 5'-fps sequence has a role in the transforming activity.

Fujinami sarcoma virus (FSV) genome codes for the gag-fps fusion protein FSV-P130. The amino acid sequence of the 3' one-third portion in v-fps is partially homologous to the 3' half of pp60src, or the kinase domain, but the sequence of the 5' portion is unique to v-fps. To identify a possible domain structure in the v-fps sequence responsible for cell transformation, we constructed various deletion mutants of FSV with molecularly cloned viral DNA. Their transforming activities were assayed by measuring focus formation on chicken embryo fibroblasts and rat 3Y1 cells and tumor formation in chickens. The mutants carrying a deletion at the 3' portion in v-fps, the kinase domain, lost transforming activity. The mutants carrying an approximately 1-kilobase deletion within the 5' portion of the v-fps sequence retained focus-forming activity and tumorigenicity in the chicken system, but the efficiency of focus formation was about 10 times lower than that of the wild type. The morphology of these transformed cells was distinct from that observed in cells infected with wild-type FSV. Furthermore, these mutants could not transform rat 3Y1 cells, although wild-type FSV DNA transformed rat 3Y1 cells at a high frequency. The mutants carrying a larger deletion in the 5' portion of fps completely lacked the transforming activity. These results suggest that the 3' portion of the v-fps sequence is necessary but not sufficient for cell transformation and that the 5' portion of v-fps has a role in the transforming activity.

Animals

The interaction of boar sperm proacrosin with its natural substrate, the zona pellucida, and with polysulfated polysaccharides.

Boar sperm acrosin is an acrosomal protease with trypsin-like specificity, and it functions in fertilization by assisting sperm passage through the zona pellucida by limited hydrolysis of this extracellular matrix. In addition to a proteolytic active site domain, acrosin binds the zona pellucida at a separate binding domain that is lost during proacrosin autolysis. In this study, we quantitate the binding of proacrosin to the physiological substrate for acrosin, the zona pellucida, and to a non-substrate, the polysulfated polysaccharide fucoidan. Binding was analogous to sea urchin sperm bindin that binds egg jelly fucan and the vitelline envelope of sea urchin eggs. Proacrosin was found to bind to fucoidan and to the zona pellucida with binding affinities similar to bindin interaction with egg jelly fucan. These interactions were competitively inhibited by similar relative molecular mass polysulfated polymers. Since bindin and proacrosin have distinctly different amino acid sequences, their interaction with acidic sulfate esters demonstrates an example of convergent evolution wherein different macromolecules localized in analogous sperm compartments have the same biological function. From cDNA sequence analysis of proacrosin, this binding may be mediated through a consensus sequence for binding sulfated glycoconjugates. Proacrosin binding to the zona pellucida may serve as both a recognition or primary sperm receptor, as well as maintaining the sperm on the zona pellucida once the acrosome reaction has occurred.

Acrosin

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans

Atriopeptin biochemical pharmacology.

Several low-molecular-weight peptides that possess potent natriuretic, diuretic, and vascular smooth muscle relaxant activity have been isolated from atrial extracts. Elucidation of their structure indicates that they consist of a 17-membered ring of amino acids formed by a cystine disulfide bond and that they differ only in the composition of the amino and carboxy termini. The 24-amino-acid peptide atriopeptin (AP) III was selected as the reference compound for structure-activity studies. Amino-terminal amino acid extensions on APIII markedly increase the natriuretic-diuretic but not the renal vasodilatory response in anesthetized dogs, which suggests a heterogeneity of AP receptors in renal tubular and vascular tissues. Radioligand (125I-labeled APIII) binding studies with fresh rat kidney slices indicate that the primary renal sites of specific AP binding are in the glomerulus and in the papillary segment of the medulla, thus implicating these structures in the natriuretic-diuretic effect. Data obtained from radioimmunoassay, chromatographic migration, vasorelaxant biological activity, and peptide sequence analysis indicate that Ser-Leu-Arg-Arg-APIII is the major circulating form of low-molecular-weight atrial peptide present in rat plasma. Circulating APs fulfill many of the criteria for involvement in the endocrine regulation of fluid and electrolyte homeostasis.

Amino Acid Sequence

CAKR: commutative algebra k-mer representations for genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Genomics

Biological and structural properties of MIP-1 alpha expressed in yeast.

The murine macrophage inflammatory proteins-1 alpha (MIP-1 alpha) and MIP-1 beta are distinct but closely related cytokines. Partially purified mixtures of the two proteins affect neutrophil function and cause local inflammation and fever. The particular properties of MIP-1 alpha have not been well studied, although it has been identified as being identical to an inhibitor of haemopoietic stem cell growth. We have expressed MIP-1 alpha in yeast cells and purified it to sequence homogeneity. Structural analysis of this biologically active material by circular dichroism and fluorescence spectroscopy confirms that MIP-1 alpha has a very similar secondary and tertiary structure to platelet factor 4 and interleukin 8 with which it shares limited sequence homology. The in-vitro stem cell inhibitory properties have been confirmed using a range of murine progenitor cells including purified bone marrow progenitor cells (FACS-1), the FDCP-mix A4 cell line, and spleen colony forming unit (CFU-S) populations. Plateau levels of inhibition of stem cell growth were achieved using concentrations of 0.15 micrograms/ml MIP-1 alpha. We have also demonstrated that MIP-1 alpha is active in vivo: 5 micrograms of MIP-1 alpha per mouse given as a bolus injection, protects stem cells from subsequent in-vitro killing by tritiated thymidine. MIP-1 alpha was also shown to enhance the proliferation of more committed progenitor granulocyte macrophage-colony forming cells (GM-CFC) in response to granulocyte macrophage-colony stimulating factor (GM-CSF).

Animals

DNA sequences amplified in cancer cells: an interface between tumor biology and human genome analysis.

There is growing evidence that amplification of specific genes is associated with tumor progression. While several proto-oncogenes are known to be activated by amplification, it is clear that not all the genes involved in DNA amplification in human tumors have been discovered. Our approach to the identification of such genes is based on the 'reverse genetics' methodology. Anonymous amplified DNA fragments are cloned by virtue of their amplification in a given tumor. These sequences are mapped in the normal genome and hence define a new genetic locus. The amplified domain is isolated by long-range cloning and analyzed along three lines of investigation: new genes are sought that can explain the biological significance of the amplification; the structure of the domain is studied in normal cells and in the amplification unit in the cancer cell; attempts are made to identify molecular probes of diagnostic value within the amplified domain. This application of genome technology to cancer biology is demonstrated in our study of a new genomic domain at chromosome 10q26 which is amplified specifically in human gastric carcinomas.

Blotting, Southern

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

A novel method for the rapid detection of specific nucleotide sequences in crude biological samples without blotting or radioactivity; application to the analysis of hepatitis B virus in human serum.

The detection of a little as 0.2 pg (60,000 molecules) of hepatitis B viral (HBV) DNA in human serum samples in 4 h has been demonstrated using a solution-hybridization and bead-capture method. An amplification method based on chemically crosslinked oligodeoxyribonucleotides was coupled with a horseradish peroxidase-labeling scheme for the ultimate detection of the analyte. Two sets of HBV complementary synthetic oligodeoxyribonucleotide probes containing one of two types of single-stranded (ss) overhangs were employed. These ss overhangs were used to capture the probe-analyte complex onto a bead and subsequently to label it. Detection was achieved with either a chemiluminescent or colorimetric output substrate for the enzyme. Only in the presence of the virus was label specifically bound to the support. The assay was relatively unaffected by either sample composition or by the presence of heterologous nucleic acids.

Base Sequence

Out-of-the-box bioinformatics capabilities of large language models (LLMs).

Large Language Models (LLMs), AI agents and co-scientists promise to accelerate scientific discovery across fields ranging from chemistry to biology. Bioinformatics- the analysis of DNA, RNA and protein sequences plays a crucial role in biological research and is especially amenable to AI-driven automation given its computational nature. Here, we assess the bioinformatics capabilities of three popular general-purpose LLMs on a set of tasks covering basic analytical questions that include code writing and multi-step reasoning in the domain. Utilizing questions from Rosalind, a bioinformatics educational platform, we compare the performance of the LLMs vs. humans on 104 questions undertaken by 110 to 68,760 individuals globally. GPT-3.5 provided correct answers for 59/104 (58%) questions, while Llama-3-70B and GPT-4o answered 49/104 (47%) correctly. GPT-3.5 was the best performing in most categories, followed by Llama-3-70B and then GPT-4o. 71% of the questions were correctly answered by at least one LLM. The best performing categories included DNA analysis, while the worst performing were sequence alignment/comparative genomics and genome assembly. Overall, LLMs performance mirrored that of humans with lower performance in tasks in which humans had low performance and vice versa. However, LLMs also failed in some instances where most humans were correct and, in a few cases, LLMs excelled where most humans failed. To the best of our knowledge, this presents the first assessment of general purpose LLMs on basic bioinformatics tasks in distinct areas relative to the performance of hundreds to thousands of humans. LLMs provide correct answers to several questions that require use of biological knowledge, reasoning, statistical analysis and computer code.

Journal Article

Isolation of multiple biologically and chemically diverse species of epidermal growth factor.

We have analyzed several lots of epidermal growth factor (EGF) purified from murine submaxillary glands including "receptor grade" EGF from Collaborative Research and EGF from Boehringer Mannheim Biochemicals. New England Nuclear uses "receptor grade" EGF to produce 125I-labeled EGF. Though these reagents are reported to be homogeneous, we found them to be a mixture of six species. A method was developed to separate this mixture into its component parts. The individual components were chemically characterized and tested for biological potency. N-terminal sequence analysis of the unfractionated EGF-mixture reveals three different sequences starting with residues 1, 2, or 3 of the mature peptide. Each component exhibited different degrees of mitogenic and EGF receptor binding activity indicating that the N-terminal region contributes to the biological response. The species representing the complete EGF peptide is the most active species in all biological assays. A rapid method for purification of homogeneous complete EGF from commercial EGF preparations is described.

Amino Acids

The biological properties of bovine parathyroid hormone (1-41), a fragment generated from the native hormone by human leukocytes.

Native bovine parathyroid hormone (bPTH) was found to be readily cleaved with human leukocyte elastase to yield the fragments bPTH(1-41) and bPTH(42-84). These were then isolated by reverse-phase HPLC and characterised by gas-phase sequencing and amino acid analysis. The biological activities of these fragments were assessed in an adenylate cyclase bioassay using the rat osteosarcoma cell line UMR106. bPTH(1-41) was found to have approximately twice the molar potency of the native hormone from which it was derived. bPTH(42-84) had no biological activity and did not modulate the adenylate cyclase response to these cells to the native hormone. The possible physiological significance of these observations is discussed.

Adenylyl Cyclases

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies.

BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19 bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Humans

A Cis-Regulatory Duplication in a Hox Hotspot Implicated in Mimetic Convergence in the Bumble Bee Bombus flavifrons.

Several species of North American bumble bees spanning the Pacific Coastal and Rocky Mountain regions converge onto distinct mimetic abdominal colour forms for each region by switching abdominal coloration from black to red. Previous genome-wide association studies (GWAS) of red and black transitions in two mimics (Bombus melanopygus and Bombus vancouverensis) revealed that black forms were generated by independently deleting a portion of the same cis-regulatory region near the Hox gene Abdominal-B (Abd-B). Here, we test the genetic basis of these mimetic colour forms in a third co-mimic, Bombus flavifrons, that has continuous variation in red and black that is shifted posteriorly one segment compared to its co-mimics. Using genome-wide association of red and black forms, we identified a structural variant <&#x2009;50&#x2009;bp away from the deletions in B. melanopygus and B. vancouverensis that was strongly associated with the colour phenotype. Sequencing across mimicry zones and closely related taxa revealed that all red forms of B. flavifrons and monomorphic red close relative Bombus centralis have a 319&#x2009;bp tandem duplication at this locus that has extensive modification to the duplicated copy. Black forms of B. flavifrons from the Cascades also have this duplication but without the modifications, while black forms in the western Rockies mostly lack this duplication, similar to ancestral black forms. This suggests independent mechanisms may regulate the black phenotypes in different populations and that ancestral sorting of variation and/or adaptive introgression generated these phenotypes. This study strengthens support for this Abd-B cis-regulatory region being a hotspot for regulating abdominal coloration in bumble bees, and features the role of regulatory region duplication in creating novel phenotypes.

Animals

Alpha 2-macroglobulin is not an acute-phase protein in the rat testis.

Earlier studies from this laboratory have shown that Sertoli cells actively synthesize and secrete a nonspecific protease inhibitor in vitro; N-terminal sequence analysis, subunit structural analysis, and other biological studies revealed that this protein is the homolog of serum alpha 2-macroglobulin. We have now quantified the relative distribution of alpha 2-macroglobulin in the reproductive compartments and their comparison with nonreproductive organs. In serum and all nonreproductive tissues examined, the concentration of alpha 2-macroglobulin progressively decreased with advancing age. However, in both the testis and epididymis, the levels of this protein increased with the age of the animals. Serum alpha 2-macroglobulin levels were consistently higher than those in any other tissues until 60 days when the concentrations of this protein were the highest in the epididymis. The distribution of alpha 2-macroglobulin in various nonreproductive tissues from female rats was similar to that observed for male rats in that its levels tended to decrease with age. However, uterine levels of alpha 2-macroglobulin increased progressively with advancing age, whereas ovarian levels of alpha 2-macroglobulin remained relatively stable with an increase in animal age. As serum alpha 2-macroglobulin is an acute-phase protein in the rat, the response of this protein in the testis to induced inflammation was examined. The concentration of alpha 2-macroglobulin in serum rose about 150-fold after injection of fermented yeast. By contrast, the levels of this protein in rete testis fluid, which is derived exclusively from seminiferous fluid, did not change in response to inflammation. These results suggest that there might be distinctive mechanisms that regulate this protein in the systemic circulation vs. the microenvironment behind the blood-testis barrier in the seminiferous epithelium.

Acute-Phase Proteins

Expression, purification and characterization of secreted recombinant human insulin-like growth factor-I (IGF-I) and the potent variant des(1-3) IGF-I in Chinese hamster ovary cells.

Recombinant human insulin-like growth factor-I (hIGF-I) and a biologically potent variant lacking the N-terminal tripeptide (des(1-3)IGF-I) were produced from transfected Chinese hamster ovary cells. The constructs encoding the signal peptide, sequence of the mature peptide and a C-terminal extension peptide were expressed under the control of a Rous sarcoma virus promoter. Successfully transfected clones secreting correctly processed recombinant hIGF-I or des(1-3)IGF-I were selected by their secretion of IGF-I-like activity into the culture medium. The recombinant peptides were purified to homogeneity as assessed by high-performance liquid chromatography and N-terminal sequence analysis. The purified recombinant peptides exhibited biological potencies equivalent to authentic IGF-I and des(1-3)IGF-I respectively.

Amino Acid Sequence

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids