PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “variant calling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

The fetal and neonatal brain protein neuronatin protects PC12 cells against certain types of toxic insult.

The protein neuronatin is expressed in the nervous system of the fetus and neonate at a much higher level than in the adult. Its function is unknown. As a result of variable splicing, neuronatin mRNA exists in two forms, alpha and beta. Wild type PC12 cells express neuronatin-alpha. We have isolated a PC12 variant, called 1.9, that retains many of the neuron-like properties of wild type PC12 cells, but it does not express neuronatin and it exhibits markedly increased sensitivity to the toxic effects of nigericin, rotenone and valinomycin. Pretreatment of the 1.9 cells with alpha-methyltyrosine, which inhibits dopamine synthesis, had little effect on the cells' sensitivity to nigericin, rotenone or valinomycin indicating that dopamine-induced oxidative stress was not involved in the toxicity of these compounds. However, flattened cell subvariants of the 1.9 cells, which do not have any neuron-specific characteristics, did not exhibit increased sensitivity to nigericin indicating that some neuronal characteristic of the 1.9 cells contributed to the toxicity of nigericin. After the neuronatin-beta gene was transfected into and expressed in the 1.9 cells, they regained wild type PC12 levels of resistance to nigericin, rotenone and valinomycin. These studies suggest that the function of neuronatin during development could be to protect developing cells from toxic insult occurring during that period.

Animals↗

The structure of an engineered domain-swapped ribonuclease dimer and its implications for the evolution of proteins toward oligomerization.

BACKGROUND: Domain swapping has been proposed as a mechanism that explains the evolution from monomeric to oligomeric proteins. Bovine and human pancreatic ribonucleases are monomers with no biological properties other than their RNA cleavage ability. In contrast, the closely related bovine seminal ribonuclease is a natural domain-swapped dimer that has special biological properties, such as cytotoxicity to tumour cells. Several recombinant ribonuclease variants are domain-swapped dimers, but a structure of this kind has not yet been reported for the human enzyme. RESULTS: The crystal structure at 2 A resolution of an engineered ribonuclease variant called PM8 reveals a new kind of domain-swapped dimer, based on the change of N-terminal domains between the two subunits. The swapping is fastened at both hinge peptides by the newly introduced Gln101, involved in two intermolecular hydrogen bonds and in a stacking interaction between residues of different chains. Two antiparallel salt bridges and water-mediated hydrogen bonds complete a new interface between subunits, while the hinge loop becomes organized in a 3(10) helix structure. CONCLUSIONS: Proteins capable of domain swapping may quickly evolve toward an oligomeric form. As shown in the present structure, a single residue substitution reinforces the quaternary structure by forming an open interface. An evolutionary advantage derived from the new oligomeric state will fix the mutation and favour others, leading to a more extended complementary dimerization surface, until domain swapping is no longer necessary for dimer formation. The newly engineered swapped dimer reported here follows this hypothetical pathway for the rapid evolution of proteins.

Amino Acid Sequence↗

Helical structures of poly(D-L-peptides). A conformational energy analysis.

Conformational energy calculations are reported for a number of possible helical structures of poly(D-L-peptides): the alpha helix, two single-stranded piDL, and five double-stranded pipiDL helices. For a poly(D-alanine-L-alanine) sequence, the energies of the various helices are found to differ by less than 1 kcal/(mol residue). For some helices (especially the piDL ones) two structural variants are predicted. These variants, called "goniomers", are characterized by reversed sequences of conformational angles but have the same screw sense and similar helical parameters. A biological implication of these goniomers is suggested, and their usefulness as a critical test for energy calculations is considered.

Alanine↗

Whole-Genome Sequencing of 54 Dengchuan Cattle (Bos taurus) from Southwest China.

Domestic cattle (Bos taurus) play a significant role in human society as they provide abundant food resources and contribute to the development of agriculture and traditional culture. Dengchuan cattle, a local breed from Yunnan, Southwest China, are known for their high-quality milk and are at risk of extinction due to crossbreeding. To preserve the superior genetic resources of Dengchuan cattle, this study conducted whole-genome sequencing of 54 Dengchuan cattle using blood DNA samples, generating approximately 3.56 TB of clean data with an average sequencing depth of 32.78X. The sequencing data were aligned to the bovine reference genome (ARS-UCD2.0), achieving an average alignment rate of 99.85%. A total of 9,950,420 SNPs and 2,476,207 indels were detected using variant calling workflow. These data were utilized to characterize genomic profile of this unique cattle breed. The data generated in this study can be incorporated into the global cattle genomic diversity database, providing valuable information for comparative studies on cattle.

Animals↗

Enhanced reaction with Vicia graminea lectin and exposed terminal N-acetyl-D-glucosaminyl residues on a sample of human red cells with Hb M-Hyde Park.

A sample of polyagglutinable red cells was obtained from a healthy individual (group O, N) possessing a hemoglobin (Hb) variant called Hb M-Hyde Park. The sialic acid content of the individual's red cells is 90 percent of normal, and his cells are agglutinated by monoclonal but not lectin anti-Tn, a panel of lectins specific for N-acetylgalactosamine (or galactose), and N-acetylglucosamine. Enhanced agglutination reactions were obtained with Vicia graminea, Ulex europaeus, and human anti-I and -i. Using various enzyme treatments and different methods of labeling cell surface components, two defective cell membrane sites have been identified: one associated with the O-linked oligosaccharides on sialoglycoproteins and the other associated with exposed N-acetylglucosaminyl residues located on membrane components of apparent molecular weights 88,000 to 130,000 and 46,000 to 73,000 (probably the Band 3 and Band 4.5 regions, respectively).

Acetylglucosamine↗

Growth hormones. 1. Polymorphism (minireview).

Pituitary growth hormone (GH) is not a single molecular species, but a whole set of similar molecules, the individual specific characteristics of which constitute the polymorphism of this hormone. The present paper deals mainly with various forms of human GH, called "variants", and touches on this polymorphism in other species as well. The 22 K variant (MW = 22,000 daltons) is the predominant form of GH to which all other variants are compared as to chemical structure and biological effect. These variants are classified into two large groups: 1) mass variants, the molecular weight of which is modified in comparison with that of 22 K; these can be subdivided into aggregated and non-aggregated forms, and 2) charge variants with modified electrophoretic mobility. Outside this classification are entities which are not yet well known; these include bioinactive GH, correctly detected by RIA but deprived of biological activity or, on the contrary, strongly bioactive GH lacking immunoreactivity and consequently difficult to study. Another outsider is the SV-hGH-2 variant encoded from a gene different from the hGH-N gene normally coding for the other variants. In this case, the product could be considered a true isohormone of 22 K and no longer a variant. The pituitary expression of this gene has never been evidenced to date, but according to recent data, it could be expressed at the placental level and be implicated in human placental growth hormone (hPGH) synthesis. hPGH is a newly-found GH in pregnant women which takes over pituitary GH from the 25th week onwards. After the GH molecules are released by the pituitary in the blood stream, they are partially taken up and carried by binding proteins. The physiological role of this phenomenon could be the setting up of a GH reservoir and a GH sparing process since the metabolic clearance rate of the complex GH-binding protein is slower than that of free GH, thus increasing the biological half-life of the hormone.

Animals↗

Specificity of prion assembly in vivo. [PSI+] and [PIN+] form separate structures in yeast.

The yeast prions [PSI+] and [PIN+] are self-propagating amyloid aggregates of the Gln/Asn-rich proteins Sup35p and Rnq1p, respectively. Like the mammalian PrP prion "strains," [PSI+] and [PIN+] exist in different conformations called variants. Here, [PSI+] and [PIN+] variants were used to model in vivo interactions between co-existing heterologous amyloid aggregates. Two levels of structural organization, like those previously described for [PSI+], were demonstrated for [PIN+]. In cells with both [PSI+] and [PIN+] the two prions formed separate structures at both levels. Also, the destabilization of [PSI+] by certain [PIN+] variants was shown not to involve alterations in the [PSI+] prion size. Finally, when two variants of the same prion that have aggregates with distinct biochemical characteristics were combined in a single cell, only one aggregate type was propagated. These studies demonstrate the intracellular organization of yeast prions and provide insight into the principles of in vivo amyloid assembly.

Amyloid↗

The IVS-II-1 (G-->a) beta0-thalassemia mutation in cis with HbA2-Troodos [delta116(G18)Arg-->Cys (CGC-->TGC)] causes a complex prenatal diagnosis in an Iranian family.

The beta-thalassemia (thal) minor phenotypes with normal Hb A2 levels and decreased MCV and MCH values are relatively rare beta-thal traits. Here, we describe a family with normal Hb A2 and decreased MCV and MCH levels. Amplification refractory mutation system-polymerase chain reaction (ARMS-PCR) revealed the IVS-II-1 (G-->A) mutation in the beta-globin gene of the proband and her father. Direct sequencing of the gamma-globin gene of the proband and her father also revealed a previously reported variant called Hb A2-Troodos [gamma116(G18)Arg-->Cys] [in cis with the IVS-II-1 (G-->A) beta0-thal mutation]. This is the first case report of Hb A2-Troodos in association with the beta0 IVS-II-1 mutation. Reduced Hb A2 expression by a concomitant Hb A2 beta-thal in cis or trans, may cause problems in carrier diagnostics, and eventually in genetic counseling and prenatal diagnosis when insufficient molecular analyses are performed.

Adult↗

Optimal amnesic probabilistic automata or how to learn and classify proteins in linear time and space.

Statistical modeling of sequences is a central paradigm of machine learning that finds multiple uses in computational molecular biology and many other domains. The probabilistic automata typically built in these contexts are subtended by uniform, fixed-memory Markov models. In practice, such automata tend to be unnecessarily bulky and computationally imposing both during their synthesis and use. Recently, D. Ron, Y. Singer, and N. Tishby built much more compact, tree-shaped variants of probabilistic automata under the assumption of an underlying Markov process of variable memory length. These variants, called Probabilistic Suffix Trees (PSTs) were subsequently adapted by G. Bejerano and G. Yona and applied successfully to learning and prediction of protein families. The process of learning the automaton from a given training set S of sequences requires theta(Ln2) worst-case time, where n is the total length of the sequences in S and L is the length of a longest substring of S to be considered for a candidate state in the automaton. Once the automaton is built, predicting the likelihood of a query sequence of m characters may cost time theta(m2) in the worst case. The main contribution of this paper is to introduce automata equivalent to PSTs but having the following properties: Learning the automaton, for any L, takes O (n) time. Prediction of a string of m symbols by the automaton takes O (m) time. Along the way, the paper presents an evolving learning scheme and addresses notions of empirical probability and related efficient computation, which is a by-product possibly of more general interest.

Algorithms↗

A transmembrane form of the prion protein contains an uncleaved signal peptide and is retained in the endoplasmic Reticulum.

Although there is considerable evidence that PrP(Sc) is the infectious form of the prion protein, it has recently been proposed that a transmembrane variant called (Ctm)PrP is the direct cause of prion-associated neurodegeneration. We report here, using a mutant form of PrP that is synthesized exclusively with the (Ctm)PrP topology, that (Ctm)PrP is retained in the endoplasmic reticulum and is degraded by the proteasome. We also demonstrate that (Ctm)PrP contains an uncleaved, N-terminal signal peptide as well as a C-terminal glycolipid anchor. These results provide insight into general mechanisms that control the topology of membrane proteins during their synthesis in the endoplasmic reticulum, and they also suggest possible cellular pathways by which (Ctm)PrP may cause disease.

Animals↗

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics↗

nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data.

MOTIVATION: Pacific Biosciences (PacBio) single-molecule, long-read sequencing enables whole genome annotation and the characterization of 20 complex repetitive repeat regions, especially relevant to neurodegenerative diseases, through their PureTarget panel. Long-read whole-genome sequencing (WGS) also allows for the detection of structural variants that would be difficult to detect with traditional short-read sequencing. However, the raw unaligned Binary Alignment Map data need to be processed before analysis. There is a need for an intuitive comprehensive bioinformatic pipeline that can analyze these data. RESULTS: We present nf-core/pacvar, a comprehensive pipeline for analyzing both PacBio single-molecule PureTarget and WGS data that demultiplexes and parallelizes pre-processing, variant calling and repeat characterization. nf-core/pacvar is compatible with little configuration and has few dependencies. This pipeline enables rapid end-to-end, parallel processing of PacBio single-molecule whole genome and targeted repeat expansion sequencing. AVAILABILITY AND IMPLEMENTATION: nf-core/pacvar is available on nf-core website (https://nf-co.re/pacvar/) and on github (https://github.com/nf-core/pacvar) under MIT License (DOI: 10.5281/zenodo.14813048).

Software↗

wgatools: an ultrafast toolkit for manipulating whole-genome alignments.

SUMMARY: With the rapid development of long-read sequencing technologies, the era of individual complete genomes is approaching. We have developed wgatools, a cross-platform, ultrafast toolkit that supports a range of whole-genome alignment formats, offering practical tools for conversion, processing, evaluation, and visualization of alignments, thereby facilitating population-level genome analysis and advancing functional and evolutionary genomics. AVAILABILITY AND IMPLEMENTATION: wgatools supports diverse formats and can process, filter, and statistically evaluate alignments, perform alignment-based variant calling, and visualize alignments both locally and genome-wide. Built with Rust for efficiency and safe memory usage, it ensures fast performance and can handle large datasets consisting of hundreds of genomes. wgatools is published as free software under the MIT open-source license, and its source code is freely available at https://github.com/wjwei-handsome/wgatools and https://zenodo.org/records/14882797.

Software↗

FuNTB: a functional network clustering tool for the analysis of genome-wide genetic variants in Mycobacterium tuberculosis.

MOTIVATION: Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mtb), still claims around 1.25 million lives each year. The growing threat of drug resistance-often driven by single‑nucleotide polymorphisms (SNPs) in Mtb genomes underscores the need for high‑quality genomic data and powerful bioinformatics tools. We present FuNTB, a python‑based pipeline that detects non‑synonymous SNPs in Mtb and builds functional network clusters to reveal genotype-phenotype relationships. RESULTS: FuNTB profiles non‑synonymous SNPs at the gene level across user‑defined phenotypes, pinpointing both shared and unique mutations. It ingests annotated Variant Call Format (VCF) files or MTBseq outputs and merges them with clinical metadata to produce network‑XML files compatible with Cytoscape and Gephi. When applied to the CRyPTIC Mtb collection, FuNTB rapidly recovered established resistance genes and surfaced novel candidates, validating its utility for mapping genotype-phenotype associations. AVAILABILITY AND IMPLEMENTATION: FuNTB is implemented in Python 3.8+ and is freely available under the MIT license at https://doi.org/10.5281/zenodo.15399917.

Mycobacterium tuberculosis↗

AncestryGeni: a novel genetic ancestry classification pipeline for small and noisy sequence data.

MOTIVATION: Efforts to address health disparities are often limited by the lack of robust computational tools for inferring genetic ancestry by calculating an individual's genetic similarity to continental groups. We have already shown that a preferred alternative to self-described race is using ancestry-informative markers (AIMs) that can be classified into ancestral components and used to estimate their similarity to those of known populations to identify continental groups. However, real-world genomic data can present challenges, including limited availability of germline DNA, a small number of AIMs for each sample, and the use of different variant calling software, limiting the application of existing solutions. RESULTS: Here, we describe a novel supervised machine-learning tool AncestryGeni, which infers genetic ancestry for samples with even a hundred markers and is applicable to any genomic data, including whole exome sequencing (WES) and RNA sequencing (RNA-Seq) data. Applying AncestryGeni to a real-world genomic dataset obtained from the Multiple Myeloma Research Foundation (MMRF) CoMMpass study, we show that it is more accurate than the commonly used FastNGSadmix when using nonstandard genomic material. We also demonstrate that when using AncestryGeni, the tumor-derived sequence obtained from WES and RNA-Seq can be a robust data source to accurately estimate an individual's genetic similarity to a continental group. AVAILABILITY AND IMPLEMENTATION: AncestryGeni pipeline is available at https://github.com/eelhaik/AncestryGeni/tree/main.

Humans↗

insilicoSV: a flexible grammar-based framework for structural variant simulation and placement.

SUMMARY: Structural variants (SVs) are key drivers of genetic variation and disease in the genome. Their discovery remains challenging, however, in large part due to the scarcity of validated SV callsets and comprehensive benchmarks, which are essential for method development and evaluation. The growing number of data-driven learning-based approaches for SV discovery, in particular, requires large, diverse, and well-balanced training datasets to achieve reliable performance. To address this need, SV simulation has served as a key tool for assessing method performance and training SV models. However, existing SV simulators only support a fixed and limited set of SV classes and do not provide fine-grained control over the placement of SVs within specific contexts of the genome. Here we present insilicoSV, a versatile framework for SV simulation, which models SVs using a simple and flexible grammar, allowing users to easily define standard and custom arbitrary genome rearrangements, as well as encode genome placement constraints. This design allows insilicoSV to naturally support new and bespoke SV types, such as the complex rearrangements of cancer genomes. In addition to grammar-based modeling, insilicoSV provides built-in support for 26 predefined SV types, placement of user-provided SVs, small variant simulation, streamlined workflows for the simulation of genome evolution and genome mixtures, read simulation, alignment, and visualization. These features enable the creation of comprehensive genomic datasets for a variety of downstream applications, such as in-depth benchmarking of alignment and variant calling methods, as well as training of data-driven learning-based approaches for SV detection. AVAILABILITY AND IMPLEMENTATION: insilicoSV is available under the MIT license at https://github.com/PopicLab/insilicoSV and https://doi.org/10.5281/zenodo.17402009.

Software↗

ALPINE: a scalable pipeline for comprehensive classification of gene-editing outcomes from long-read amplicon sequencing.

SUMMARY: CRISPR genome editing has enabled precise genetic modification for gene and cell therapies, but edits often produce heterogeneous on-target outcomes, including homology-directed repair (HDR) knock-ins, DNA repair template integrations, and structural variants. Existing tools are frequently limited to short reads or lack viral vector-specific integration categories needed for therapeutic development. Here, we present ALPINE (Amplicon Long-read Pipeline for INtegration Evaluation), a scalable and reproducible pipeline for classifying and quantifying gene-editing outcomes from long-read amplicon sequencing supporting both PacBio HiFi and Oxford Nanopore platforms. ALPINE classifies reads into 10+ categories, including DNA repair vector integration subtypes, and performs variant calling near the gene-edited site with batch, multi-sample reporting. Uniquely, ALPINE can distinguish between cells treated with multiple DNA repair vectors and identify distinct molecular features, such as inverted terminal repeats (ITRs), enabling comprehensive characterization of complex gene editing outcomes. Dual-target benchmarking on simulated datasets demonstrated high accuracy for transgene integration events. Independent validation on public crosslinked-HDR dataset confirmed ALPINE's integration detection capabilities, and application to edited T cell samples demonstrated comprehensive gene-editing outcome profiling. AVAILABILITY: ALPINE is available under MIT license at https://github.com/Maggi-Chen/ALPINE and https://doi.org/10.5281/zenodo.20272510. All analysis scripts and visualization code used in this manuscript are available at https://github.com/Maggi-Chen/ALPINE-manuscript-analysis. Simulated datasets are deposited at Zenodo (https://doi.org/10.5281/zenodo.20260865). Public dataset PRJNA913199 is available through NCBI SRA.

Gene Editing↗

Nonhomologous pairing in mice heterozygous for a t haplotype can produce recombinant chromosomes with duplications and deletions.

We have investigated the structure and properties of a chromosomal product recovered from a rare recombination event between a t haplotype and a wild-type form of mouse chromosome 17. Our embryological and molecular studies indicate that this chromosome (twLub2) is characterized by both a deletion and duplication of adjacent genetic material. The deletion appears to be responsible for a dominant lethal maternal effect and a recessive embryonic lethality. The duplication provides an explanation for the twLub2 suppression of the dominant T locus phenotype. A reanalysis of previously described results with another chromosome 17 variant called TtOrl indicates a structure for this chromosome that is reciprocal to that observed for twLub2. We have postulated the existence of an inversion over the proximal portion of all complete t haplotypes in order to explain the generation of the partial t haplotypes twLub2 and TtOrl. This proximal inversion and the previously described distal inversion are sufficient to account for all of the recombination properties that are characteristic of complete t haplotypes. The structures determined for twLub2 and TtOrl indicate that rare recombination can occur between nonequivalent genomic sequences within the inverted proximal t region when wild-type and t chromosomes are paired in a linear, nonhomologous configuration.

Alleles↗