PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference mapping”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The phenotype-genotype reference map: Improving biobank data science through replication.

Population-scale biobanks linked to electronic health record data provide vast opportunities to extend our knowledge of human genetics and discover new phenotype-genotype associations. Given their dense phenotype data, biobanks can also facilitate replication studies on a phenome-wide scale. Here, we introduce the phenotype-genotype reference map (PGRM), a set of 5,879 genetic associations from 523 GWAS publications that can be used for high-throughput replication experiments. PGRM phenotypes are standardized as phecodes, ensuring interoperability between biobanks. We applied the PGRM to five ancestry-specific cohorts from four independent biobanks and found evidence of robust replications across a wide array of phenotypes. We show how the PGRM can be used to detect data corruption and to empirically assess parameters for phenome-wide studies. Finally, we use the PGRM to explore factors associated with replicability of GWAS results.

Humans

Human liver protein map: a reference database established by microsequencing and gel comparison.

This publication establishes a reference human liver protein map obtained with immobilized pH gradients. By microsequencing, 57 spots or 42 polypeptide chains were identified. By protein map comparison and matching (liver, red blood cell and plasma sample maps), 8 additional proteins were identified. The new polypeptides and previously known proteins are listed in a table and/or labeled on the protein map, thus providing a human liver two-dimensional gel database. This reference map can be used to identify protein spots on other samples such as rectal cancer biopsies.

Amino Acid Sequence

Genetic analysis of bacteriophage phi 29 of Bacillus subtilis: integration and mapping of reference mutants of two collections.

Reference mutants of Bacillus subtilis phage phi 29 of the Madrid and Minneapolis collections were employed to construct a genetic map. Suppressor-sensitive and temperature-sensitive mutants were assigned to 17 cistrons by quantitative complementation. Three-factor crosses were used to assign an unambiguous order for the 17 cistrons. Recombination frequencies determined by two-factor crosses were used to construct a linear genetic map of 24.4 recombination units. The genes were numbered sequentially from left to right (1 to 17) according to their relative map position.

Bacillus subtilis

A modular class-aware workflow for small RNA sequencing analysis using mouse sperm as a case study.

BACKGROUND: Small RNA sequencing analysis is challenging because RNA classes differ in biogenesis, sequence redundancy, genomic organization, and annotation reliability. Integrated workflows accommodating these constraints remain limited, particularly for fragment-level and cluster-level analysis. METHODS: We present a reproducible, containerized, class-aware workflow for small RNA sequencing analysis, using mouse sperm as a case study. The workflow combines standardized preprocessing with complementary annotation and quantification strategies for microRNAs (miRNAs), transfer RNA-derived small RNAs (tsRNAs), ribosomal RNA-derived small RNAs (rsRNAs), and PIWI-interacting RNA (piRNA)-enriched genomic clusters. Using sperm small RNA data from offspring of lipopolysaccharide (LPS)-exposed male mice, we compared integrated-reference mapping, multi-class annotation, fragment-level tsRNA profiling, and genome-based piRNA cluster analysis, with custom modules for locus-aware harmonization and condition-specific cluster analysis. RESULTS: Integrated-reference mapping aligned 88.17% of reads and retained 690 features after filtering. It identified 11 differentially expressed miRNAs between LPS and controls, while other classes showed limited signal. Fragment-level profiling improved tsRNA resolution. piRNA cluster analysis identified 958 control and 940 LPS clusters, with 18 control-specific and no LPS-specific clusters. CONCLUSION: This workflow supports transparent, reproducible, class-aware interpretation of small RNA sequencing data while emphasizing cautious interpretation of piRNA-enriched signals from total small RNA sequencing.

Small non-coding RNA analysis

Using Mapping-Profiles to Refine Strain-Level Metagenomic Classification.

Metagenomic classification at the strain level remains challenging due to high sequence similarity among closely related genomes, which leads to ambiguous read mappings and frequent false-positive strain detections. Reducing such errors improves the reliability of strain-level analyses, which is critical for applications such as pathogen detection. We introduce StrainRefine, a post-mapping refinement method that analyzes read-reference mapping profiles to resolve ambiguous assignments among highly similar genomes. The method represents candidate reference genomes using binary profiles that capture read-support patterns and measures similarity between references based on profile overlap. The method clusters references based on similar mapping profiles, filters weakly supported genomes, and reassigns reads to representative references, reducing redundant reporting of near-identical strains. StrainRefine substantially reduces false-positive strain detections while preserving recall and improving agreement between predicted and true abundance profiles. On large-scale metagenomic datasets, it achieves a substantially improved precision-recall balance compared with existing mapping-based approaches, with the standalone method obtaining the highest read-level classification accuracy on the most complex evaluated dataset. Unlike many strain-level tools designed for individual species, StrainRefine operates without prior assumptions about sample composition or curated species-specific reference collections, while still achieving comparable performance in single-species settings on species-specific reference databases. These results highlight mapping-profile similarity as an effective signal for improving strain-level metagenomic classification.

false-positive reduction

GRNContext: an interactive web platform for contextualized gene regulatory networks visualization across human cancers.

SUMMARY: While current Gene Regulatory Network (GRN) databases provide comprehensive reference maps of potential interactions between transcription factors and target genes, they do not specify which regulatory interactions are active within specific biological contexts. This limitation is particularly critical in cancer, where transcriptional programs are inherently tissue-specific. To address this gap, we developed GRNContext, an interactive web platform designed for the visualization, exploration, and comparative analysis of gene regulatory networks contextualized across 33 cancer types from The Cancer Genome Atlas (TCGA). Our approach uses the TFLink human reference GRN as a starting point and integrates TCGA transcriptomic profiles to infer cancer-specific regulatory activity. Regulatory relevance was assessed using complementary machine learning and statistical methods, which were unified into a consensus score to prioritize and filter the most relevant candidate regulators for each target gene. By providing both curated context-specific GRNs and a user-friendly platform, GRNContext constitutes a comprehensive and accessible resource that supports mechanistic investigations, hypothesis generation, and translational research focused on transcriptional regulation in cancer. AVAILABILITY AND IMPLEMENTATION: GRNContext is supported by all major browsers and freely available on the web at https://apps.cienciavida.org/grncontext. It is implemented as a client-server web application featuring a FastAPI backend and a React frontend utilizing Cytoscape.js for interactive network visualization, all containerized via Docker for cross-platform compatibility.

Humans

Architectural transcription factors collectively shape nuclear radial positioning of chromatin contacts.

The measurement of three-dimensional genome folding in the nucleus, mostly through Hi-C methods, is expressed as contact frequencies between genomic segments, without anchoring to physical axes of the spherical nucleus. Here, we mapped the chromatin contacts along nuclear radial axis and built radial score by factoring in contact frequencies. The chromatin high-order structures exhibit rich diversity along radial axis. Furthermore, the proximal trans contacts retrieved by radial score reveal conserved active/inactive chromatin segregation across intra- and interchromosomal interactions. Ablation of CTCF proteins disrupts chromatin loops with mild changes to chromatin radial positioning. By acutely perturbing multiple transcription factor (TF) occupancy, chromatin loop dissolutions are often accompanied by radial dissociations between two anchors. Our work provides a genome architecture reference map adhering to nuclear physical axis and suggests that multiple architectural TFs collectively shape nuclear positioning of chromatin and their contacts, with contacts serving as forces on chromatin positioning as well.

Chromatin

CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigation.

SUMMARY: Next Generation Sequencing is widely deployed in cholera-endemic regions, yet an end-to-end reproducible pipeline that unifies read QC, filtering, reference mapping, variant calling/annotation, recombination screening, and extraction of parsimony informative sites/variant codons, phylogenetic inference for downstream phylodynamic and epidemiological analyses have been lacking, slowing outbreak investigation and public health response. CholeraSeq is a high-throughput genomics pipeline for cholera genomic surveillance. It ingests consensus genomes, short read sequence data, draft assemblies, and scales seamlessly from local to cloud environments. To accelerate epidemiological context placement of new outbreak strains, we provide a curated ready-to-use core genome alignment compiled from public data, enabling flexible, fast, integration of new samples for outbreak investigations. AVAILABILITY AND IMPLEMENTATION: CholeraSeq is freely available on the GitHub platform https://github.com/CERI-KRISP/CholeraSeq. CholeraSeq is implemented in Nextflow with a modular design building upon the nf-core community standards.

Cholera

MetaStrainer: accurate reconstruction of bacterial strain genotypes from short-read metagenomic samples.

MOTIVATION: Metagenomics provides broad insights from microbial communities, but more biological relevant phenotypes are attributed to subtle changes at the strain-level rather than species. Despite development of several tools using different algorithms, resolving individual strains from short-read pair-end sequencing data remains challenging. RESULTS: Here we present MetaStrainer, a tool capable of reconstructing strain genotypes from metagenomic data. Compared with existing approaches, MetaStrainer substantially increases genotype accuracy, correctly identifies the number of strains, and accurately estimates their relative abundances. Accuracy of reconstructed genotypes is robust to choice of mapping reference. AVAILABILITY: MetaStrainer is implemented in Python 3. Source code and instructions are available on GitHub at www.github.com/lbobay/MetaStrainer and on Zenodo: 10.5281/zenodo.17872331.

Metagenomics

The mitotic, polytene, and meiotic chromosomes of Drosophila ananassae.

The mitotic chromosome complement of D. ananassae consists of four structurally distinguishable submetacentric pairs and all four have been identified with their linkage groups. For the polytene chromosome complement of six arms representing the X, second and third chromosomes, an improved reference map has been constructed and used to describe selected cytogenetically useful rearrangements. In meiotic prophase of spermatocytes, chromosomes 2 and 3 form pachytene-diplotene bivalents whose arms may be associated by chiasmata in postdiplotene stages, but the X, Y and fourth chromosomes participate in a complex multivalent. No correlation was detected between meiotic chromosome behavior and specific genes that regulate crossing over in males. In male inversion heterozygotes having high levels of genetically monitored crossing over, no unequivocal evidence was found for formation of either pachytene inversion loops or anaphase bridges and fragments.

Animals

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

The use and limitations of chiasma scoring with reference to human genetic mapping.

Human chiasma data are summarized, and some preliminary new observations in fetal oocytes are presented. Male chiasma data may give reliable estimates of genetic lengths, both for individual chromosome arms and for the total autosomal complement. Female data are as yet less accurate and give information according to chromosome group only. Movement of chiasmata before they can be reliably scored is unlikely. In both sexes, chiasmata are seen to be clustered along the length of the chromosomes, which may reflect crossingover interference and a tendency for crossingover to more often take place in certain chromosome segments; there are some indications of sex differences in these preferences.

Chromatids

Microcarcinoma of the endometrium: a mapping study with special reference to cytologic atypia in the endometrium.

In order to elucidate the basis for the development of an endometrial carcinoma, we looked for microcarcinomas measuring < 5 mm in greatest diameter, and studied their histologic characteristics and those of the neighboring endometrium. Using serial step section methods, two microcarcinomas were detected. A microcarcinoma was found in one of 14 uteri resected for atypical hyperplasia and the other was found in one of 114 uteri resected for endometrial carcinoma. The neighboring endometrium of the former was adenomatous and had atypical hyperplasia and that of the latter was atrophic and contained atypical glands characterized by cytologic atypia and not by architectural changes. The findings may suggest endometrial carcinomas to have two pathogenetic forms: a carcinoma associated with hyperplasia and occurring in premenopausal women, a second carcinoma associated with atrophic endometrium and occurring in postmenopausal women. Atypical glands in atrophic endometria may indicate that endometrial specimens from postmenopausal women should be carefully screened for cytologic atypia.

Adenocarcinoma

Similarity between average distance maps of structurally homologous proteins.

A similarity between average distance maps (Kikuchi et al., 1988a)--that is, predicted contact maps of two tertiary structurally homologous proteins--is examined. Comparisons of shapes of average distance maps (we refer to this as ADM) are made by superpositions of ADMs for two homologous proteins. Also, we compare shapes of actual contact maps for the pair of proteins. We search a optimal superposition mode of each pair of maps showing that two proteins are most similar. It is concluded that two ADMs are also similar when actual tertiary structures between two proteins show similarity. A criterion for similarity of maps is also proposed. The possibility of application of this method to detect weak homology between protein structures is discussed.

Protein Conformation