PubMed HealthSearch

SEARCH · PubMed Health

Results for “data mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models

Theoretical models for determining 222Rn and 220Rn progeny levels in Canadian underground U mines--a comparison with experimental data.

Use has been made of several theoretical models to predict radiation levels in underground U mines. The models used are the Evans model, the Thomas-Epps mine model, the isolated mine model, the mine tunnel with no air flow and the mine tunnel with air flow (Beckman and Holub). Calculations based on the above models have been extended to include 220Rn gas and its progeny, a common occurrence in some Canadian U mines. Theoretical predictions include 222Rn and 220Rn progenies working levels and concentrations, as well as some ratios of great practical interest. The main differences between the models are pointed out and comparison is made with experimental data gathered during the last 3 yr in several Canadian underground U mines. In general terms, the mine models reported in the literature and presented here are neither satisfactory nor clearly distinguishable for practical application on the basis of the experimental data so far collected. The reason for such lack of agreement lies mainly in the unrealistic assumptions on which the models are based. Compounded is the gross oversimplification of quite complex dynamic situations encountered in actual practice. Deficiencies inherent to the models are noted and suggestions to improve the applicability of mine models to practical situations are indicated.

Air Pollutants

Metabolic Dysregulation of the Lysophospholipid/Autotaxin Axis in the Chromosome 9p21 Gene SNP rs10757274.

BACKGROUND: Common chromosome 9p21 single nucleotide polymorphisms (SNPs) increase coronary heart disease risk, independent of traditional lipid risk factors. However, lipids comprise large numbers of structurally related molecules not measured in traditional risk measurements, and many have inflammatory bioactivities. Here, we applied lipidomic and genomic approaches to 3 model systems to characterize lipid metabolic changes in common Chr9p21 SNPs, which confer ≈30% elevated coronary heart disease risk associated with altered expression of ANRIL, a long ncRNA. METHODS: Untargeted and targeted lipidomics was applied to plasma from NPHSII (Northwick Park Heart Study II) homozygotes for AA or GG in rs10757274, followed by correlation and network analysis. To identify candidate genes, transcriptomic data from shRNA downregulation of ANRIL in HEK-293 cells was mined. Transcriptional data from vascular smooth muscle cells differentiated from induced pluripotent stem cells of individuals with/without Chr9p21 risk, nonrisk alleles, and corresponding knockout isogenic lines were next examined. Last, an in-silico analysis of miRNAs was conducted to identify how ANRIL might control lysoPL (lysophosphospholipid)/lysoPA (lysophosphatidic acid) genes. RESULTS: Elevated risk GG correlated with reduced lysoPLs, lysoPA, and ATX (autotaxin). Five other risk SNPs did not show this phenotype. LysoPL-lysoPA interconversion was uncoupled from ATX in GG plasma, suggesting metabolic dysregulation. Significantly altered expression of several lysoPL/lysoPA metabolizing enzymes was found in HEK cells lacking ANRIL. In the vascular smooth muscle cells data set, the presence of risk alleles associated with altered expression of several lysoPL/lysoPA enzymes. Deletion of the risk locus reversed the expression of several lysoPL/lysoPA genes to nonrisk haplotype levels. Genes that were altered across both cell data sets were DGKA, MBOAT2, PLPP1, and LPL. The in-silico analysis identified 4 ANRIL-regulated miRNAs that control lysoPL genes as miR-186-3p, miR-34a-3p, miR-122-5p, and miR-34a-5p. CONCLUSIONS: A Chr9p21 risk SNP associates with complex alterations in immune-bioactive phospholipids and their metabolism. Lipid metabolites and genomic pathways associated with coronary heart disease pathogenesis in Chr9p21 and ANRIL-associated disease are demonstrated.

Chromosomes, Human, Pair 9

Lung carcinoma by histologic type in coal miners.

Histologic types of lung carcinoma were studied in 171 coal miners in the National Coal Workers' Autopsy Study. These miners had an average underground mining tenure of 29 +/- 14 years and an average smoking history of 31 +/- 23 pack-years. The proportion of carcinomas by cell type were: squamous cell carcinoma, 30%; adenocarcinoma, 27%; small-cell undifferentiated carcinoma, 26%; large-cell undifferentiated carcinoma, 9%; and other carcinomas, 8%. More tumors were observed in the right lung and in the upper lobes of both lungs than in the left lung and in the lower lobes of both lungs, respectively. The majority of the tumors were centered on cartilaginous airways (81%) as compared with the peripheral regions of the lung (19%). Squamous cell carcinomas predominated in the older miners and in larger airways. Adenocarcinomas were more common in the peripheral lung. No significant interaction was demonstrated between cell type and years of underground mining. The data indicate that lung carcinoma in coal miners differs little in its pathologic features from men in the general population who smoke cigarettes. No effect of coal mine dust exposure on lung carcinoma histogenesis was demonstrated.

Adenocarcinoma

[Evaluation of effects of dust prevention in the principal tungsten mines in Jiangxi].

An evaluation of the effects of dust prevention in the Jiangxi tungsten mines has been carried out. The rate of silicosis morbidity in most mines was under 1%. Up to 1983, the rate in individual mines is 1.95%. According to the data from those mines, the forecasting of cumulative probability of morbidity of mine workers having been in contact with dust for 30 years is up to 7.5%. From those data, the authors suggest that the maximum permissible concentration of dust should be 1.0 mg/m3 in the mines with concentration of silicon dioxide dust over 70%.

Dust

MADCAP: isolation of novel nAb-naïve AAV capsids from metagenomic data.

UNLABELLED: Gene therapy using adeno-associated virus (AAV) vectors offers promising treatment for genetic disorders, but significant limitations restrict clinical application. Current AAV serotypes exhibit strong liver tropism and require high doses for extra-hepatic targeting, and pre-existing antibodies (NAbs) exclude up to 50% of potential patients. Evolutionarily distant isolates can evade neutralization but typically transduce human tissues poorly and require extensive engineering. We developed MADCAP (Metagenomic AAV Discovery and Capsid Annotation Pipeline) to systematically mine metagenomic data for functional, clinically relevant AAV capsids. We hypothesized that these sources might contain capsids that do not circulate widely in humans, can transduce human cells, and avoid neutralization. We screened 4.2 million metagenomic samples and identified 139 novel AAV capsid isolates which were tested for viral capsid assembly, viability, neutralization evasion, and tissue transduction in non-human primates. While natural serotypes (AAV1, AAV2, AAV9) were neutralized at low dilutions of pooled human immunoglobulin (IVIG), 68% of tested MADCAP capsids exhibited minimal to undetectable neutralization even at supra-physiological IVIG concentrations. Systemically delivered MADCAP capsids effectively transduced multiple clinically relevant tissues in non-human primates. Two capsids, MC46 and MC55, demonstrated improved CNS tropism compared to AAV9 while maintaining comparable production yields. In passive transfer studies, MC46 retained full transduction efficiency in the presence of human antibodies, while AAV9 transduction was completely lost. This work establishes metagenomic mining as a powerful tool for accelerating AAV capsid discovery, identifying isolates with favorable tissue tropisms and resistance to broadly neutralizing antibodies. IMPORTANCE: This work provides proof of concept that potentially clinically relevant AAVs can be isolated from metagenomic data. Our findings lay the groundwork for accelerated discovery of AAV capsids which could potentially increase the accessibility and effectiveness of AAV gene therapy.

AAV

Prevalence of pneumonoconiosis among coal and heavy metal miners in Zimbabwe.

No prevalence data on pneumonoconiosis among Zimbabwe's 30,000 miners have been available. Passage of a 1984 law requiring examination of all miners has provided a data base to assess this, but the records had not been previously evaluated or stored in a manner to facilitate this. In 1988 we developed a strategy to utilize the existing records to estimate cross-sectional rates rapidly. In this report, we describe the approach and demonstrate high rates of simple pneumonoconiosis among long-term workers in the coal, nickel, copper, and gold mines. These data, though limited, provide a rationale for more detailed investigations in these workforces and an impetus to establish an ongoing surveillance plan for the nation's miners.

Analysis of Variance

Modtector: ultra-fast modification signal mining on mapped sequencing reads.

SUMMARY: Existing tools for RNA epitranscriptomic modification and structural signal analysis are often fragmented, inefficiency, and limited to single signal types. We developed Modtector, an unified tool for extracting mutation and reverse-transcription stop signals from aligned sequencing reads. By using a "count-then-correct" strategy, Modtector reduces computational complexity and enables efficient dual-signal analysis. It achieves multi-fold speedups on large-genome and high-coverage datasets, including completing HEK293 22G data analysis in 5 minutes, and show strong scalability on single-cell datasets with speedups exceeding 50-fold. AVAILABILITY: The source code is available at GitHub (https://github.com/TongZhou2017/modtector) and Crates.io (https://crates.io/crates/modtector). The archived source-code snapshot used in this study is available at Zenodo (DOI: 10.5281/zenodo.20967747), corresponding to GitHub commit 7c60e9d. Workflow examples, datasets, and analysis scripts are available at Zenodo (DOI: 10.5281/zenodo.17316476 and 10.5281/zenodo.18523297).

Humans

ConceptDrift: leveraging spatial, temporal and semantic evolution of biomedical concepts for hypothesis generation.

MOTIVATION: Hypothesis generation is a fundamental problem in biomedical text mining that aims to generate ideas that are new, interesting, and plausible by discovering unexplored links between biomedical concepts. Despite significant advances made by existing approaches, they do not fully leverage the evolutionary properties of biomedical concepts. This is limiting because scientific knowledge continually evolves over time, with new facts being added and old ones becoming obsolete. Thus, it is crucial to capture the evolutionary properties of biomedical concepts from multiple perspectives (e.g. spatial, temporal, and semantic) to generate hypotheses that reflect the up-to-date information landscape of the biomedical domain. RESULTS: We introduce a novel framework, ConceptDrift, that models the hypothesis generation task as a sequence of temporal graphlets and simultaneously encodes spatial, temporal, and semantic change. Unlike existing approaches that treat these dimensions independently, ConceptDrift is the first to provide a holistic understanding of concept evolution by integrating them into a unified framework. Grounded in the theories of the Distributional Hypothesis and Conceptual Change, our method adapts these principles to the unique challenges of large-scale biomedical literature. We conduct extensive experiments across multiple datasets and demonstrate that ConceptDrift consistently outperforms state-of-the-art baselines in generating accurate and meaningful hypotheses. Our framework shows immediate practical benefits for web-based literature mining tools in life sciences and biomedicine, offering more robust and predictive feature representations. AVAILABILITY AND IMPLEMENTATION: https://github.com/amir-hassan25/ConceptDrift (DOI: 10.6084/m9.figshare.29975476).

Semantics

Systematic mining and quantification reveal the dominant contribution of non-HLA variations to acute graft-versus-host disease.

Human leukocyte antigen (HLA) disparity between donors and recipients is a key determinant triggering intense alloreactivity, leading to a lethal complication, namely, acute graft-versus-host disease (aGVHD), after allogeneic transplantation. Moreover, aGVHD remains a cause of mortality after HLA-matched allogeneic transplantation. Protocols for HLA-haploidentical hematopoietic cell transplantation (haploHCT) have been established successfully and widely applied, further highlighting the urgency of performing panoramic screening of non-HLA variations correlated with aGVHD. On the basis of our time-consecutive large haploHCT cohort (with a homogenous discovery set and an extended confirmatory set), we first delineated the genetic landscape of 1366 samples to quantitatively model aGVHD risk by assessing the contributions of HLA and non-HLA genes together with clinical factors. In addition to identifying multiple loss-of-function (LoF) risk variations in non-HLA coding genes, our data-driven study revealed that non-HLA genetic variations, independent of HLA disparity, contributed the most to the occurrence of aGVHD. This unexpected major effect was verified in an independent cohort that received HLA-identical sibling HCT. Subsequent functional experiments further revealed the roles of a representative non-HLA LoF gene and LoF gene pair in regulating the alloreactivity of primary human T cells. Our findings highlight the importance of non-HLA genetic risk in the new era of transplantation and propose a new direction to explore the immunogenetic mechanism of alloreactivity and to optimize donor selection strategies for allogeneic transplantation.

Humans

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs

Mining Stored-Specimen Studies for Information about Cancer Natural History.

The advent of new multicancer early detection tests and publication of early diagnostic results have generated expectations of clinical benefit from multicancer screening. The clinical benefit of a cancer screening test depends critically on disease natural history, which is typically learned from prospective screening studies. Retrospective studies of stored blood specimens are important in learning about a test's preclinical diagnostic performance but have rarely been used to infer natural history. The extent to which these studies might be harnessed to also learn natural history is discussed in the context of an article in this issue that infers the combined natural history of a range of cancers targeted by a multicancer early detection test using a case-control subsample of specimens from a large cohort study. The critical question concerns the identifiability of key transition rates in multistate models of natural history alongside state-specific sensitivities. The article suggests that these parameters are estimable within a Bayesian framework that leverages prior information about test sensitivity from diagnostic studies. We offer a heuristic discussion of identifiability in this setting and encourage formal study to determine the extent to which models with varying degrees of complexity may be learned from stored-specimen studies. See related article by Dai et al., p. 1535.

Humans

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus

An approach to the characterization of silica exposure in U.S. industry.

Quantitative evaluation of worker exposure to silica in nine Standard Industrial Classification (SIC) codes was conducted, using data derived from OSHA compliance inspections, in order to assess the silica exposure problem in the U.S. The nine SICs studied were those in which OSHA inspections were concentrated. They include: construction; chemical manufacture; stone, glass, and clay manufacturing; primary metal industries; metal fabrication; machinery; transportation; and miscellaneous manufacturing industries. High exposures to silica were documented in each industry, with the number of test samples over the permissible exposure limit ranging from 14% (aluminum foundries) to 73% (pottery). An estimation is made that 24,889 workers employed in ferrous and nonferrous foundries are at risk of silica-related pulmonary effects. The data developed in this analysis also indicate the need to investigate certain industries that had high exposures but few inspections. The limitations of the data base for estimating the scope of the silica problem, including lack of data on mining and milling, are discussed. We conclude that exposure to silica represents a continuing and significant problem in a number of U.S. industries.

Air Pollutants, Occupational

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models

Lit-OTAR framework for extracting biological evidences from literature.

SUMMARY: The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION: The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.

Drug Discovery