PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Ergonomics in mining.

Like other areas of occupational health and safety (OHS) ergonomics is evolving and becoming more integrated into overall work management systems. As we learn more about the complex interaction between psychosocial and physical factors in the aetiology of work-related illness and injury the more we rely on managers to 'get it right' if we are to prevent these conditions. Risks to health and safety in the mining industry posed by longer shift lengths, higher work loads, less task variation and decision latitude have not really been well researched. Heavy physical workloads and stresses are still areas of concern, but are likely to be intermittent rather than constant. Recent research confirms current thinking rather than shedding new light on the subject. The contribution of slips, trips and falls and increasing age of miners to manual handling injuries is still not clear. In some cases sedentary work and the operation of machinery has completely replaced heavy physical work. The issues of machinery design for operations and maintenance and whole-body vibration exposures when operating machines and vehicles are becoming more critical. The link between prolonged sitting, poor cab design and vibration with back and neck pain is being recognized but has yet to be addressed in any systematic way by the mining industry. On the plus side some mining companies have well-developed participative approaches to problem solving and these need to be extended to areas such as ergonomics.

Age Factors↗

Mapping of neural networks onto the memory-processor integrated architecture.

In this paper, an effective memory-processor integrated architecture, called memory-based processor array for artificial neural networks (MPAA), is proposed. The MPAA can be easily integrated into any host system via memory interface. Specifically, the MPA system provides an efficient mechanism for its local memory accesses allowed by row and column bases, using hybrid row and column decoding, which is suitable for computation models of ANNs such as the accessing and alignment patterns given for matrix-by-vector operations. Mapping algorithms to implement the multilayer perceptron with backpropagation learning on the MPAA system are also provided. The proposed algorithms support both neuron and layer level parallelisms which allow the MPAA system to operate the learning phase as well as the recall phase in the pipelined fashion. Performance evaluation is provided by detailed comparison in terms of two metrics such as the cost and number of computation steps. The results show that the performance of the proposed architecture and algorithms is superior to those of the previous approaches, such as one-dimensional single-instruction multiple data (SIMD) arrays, two-dimensional SIMD arrays, systolic ring structures, and hypercube machines.

Journal Article↗

SeqQC-former: A sequence-quality fusion framework for QC-aware review prioritization of candidate somatic SNVs in cancer genomics.

The accurate prioritization of candidate somatic single-nucleotide variants (SNVs) remains a challenge due to the substantial variability in sequencing quality across genomic loci. SeqQC-Former is a sequence-quality fusion framework that integrates the local nucleotide context with read-level quality-control (QC) covariates derived from matched tumor-normal sequencing data. This integration generates QC-aware prioritization scores for the downstream review of candidate variants. Unlike conventional variant callers, SeqQC-Former is designed not to infer biological truth but to support post-calling review and prioritization under heterogeneous sequencing conditions. The framework was trained and evaluated on a SEQC2-derived dataset comprising 89,447 candidate loci, including 1378 positive and 88,069 negative loci. In chromosome-held-out validation, which aims to reduce potential genomic-position leakage, SeqQC-Former demonstrated strong discrimination (AUROC = 0.9479; AUPRC = 0.9448), indicating good generalization to previously unseen chromosomes. Given that the SEQC2-derived labels contain QC-associated information; these results should be interpreted as an evaluation of QC-aware prioritization capability rather than an independent validation of biological variant correctness. Ablation analyses revealed that structured QC covariates provided the dominant predictive signal under the current SEQC2-derived labeling regime. SeqQC-Former achieved a significantly higher AUROC than classical machine-learning baselines, as determined by DeLong's test (p&#x202f;<&#x202f;0.01). Application to 53,164 glioblastoma variants demonstrated that external predictions were sensitive to QC scaling and threshold selection, underscoring that model outputs should be interpreted as QC-dependent prioritization scores rather than calibrated probabilities or definitive biological classifications. Overall, SeqQC-Former offers a reproducible post-calling QC-aware prioritization framework for large-scale somatic SNV review and underscores the importance of explicitly modeling sequencing-quality information when interpreting structured cancer genomics datasets.

Humans↗

Pavlovian feed-forward mechanisms in the control of social behavior.

The conceptual and investigative tools for the analysis of social behavior can be expanded by integrating biological theory, control systems theory, and Pavlovian conditioning. Biological theory has focused on the costs and benefits of social behavior from ecological and evolutionary perspectives. In contrast, control systems theory is concerned with how machines achieve a particular goal or purpose. The accurate operation of a system often requires feed-forward mechanisms that adjust system performance in anticipation of future inputs. Pavlovian conditioning is ideally suited to subserve this function in behavioral systems. Pavlovian mechanisms have been demonstrated in various aspects of sexual behavior, maternal lactation, and infant suckling. Pavlovian conditioning of agonistic behavior has been also reported, and Pavlovian processes may likewise be involved in social play and social grooming. Several further lines of evidence indicate that Pavlovian conditioning can increase the efficiency and effectiveness of social interactions, thereby improving their cost/benefit ratio. We extend Pavlovian concepts beyond the traditional domain of discrete secretory and other physiological reflexes to complex real-world behavioral interactions and apply abstract laboratory analyses of the mechanisms of associative learning to the daily challenges animals face as they interact with one another in their natural environments.

Aggression↗

Normal mode analysis of macromolecular motions in a database framework: developing mode concentration as a useful classifying statistic.

We investigated protein motions using normal modes within a database framework, determining on a large sample the degree to which normal modes anticipate the direction of the observed motion and were useful for motions classification. As a starting point for our analysis, we identified a large number of examples of protein flexibility from a comprehensive set of structural alignments of the proteins in the PDB. Each example consisted of a pair of proteins that were considerably different in structure given their sequence similarity. On each pair, we performed geometric comparisons and adiabatic-mapping interpolations in a high-throughput pipeline, arriving at a final list of 3,814 putative motions and standardized statistics for each. We then computed the normal modes of each motion in this list, determining the linear combination of modes that best approximated the direction of the observed motion. We integrated our new motions and normal mode calculations in the Macromolecular Motions Database, through a new ranking interface at http://molmovdb.org. Based on the normal mode calculations and the interpolations, we identified a new statistic, mode concentration, related to the mathematical concept of information content, which describes the degree to which the direction of the observed motion can be summarized by a few modes. Using this statistic, we were able to determine the fraction of the 3,814 motions where one could anticipate the direction of the actual motion from only a few modes. We also investigated mode concentration in comparison to related statistics on combinations of normal modes and correlated it with quantities characterizing protein flexibility (e.g., maximum backbone displacement or number of mobile atoms). Finally, we evaluated the ability of mode concentration to automatically classify motions into a variety of simple categories (e.g., whether or not they are "fragment-like"), in comparison to motion statistics. This involved the application of decision trees and feature selection (particular machine-learning techniques) to training and testing sets derived from merging the "list" of motions with manually classified ones.

Databases, Protein↗

The AcCell series 2000 as a support system for training and evaluation in educational and clinical settings.

Providing effective training, retraining and evaluation programs, including proficiency testing programs, for cytoprofessionals is a challenge shared by many academic and clinical educators internationally. In cytopathology the quality of training has immediately transferable and critically important impacts on satisfactory performance in the clinical setting. Well-designed interactive computer-assisted instruction and testing programs have been shown to enhance initial learning and to reinforce factual and conceptual knowledge. Computer systems designed not only to promote diagnostic accuracy but to integrate and streamline work flow in clinical service settings are candidates for educational adaptation. The AcCell 2000 system, designed as a diagnostic screening support system, offers technology that is adaptable to educational needs during basic and in-service training as well as testing of screening proficiency in both locator and identification skills. We describe the considerations, approaches and applications of the AcCell 2000 system in education programs for both training and evaluation of gynecologic diagnostic screening proficiency.

Automation↗

Integrative dual-track transcriptomics reveals stage-specific coordination, regulatory divergence, and HSP90AA1-associated remodeling in human folliculogenesis.

Human folliculogenesis depends on coordinated yet non-identical developmental remodeling in the oocyte and its surrounding granulosa cells. When these two compartments remain synchronized and when they diverge into lineage-specific regulatory states, however, remains incompletely resolved. Here we performed an integrative dual-track re-analysis of the human RNA-seq dataset GSE107746, modeling oocytes and granulosa cells as distinct but developmentally linked compartments across follicular progression. Analysis of 148 sequencing libraries showed that compartment identity was the dominant source of transcriptomic variation, supporting compartment-aware downstream interpretation. Within this framework, oocytes followed a relatively continuous developmental trajectory, with substantial transcriptional remodeling already evident across adjacent stages, whereas granulosa cells showed weaker early-stage contrasts but markedly stronger late-stage reorganization, particularly around the antral and preovulatory transitions. Functional enrichment indicated that oocyte maturation was associated with RNA-processing and broader genome-regulatory remodeling, whereas granulosa maturation was dominated by progressive mitochondrial and bioenergetic activation. Co-expression analysis showed that both compartments contained strong late-stage programmes together with inverse early-state modules, indicating a shared systems-level architecture of maturation, although the hub-gene composition and biological content of these programmes were largely compartment-specific. Machine-learning validation reinforced this asymmetry: oocyte stage classification was best recovered from a compact eigengene-based representation, whereas granulosa stage discrimination was better resolved by a broader differential-expression-derived feature set. At the gene level, HSP90AA1 emerged as a stage-associated marker with compartment-specific behavior, showing progressive attenuation across oocyte development, assignment to the selected oocyte blue module, and sharper transitional dynamics in granulosa cells. Together, these findings support a model in which human folliculogenesis proceeds through coordinated but non-equivalent transcriptomic remodeling, with shared developmental logic at the systems level but distinct molecular execution in germline and somatic compartments.

Co-expression networks↗

Machine learning based pattern recognition applied to microarray data.

MOTIVATION: Microarrays have allowed the expression level of thousands of genes or proteins to be measured simultaneously. Data sets generated by these arrays consist of a small number of observations (e.g., 20-100 samples) on a very large number of variables (e.g., 10,000 genes or proteins). The observations in these data sets often have other attributes associated with them such as a class label denoting the pathology of the subject. Finding the genes or proteins that are correlated to these attributes is often a difficult task since most of the variables do not contain information about the pathology and as such can mask the identity of the relevant features. We describe a genetic algorithm (GA) that employs both supervised and unsupervised learning to mine gene expression and proteomic data. The pattern recognition GA selects features that increase clustering, while simultaneously searching for features that optimize the separation of the classes in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the features chosen by the GA contain information primarily about differences between classes in the data set. The principal component analysis routine embedded in the fitness function of the GA acts as an information filter, significantly reducing the size of the search space since it restricts the search to feature sets whose principal component plots show clustering on the basis of class. The algorithm integrates aspects of artificial intelligence and evolutionary computations to yield a smart one pass procedure for feature selection, clustering, classification, and prediction.

Algorithms↗

Integrated unit performance testing of powered, air-purifying particulate respirators using a DOP challenge aerosol.

Although workplace protection factor (WPF) and simulated workplace protection factor (SWPF) studies provide useful information regarding the performance capabilities of powered air-purifying respirators (PAPRs) under certain workplace or simulated workplace conditions, some fail to address the issue of total PAPR unit performance over extended time. PAPR unit performance over time is of paramount importance in protecting worker health over the course of a work shift or at least for the recommended service lifetime of the PAPR battery pack, whichever is shorter. The need for PAPR unit performance testing has become even more important with the inception of 42 CFR 84 and the recent introduction of electrostatic respirator filter media into the PAPR market. This study was conducted to learn how current PAPRs certified by the National Institute for Occupational Safety and Health would perform under an 8-hour unit performance test similar to the dioctyl phthalate (DOP) loading test described in 42 CFR 84 for R- and P-series filters for nonpowered, air-purifying particulate respirators. In this study, entire PAPR units, four with mechanical filters and one with an electrostatic filter, were tested using a TSI Model 8122 Automated Respirator Tester, with and without the built-in breathing machine. The two, tight-fitting PAPRs, both with mechanical filters, showed little effect on performance resulting from the breathing machine. The two loose-fitting helmet PAPRs indicate that unit performance testing without the breathing machine is a more stringent test than testing with the breathing machine under the conditions used. The PAPR with a loose-fitting hood gave inconclusive results as to which testing condition is more stringent. The PAPR unit equipped with electrostatic filters gave the highest maximum penetration values during unit performance testing.

Aerosols↗

A Meta-learning-driven strategy for adulteration detection in sweet potato starch and vermicelli using Raman spectroscopy.

To address the widespread adulteration of sweet potato starch and its vermicelli with cheaper starches and overcome conventional supervised learning's dependency on large labeled datasets, this study developed a few-shot discrimination method integrating Raman spectroscopy with meta-learning. We constructed a meta-learning framework using cassava- and wheat-adulterated sweet potato starch as the source domain for training, with potato-adulterated sweet potato starch and cassava-adulterated sweet potato vermicelli as two target domains for testing. Raman spectra showed high consistency between sweet potato vermicelli and its raw starch, laying the foundation for cross-domain detection. Testing yielded comprehensive classification accuracies of 95.33% and 98.00% for the two target domains, significantly outperforming SVM, RF, and CNN (max. 85.24%). This approach effectively identifies subtle starch variety differences in complex adulteration, providing novel food quality inspection solutions and verifying the feasibility of raw material-to-finished product cross-domain detection.

Ipomoea batatas↗

Does computation provide a model for creativity? An epistemological perspective in neuroscience.

In 1939 Alan Turing, a major scholar in the field of mechanical computation, described a system whose computational power was beyond that of a discrete, finite state machine (Turing Machine). The composition of this system was likely the first example of what is now called an hybrid computational system. Since then, development of neural networks and brain automata has made aware that forms of computation might exist that are likely to go beyond Turing's limits. Natural systems, like the central nervous system in Mammals and man, are likely to use such a type of computation, especially to perform highly integrating activities, like feedback controls and mental creative processes. The latter are usually understood as processes that involve infinitary procedures, ending up in a complex information network, the computational maps, in which both digital, Turing-like computation and continuous, analog forms of calculus are expected to occur. Pictorial representation may be a fruitful example, mostly metaphorical, to analyze the use of this hybrid forms of computation by higher order computational maps, and the possible role of these types of computational processes in painting creativity is briefly analyzed in comparing 15th vs 16th century Renaissance Art. An open challenge for neuroscience in the 21st century is to clarify whether a hybrid neural learning network might represent a reasonable clue to scientifically interpret the theme of "creativity".

Art↗

PMGen: from peptide-MHC structure prediction to peptide generation.

MOTIVATION: Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS: We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core C&#x3b1; RMSDs of 0.62&#xa0;&#xc5; for MHC-I and 0.33&#xa0;&#xc5; for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63&#xa0;817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION: PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

Peptides↗

Human-AI Interaction With AI-Assisted Tumor Overlays in Pediatric Whole-Body Magnetic Resonance Imaging: Exploratory Reader Study.

BACKGROUND: AI tools have the potential to enhance personalized clinical care, particularly in radiology. However, their integration into clinical workflows remains complex, especially in pediatric oncology, where early cancer detection is critical. Children with Li-Fraumeni syndrome (LFS), a rare cancer predisposition disorder, undergo regular surveillance whole-body magnetic resonance imaging (wbMRI), which presents an opportunity for AI-assisted tumor detection. OBJECTIVE: We evaluated the feasibility of an AI-assisted overlay for highlighting tumor-like regions in pediatric surveillance wbMRI and explored how access to the overlay influenced radiologist workflow, candidate-lesion marking behavior, follow-up recommendations, and perceived workload. METHODS: We developed a patch-based AI segmentation model trained on augmented 2D slices from 675 surveillance wbMRI volumes of pediatric patients with LFS. The model was designed to highlight regions with high tumor probability. A reader study was conducted with 2 radiologists who independently reviewed wbMRI cases both with and without AI assistance. We measured evaluation time, number and location of reader-marked candidate lesions, type of follow-up recommendation, and subjective feedback using structured questionnaires. RESULTS: AI assistance altered interpretation workflows for both radiologists, with mixed effects. On average, the time required to evaluate each case increased when using the AI tool for both radiologists. However, one radiologist had an increase in the number of candidate lesion locations selected with the tool, and one had a decrease in the number of candidate lesion locations selected with the tool. Subjective feedback indicated that one of the radiologists reported lower mental demand with the AI tool, while both radiologists reported lower stress with the AI tool. Interrater variability was evident, underscoring the need for personalized calibration of AI tools. CONCLUSIONS: AI-assisted wbMRI interpretation can improve tumor detection in pediatric cancer surveillance by reducing false negatives. However, its influence on workflow efficiency and interradiologist variability highlights the importance of careful implementation. Successful integration requires addressing challenges such as improving the predictive precision of AI models, offering intuitive end-user designs and instructions, and building trust in AI outputs. AI outputs can influence workflow and behavior in reader-specific ways. Clinical translation will require larger, randomized, multireader studies and model refinement to reduce false positives and quantify lesion-level reader performance. This can help ensure better patient outcomes in addition to reduced clinician burnout.

Humans↗

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n&#xa0;=&#xa0;549) and a validation set (n&#xa0;=&#xa0;236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60&#xa0;mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60&#xa0;mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans↗

Demographics, Overlap, and Latency of Severe Cutaneous Adverse Reactions in an FDA Database.

IMPORTANCE: Severe cutaneous adverse reactions (SCARs), including Stevens-Johnson syndrome/toxic epidermal necrolysis (SJS-TEN), drug reaction with eosinophilia and systemic symptoms (DRESS), acute generalized exanthematous pustulosis (AGEP), and generalized bullous fixed drug eruption (GBFDE), are rare but life-threatening drug hypersensitivity syndromes. Due to their low incidence and diagnostic complexity, large-scale characterization of SCAR is challenging. OBJECTIVE: To characterize the demographics, causative agents, trends, latency, and phenotypic overlap of SCAR using a large-scale, sanitized pharmacovigilance dataset from FAERS (FDA Adverse Event Reporting System). DESIGN: Cross-sectional study of spontaneous adverse event reports. Cases were drawn from the U.S. Food and Drug Administration Adverse Event Reporting System (FDA FAERS) from January 2004 to December 2023 and subjected to sanitization and deduplication. Disproportionality analysis was used to characterize causative agents. Machine learning (random forest classifiers) was used to analyze predictors of drug latency and mortality. SETTING: Global pharmacovigilance reports submitted to FAERS. PARTICIPANTS: A total of 56,683 deduplicated SCAR reports were identified, representing 0.33% of reports during the study period. EXPOSURES: Suspected causative drugs, including both small molecules and biologics. MAIN OUTCOMES AND MEASURES: Main outcomes included the frequency and distribution of SCAR syndromes, reporting trends over time, latency from drug start to reaction onset, drug-specific disproportionality (PRR, ROR, IC), and co-reporting between SCAR types and related conditions. RESULTS: A total of 56,683 unique SCAR reports were identified, including SJS-TEN (28,871), DRESS (22,444), AGEP (6,183), and GBFDE (150). We identified 237 drugs with significant disproportionality for SCAR overall. Co-reporting between SCARs was significantly enriched (p < 1e-200), suggesting overlapping phenotypes. Latency varied by drug and syndrome (median: GBFDE 3 days, AGEP 4 days, SJS-TEN 12 days, DRESS 20 days). CONCLUSIONS AND RELEVANCE: SCAR syndromes display distinct but overlapping phenotypes, with variable latency and diverse causative agents. These findings, based on the largest SCAR dataset to date, highlight the need for improved classification frameworks and molecular validation. Large-scale pharmacovigilance, integrated with genomic and histopathologic data, will be critical to improving diagnosis, mechanistic understanding, and clinical management of SCAR.

Acute Generalized Exanthematous Pustulosis↗

Integrated analysis of established and novel microbial and chemical methods for microbial source tracking.

Several microbes and chemicals have been considered as potential tracers to identify fecal sources in the environment. However, to date, no one approach has been shown to accurately identify the origins of fecal pollution in aquatic environments. In this multilaboratory study, different microbial and chemical indicators were analyzed in order to distinguish human fecal sources from nonhuman fecal sources using wastewaters and slurries from diverse geographical areas within Europe. Twenty-six parameters, which were later combined to form derived variables for statistical analyses, were obtained by performing methods that were achievable in all the participant laboratories: enumeration of fecal coliform bacteria, enterococci, clostridia, somatic coliphages, F-specific RNA phages, bacteriophages infecting Bacteroides fragilis RYC2056 and Bacteroides thetaiotaomicron GA17, and total and sorbitol-fermenting bifidobacteria; genotyping of F-specific RNA phages; biochemical phenotyping of fecal coliform bacteria and enterococci using miniaturized tests; specific detection of Bifidobacterium adolescentis and Bifidobacterium dentium; and measurement of four fecal sterols. A number of potentially useful source indicators were detected (bacteriophages infecting B. thetaiotaomicron, certain genotypes of F-specific bacteriophages, sorbitol-fermenting bifidobacteria, 24-ethylcoprostanol, and epycoprostanol), although no one source identifier alone provided 100% correct classification of the fecal source. Subsequently, 38 variables (both single and derived) were defined from the measured microbial and chemical parameters in order to find the best subset of variables to develop predictive models using the lowest possible number of measured parameters. To this end, several statistical or machine learning methods were evaluated and provided two successful predictive models based on just two variables, giving 100% correct classification: the ratio of the densities of somatic coliphages and phages infecting Bacteroides thetaiotaomicron to the density of somatic coliphages and the ratio of the densities of fecal coliform bacteria and phages infecting Bacteroides thetaiotaomicron to the density of fecal coliform bacteria. Other models with high rates of correct classification were developed, but in these cases, higher numbers of variables were required.

Animals↗

SPINE: an integrated tracking database and data mining approach for identifying feasible targets in high-throughput structural proteomics.

High-throughput structural proteomics is expected to generate considerable amounts of data on the progress of structure determination for many proteins. For each protein this includes information about cloning, expression, purification, biophysical characterization and structure determination via NMR spectroscopy or X-ray crystallography. It will be essential to develop specifications and ontologies for standardizing this information to make it amenable to retrospective analysis. To this end we created the SPINE database and analysis system for the Northeast Structural Genomics Consortium. SPINE, which is available at bioinfo.mbb.yale.edu/nesg or nesg.org, is specifically designed to enable distributed scientific collaboration via the Internet. It was designed not just as an information repository but as an active vehicle to standardize proteomics data in a form that would enable systematic data mining. The system features an intuitive user interface for interactive retrieval and modification of expression construct data, query forms designed to track global project progress and external links to many other resources. Currently the database contains experimental data on 985 constructs, of which 740 are drawn from Methanobacterium thermoautotrophicum, 123 from Saccharomyces cerevisiae, 93 from Caenorhabditis elegans and the remainder from other organisms. We developed a comprehensive set of data mining features for each protein, including several related to experimental progress (e.g. expression level, solubility and crystallization) and 42 based on the underlying protein sequence (e.g. amino acid composition, secondary structure and occurrence of low complexity regions). We demonstrate in detail the application of a particular machine learning approach, decision trees, to the tasks of predicting a protein's solubility and propensity to crystallize based on sequence features. We are able to extract a number of key rules from our trees, in particular that soluble proteins tend to have significantly more acidic residues and fewer hydrophobic stretches than insoluble ones. One of the characteristics of proteomics data sets, currently and in the foreseeable future, is their intermediate size ( approximately 500-5000 data points). This creates a number of issues in relation to error estimation. Initially we estimate the overall error in our trees based on standard cross-validation. However, this leaves out a significant fraction of the data in model construction and does not give error estimates on individual rules. Therefore, we present alternative methods to estimate the error in particular rules.

Animals↗

Attribution of PM2.5-Induced Transcriptomic Perturbation to Toxic Components.

Ambient fine particulate matter (PM2.5) is a chemically complex mixture whose health impacts are not fully captured by particle mass. Here, we developed an interpretable chemotranscriptomic framework to attribute PM2.5-induced molecular perturbations to toxicity-relevant components. PM2.5 collected from urban roadside and coastal environments was separated into whole, extractable, and unextractable fractions, characterized by LC/GC &#xd7; GC-HRMS-based nontarget analysis and inductively coupled plasma mass spectrometry (ICP-MS), and evaluated using cytotoxicity testing and transcriptomic profiling in human bronchial epithelial cells. Urban PM2.5 exhibited greater cytotoxic potency per unit mass than coastal PM2.5, with extractable fractions accounting for most cytotoxic and pathway-level responses. Transcriptomics revealed distinct site-specific modes of action: urban PM2.5 preferentially induced oxidative stress, xenobiotic metabolism, and cell cycle suppression, consistent with acute, nonapoptotic injury, whereas coastal PM2.5 elicited weaker cytotoxicity but stronger interferon-mediated immune and apoptosis-related signaling. Integrating chemical abundance with pathway activity using random forest regression, SHAP interpretation, and mechanistic corroboration reduced 5,033 detected features to 444 pathway-linked candidate drivers. Fewer than 5% of features explained &#x223c;95% of cumulative model contribution. Standard-confirmed contributors included plasticizer-related compounds, aromatic and heteroaromatic combustion products, and copper for urban PM2.5 and secondary/aged organics and nickel for coastal PM2.5. These findings support mechanism-informed prioritization of hazardous PM2.5 components beyond mass-based assessment.

Particulate Matter↗