PubMed HealthSearch

SEARCH · PubMed Health

Results for “Process optimization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Engineering the Vero Cell Lineage: Toward a Programmable Vaccine Manufacturing Platform.

Vero cells remain an indispensable continuous substrate for human viral vaccine manufacturing. Despite decades of empirical process optimization, intrinsic genomic instability, including segmental aneuploidy and dynamic chromatin rearrangements, continues to limit the durability of engineered phenotypes under sustained viral burden and bioreactor stress. Here, we review the expanding engineering toolkit for the Vero lineage across a three-layered functional framework: the membrane interface, cytoplasmic foundry, and nuclear blueprint, evaluating translational prospects at each level. Receptor transplantation and morphological reprogramming have broadened viral entry range and enabled suspension-adapted culture formats, while metabolic flux management and temporally controlled apoptosis modulation have addressed intracellular production bottlenecks, albeit often with trade-offs between productivity, biosafety, and long-term population stability. At the genomic level, targeted perturbations of transcriptional regulators and emerging epigenetic interventions offer more durable gains, yet expression drift, clonal heterogeneity, and karyotypic instability during extended passaging highlight the need for locus-level precision rather than constitutive trait installation. Looking forward, infection-responsive dynamic logic circuits and the systematic identification of Vero-specific genomic safe harbors could shift the paradigm toward a conditionally responsive manufacturing architecture. Collectively, these advances suggest a pathway for transitioning the Vero lineage from a passive, empirically optimized biological substrate into a conditionally responsive, genomically stable, and programmable platform for modern vaccine preparedness.

Vero cells

Comparative profiling of microbial community structure, enzyme potential, metabolic features, and volatile composition in craft and Jiafan Huangjiu processes.

Craft Huangjiu and Jiafan Huangjiu represent two distinct industrial Huangjiu product outcomes with contrasting volatile profiles. This study compared craft Huangjiu (L70) and Jiafan Huangjiu (L79) to characterize their physicochemical, microbial, gene-level functional, metabolic, and volatile features. Because L70 involved mid-fermentation addition of finished Huangjiu, this comparison was not intended to isolate the sole effect of fermentation interruption versus continued fermentation. L79 showed more extensive carbon and nitrogen utilization, with lower residual substrates and higher ethanol and acetic acid contents than L70, whereas L70 retained a less complete fermentation state. At the volatile level, GC-MS and volatile metabolomics consistently showed an ester-enriched profile in L79 and a more alcohol-dominant profile in L70. FlavorDB-based putative annotation and threshold-based OAV analysis further indicated distinct database-assigned descriptor distributions and potential odor-active compounds, with more OAV > 1 ester-related compounds in L79. Metagenomic analysis showed that L70 was dominated by Lactobacillus acetotolerans, whereas L79 contained higher relative abundances of Saccharomyces cerevisiae, Aspergillus oryzae, Aspergillus flavus, and Fructilactobacillus fructivorans. Metagenomic functional annotation showed higher representation of hydrolysis-related CAZy genes and ester-related enzyme annotations in L79. KEGG-based pathway mapping further indicated greater gene-level potential for ethanol-, acetate-, and acetyl-CoA-related metabolism in L79. Accordingly, the L70 profile should be interpreted as the integrated final-product outcome of process intervention, exogenous input, and subsequent fermentation. The findings provide a comparative basis for future flavor regulation and process optimization in Huangjiu and other fermented alcoholic beverages.

Volatile Organic Compounds

Proteomic characterization of acidic aqueous extracts from Vicia faba L. pod valves identifies chitinase as a major co-extracted protein macromolecule.

Naturally acidic aqueous extracts from Vicia faba L. pod valves are being explored as sustainable, L-DOPA-oriented plant preparations. Pod valves represent an underutilized processing by-product reported to contain L-DOPA, a compound widely used in Parkinson's disease therapy, while acidic aqueous media may help preserve its physicochemical stability. However, the protein macromolecules co-extracted from V. faba pod valves under these conditions remain poorly characterized. This information is relevant because persistent plant proteins may influence extract composition, stability, susceptibility to degradation, and downstream processing requirements. Here, we characterized co-extracted V. faba protein macromolecules in aqueous pod-valve extracts prepared in ultrapure water or naturally acidic media, including 2% Phyllanthus emblica, 5% Punica granatum, and 2% Ribes rubrum. Protein profiles were first evaluated by SDS-PAGE and subsequently analyzed by nanoflow liquid chromatography coupled to high-resolution tandem mass spectrometry (nLC-MS/MS). Protein identifications were complemented with Gene Ontology annotation and a descriptive semi-quantitative assessment of relative protein representation across extraction media. Chitinase was the most represented V. faba-assigned protein macromolecule across the extracts, with additional highly represented proteins including glucan endo-1,3-beta-D-glucosidase, pathogenesis-related proteins, and polyphenol oxidase A1. These co-extracted proteins are mainly associated with plant defense, stress responses, cell-wall remodeling, and oxidative processing, suggesting that they may be relevant for extract quality attributes during handling and storage. This study provides a compositional proteomic reference for the co-extracted protein macromolecules present in acidic aqueous extracts from V. faba pod valves, supporting future studies on extract stability, processing optimization, and the development of standardized plant-based preparations.

Vicia faba

Selective monitoring of trace-level catechin and myricetin in herbal and aqueous matrices using magnetic MIP-DSPME: Optimization via design of experiments.

A novel dispersive solid-phase microextraction approach utilizing a magnetic molecularly imprinted polymer (MMIP) integrated with HPLC-UV detection was developed for the concurrent quantification of catechin and myricetin in herbal extracts and aqueous samples. The sorbent was engineered as a core-shell nanocomposite, consisting of a selective polymer layer deposited onto Fe3O4@SiO2-APTMS magnetic nanoparticles. Dual-template imprinting using catechin and myricetin generated complementary binding cavities within the polymer framework. Experimental variables influencing extraction were systematically screened and subsequently optimized. A Plackett-Burman design was first applied to identify the most influential factors, with pH and sorption time identified as the dominant variables. These parameters were subsequently fine-tuned using a central composite design, and the optimization process was completed in only 30 experimental runs. The sorption characteristics of the imprinted sorbent (MMIP) were compared with those of its non-imprinted counterpart (MNIP). The MMIP demonstrated markedly higher maximum binding capacities (Qmax), reaching 119.3 mg g-1 for myricetin and 112.1 mg g-1 for catechin, whereas the corresponding values for the MNIP were 32.55 and 32.08 mg g-1, respectively. Moreover, the affinity constants (KL = 0.760-0.950 L mg-1) were approximately 2.3-fold higher for the MMIP, confirming its stronger and more selective interactions with the target analytes. The selectivity coefficients for the targeted flavonoids relative to structurally related compounds, including ferulic acid, p-coumaric acid, melatonin, and curcumin, exceeded 3.5 for the MMIP, whereas the corresponding values for the MNIP were close to 1.1, demonstrating the high molecular recognition capability of the imprinted sorbent. Method validation demonstrated limits of detection (LODs) of 0.33-0.59 ng mL-1 and limits of quantification (LOQs) of 1.10-1.96 ng mL-1, and excellent linearity over the concentration range of 5.0-5500 ng mL-1 (R2 > 0.998). The method achieved recoveries of 93.96% to 105.69% with RSDs below 5.5%, while the preconcentration factors ranged from 209 to 229. Furthermore, the sorbent retained more than 95% of its extraction efficiency after four consecutive reuse cycles and more than 80% after six cycles, demonstrating excellent stability and reusability. The proposed method was successfully applied to the analysis of six medicinal plant extracts and water samples, showing negligible matrix interference and superior sensitivity, selectivity, and operational simplicity compared with conventional solid-phase extraction methods.

Flavonoids

Unicorn: enhancing single-cell Hi-C data with blind super-resolution for 3D genome structure reconstruction.

MOTIVATION: Single-cell Hi-C (scHi-C) data provide critical insights into chromatin interactions at individual cell levels, uncovering unique genomic 3D structures. However, scHi-C datasets are characterized by sparsity and noise, complicating efforts to accurately reconstruct high-resolution chromosomal structures. In this study, we present ScUnicorn, a novel blind super-resolution framework for scHi-C data enhancement. ScUnicorn uses an iterative degradation kernel optimization process, unlike traditional super-resolution approaches, which rely on downsampling, predefined degradation ratios, or constant assumptions about the input data to reconstruct high-resolution interaction matrices. Hence, our approach more reliably preserves critical biological patterns and minimizes noise. Additionally, we propose 3DUnicorn, a maximum likelihood algorithm that leverages the enhanced scHi-C data to infer precise 3D chromosomal structures. RESULTS: Our evaluation demonstrates that ScUnicorn achieves superior performance over the state-of-the-art methods in terms of Peak Signal-to-Noise Ratio, Structural Similarity Index Measure, and GenomeDisco scores. Moreover, 3DUnicorn's reconstructed structures align closely with experimental 3D-FISH data, underscoring its biological relevance. Together, ScUnicorn and 3DUnicorn provide a robust framework for advancing genomic research by enhancing scHi-C data fidelity and enabling accurate 3D genome structure reconstruction. AVAILABILITY AND IMPLEMENTATION: Unicorn implementation is publicly accessible at https://github.com/OluwadareLab/Unicorn.

Single-Cell Analysis

Induced pluripotent stem cell reprogramming: methodological evolution and challenges in clinical translation.

Cell reprogramming can transform somatic cells into induced pluripotent stem cells providing a platform for patient-specific disease modeling, drug screening and regenerative medicine research. Since the advent of OKSM-mediated reprogramming, the system of technical approaches has evolved continuously - from integrated viral vectors to non-integrated episomal systems and, more recently, chemical reprogramming and CRISPR approaches. The simultaneous advances in single-cell multi-omics, biomaterials engineering, and artificial intelligence have further refined the controllability and precision of the reprogramming process. Despite these innovations, problems persist that hinder clinical translation: incomplete epigenetic resetting, ongoing clonal heterogeneity, genomic instability in long-term culture, and the lack of standardized Good Manufacturing Practice protocols for large-scale manufacturing. This review summarizes the trajectory of iPSC reprogramming technologies, with special emphasis on the translational applicability of each modality. We evaluated viral and nonviral delivery systems, chemical reprogramming, strategies that aid gene editing, and emerging engineering platforms, including microfluidics, smart biomaterials, and artificial-intelligence-driven process optimization. We further identify the core "translational triltrilas", namely, the inherent tradeoffs between security, homogeneity, and scalability, and propose a comprehensive strategy to overcome these bottlenecks. By linking basic mechanistic understandings with industrial and regulatory considerations, this review aims to provide a route for transitioning iPSC technology from a laboratory tool to a clinically viable manufacturing platform.

clinical translation

Marine endophytes: biosynthetic engines for novel bioactive metabolites.

Marine endophytes are prolific sources of structurally diverse secondary metabolites with significant pharmaceutical potential, including anticancer, antimicrobial, and antioxidant agents. However, their commercial utilization is hindered by genomic instability in axenic cultures and inconsistent metabolite yields. While current studies focus on symbiotic interactions and compound discover, critical gaps persist in harnessing their biosynthetic capabilities. This review synthesizes knowledge on marine fungal metabolites and proposes a paradigm shift toward resource-driven research. It addresses strain improvement limitations and suggests strategies like mutagenesis, protoplast fusion, and metabolic engineering to bolster production stability and efficiency. The paper also discusses biological process optimization, including fermentation tuning, inducer and precursor addition, and adsorbent use, to enhance natural product synthesis. By identifying these research gaps and proposing a strategic roadmap, the review advances the stable and scalable production of bioactive metabolites, unlocking the commercial and therapeutic potential of marine endophytic fungi.

bioactive metabolites

[Ways to optimize the technological process of oleandomycin reextraction].

Dependence of the oleandomycin distribution coefficient on pH of the acqueous phase and temperature in the system of butylacetate extract-water acidified with orthophosphoric acid was studied. With a purpose of intensification of the process of oleandomycin reextraction, decreasing the antibiotic inactivation and evaporation of the organic solvent it was proposed to perfom oleandomycin extraction at pH 4.0--5.0 accompanied by simultaneous decreasing of the temperature.

Acetates

Validation of an integrated metagenomic pipeline combining optimized wet-lab processing and tiered reporting for CSF pathogen detection.

UNLABELLED: Metagenomic next-generation sequencing (mNGS) in the infectious disease diagnostic space has been gaining traction and is popular for aiding in the diagnosis of central nervous system infections. However, many challenges and obstacles remain in making this technology a gold standard for infectious disease diagnostic testing. One major challenge is being able to distinguish between the clinically relevant organisms from background contamination. We performed a validation study for mNGS on cerebrospinal fluid (CSF) that utilized positive clinical samples and contrived samples that incorporated a bioinformatics pipeline that can better distinguish between background contamination and clinically relevant organisms and used a three-tiered reporting algorithm meant to decrease the inherent subjectivity that comes with interpreting and reporting data from clinical metagenomic sequencing. The validation of this assay and category-based reporting pipeline revealed an overall concordance of 91.8%, with a sensitivity of 100% and a specificity of 72.4%. In addition, we improved the detection of clinically relevant RNA viruses to almost 100% in the CSF by modifying the wet lab processing of the sample. This bioinformatics pipeline with a category-based reporting algorithm will provide more confidence in reporting microorganisms detected with this technology, mNGS, and improving patient care. IMPORTANCE: Metagenomic next-generation sequencing (mNGS) can offer a broad, unbiased approach for the detection of infectious pathogens and has shown promise in diagnosing central nervous system infections. Despite its potential, clinical implementation remains limited by challenges in distinguishing clinically relevant organisms from background contamination. This study validated an mNGS assay for cerebrospinal fluid that incorporates an optimized bioinformatics pipeline with a three-tiered reporting algorithm designed to reduce subjectivity and enhance diagnostic confidence. The assay also has improved detection of clinically relevant RNA viruses through modified wet-lab processing. These findings support the clinical utility of a structured, category-based reporting approach for mNGS, advancing its reliability as a diagnostic tool in infectious disease testing.

Metagenomics

Integrated widely targeted metabolomics and GC-IMS reveal dynamic flavor, nutritional, functional, and metabolic profiles in macadamia kernels during processing.

Different processing stages influence the color, flavor, and antioxidant activities of macadamia kernels. However, the biochemical mechanisms that occur during processing are not well known. This study integrated widely targeted metabolomics (UPLC-MS/MS) with GC-IMS to systematically characterize non-volatile and volatile compounds in macadamia kernels across key three sample groups: fresh kernels (FMN), low-temperature-dried kernels (DMN), and roasted kernels (BMN). A total of 622 non-volatile metabolites and 52 volatile compounds were identified. Low-temperature drying promoted the accumulation of phenolic acids and flavonoids, enhancing antioxidant capacity. Roasting degraded heat-sensitive nutrients but generated flavor compounds via Maillard reaction and lipid oxidation, shifting aroma from green to nutty notes. Nutritional assessment confirmed that roasting significantly reduced antioxidant activities and bile acid binding capacity. Pearson correlation analysis verified the key metabolite-antioxidant relationships. These findings provide critical insights into metabolic dynamics during nut processing and establish a scientific basis for optimizing thermal processing strategies.

Metabolomics

A correction factor for bridging compaction simulator and different roller compactors.

Roller compaction (RC) is an important dry granulation technique. Since pilot and commercial scale roller compactors, which operate continuously on a large scale, usually require kilograms of material per run, formulation and process development directly on such roller compactors is not practical. In contrast, a compaction simulator (CS) can produce ribblets, also known as "slugs", using only a few grams of material with sinusoidal displacement profile replicating the motion of a specific point on the roll surface. Thus, it is possible to develop RC formulation and process in laboratory using a CS-based material-sparing approach. However, because of the inherently different configurations for applying pressure between die compression and roll compression, translating uniaxial pressure from CS experiments to roll pressure during RC is often unreliable, leading to significant uncertainties in the critical quality attributes of ribbons, such as ribbon solid fraction (or porosity) and mechanical strength. The objective of this study was to identify a correction factor (Kp = uniaxial die compression pressure/roll pressure), by correlating the compressibility profiles from CS and a roller compactor of interest, to enable more reliable process translation from CS to roller compactor. In this study, a Kp value of 0.5 was determined for Alexanderwerk WP120 and validated for Gerteis Mini-Pactor and Bepex Pharmapactor. This value may serve as a starting point for translating the optimal compaction pressure identified based on CS investigation to common roller compactors, requiring only minor adjustments to attain optimal RC process parameters (i.e., roll force and roll gap) for a chosen roller compactor.

Drug Compounding

Improving recombinant protein productivity in CHO cells via multi-omics data integration.

Chinese hamster ovary (CHO) cells represent the dominant host system for the production of recombinant therapeutic proteins. In recent decades, extensive research has focused on process/media optimization and cell line engineering to improve both the productivity and quality of biopharmaceutical proteins produced in CHO cells. Nevertheless, the inherent complexity of biological pathways and the heterogeneous cellular responses to different environmental conditions have posed substantial challenges to traditional methodologies. Recent advances in omics technologies have enabled comprehensive characterization of CHO cell physiology, providing multidimensional molecular and phenotypic insights that facilitate the enhancement of recombinant protein production. This review first summarizes the methodologies and advances in CHO omics research, including genomics, transcriptomics, proteomics, metabolomics, and epigenomics. It then examines contemporary approaches to integrate and analyze multi-omics data in CHO cells. The review further elucidates how these multi-omics datasets can be strategically applied across various developmental stages, including cell line selection, genetic engineering, expression vector design, and bioprocess optimization. Finally, we explore the transformative potential of integrating multi-omics with artificial intelligence and discuss promising future research directions in CHO cell studies. These emerging paradigms offer novel opportunities for data-driven cell engineering and bioprocess optimization in CHO-based biomanufacturing.

Bioprocessing

Identification of Plant Chromatin Interaction Networks Using IP-MS and co-IP.

Proteins often act in concert to perform their function. Thus, the identification of protein complexes is crucial if we want to understand how they work. In this chapter, we present a highly sensitive protocol for the immunoprecipitation of nuclear chromatin-linked proteins in Arabidopsis thaliana that does not rely on time-consuming nuclei extraction. Interaction partners are identified using mass spectrometry and confirmed by co-immunoprecipitation. To help solubilize chromatin-bound proteins and eliminate nonspecific interactions of proteins binding the same DNA stretch, we include an enzymatic digestion step to remove DNA before immunoprecipitation. Our protocol offers a simplified process using optimized buffers, which facilitates quick and effective immunoprecipitation. The outcome is high-quality eluates that are ideal for identifying proteins through MS.

Chromatin

Plant-derived and microbial biostimulants in sustainable agriculture: mechanisms, applications, and challenges.

Plant biostimulants have emerged as transformative and sustainable tools for improving crop productivity, resource-use efficiency, and resilience under rapidly intensifying environmental stresses. Unlike conventional agrochemicals, biostimulants function by activating physiological, biochemical, and molecular processes that optimize plant performance without directly supplying nutrients or exerting pesticidal effects. This review comprehensively examines the integrated roles of plant-derived and microbial biostimulants in sustainable agriculture, with particular emphasis on microbial-mediated mechanisms underlying plant stress adaptation and rhizosphere functioning. Plant-derived biostimulants, including seaweed extracts, humic substances, protein hydrolysates, amino acids, and chitosan, enhance nutrient acquisition, root architecture, hormonal regulation, and antioxidant defense systems. More importantly, microbial biostimulants, such as plant growth-promoting rhizobacteria (PGPR), endophytic microorganisms, mycorrhizal fungi, actinomycetes, yeasts, and cyanobacteria, exert multifunctional effects through biological nitrogen fixation, mineral solubilization, phytohormone biosynthesis, volatile signaling, osmolyte accumulation, pathogen suppression, and modulation of stress-responsive genes. These beneficial microorganisms reshape rhizosphere microbial communities, improve nutrient cycling, and enhance plant tolerance to drought, salinity, heat, and heavy metal toxicity. Emerging evidence from genomics, transcriptomics, metabolomics, and microbiome-based investigations has further revealed the molecular networks and signaling pathways governing biostimulant-induced resilience and plant-microbe interactions. Despite their substantial promise, inconsistent field performance, formulation instability, regulatory limitations, and inadequate mechanistic understanding continue to restrict their large-scale adoption. This review highlights recent advances in microbial and plant-derived biostimulants while identifying critical knowledge gaps and future opportunities for precision biostimulant engineering, microbiome manipulation, and climate-resilient crop management. The integration of next generation biostimulant technologies into sustainable agricultural systems may significantly reduce dependence on agrochemicals while improving crop productivity, environmental sustainability, and global food security.

Agriculture

Inducible flocculation in Komagataella phaffii enables enhanced biomass separation for biopharmaceutical production.

Biomass separation represents a critical bottleneck in Komagataella phaffii-based biopharmaceutical processes, as typically high cell densities of 40 - 50 % create significant operational, technical and economic challenges for harvest operations. Yeast cell aggregation (flocculation) provides a solution to accelerate cell sedimentation by increasing particle size, thus allowing to improve biomass-supernatant separation efficiency during both natural gravity settling and (continuous) centrifugation operations. This study demonstrates successful engineering of K. phaffii strains with an inducible flocculation phenotype using CRISPR/Cas9-based genome editing to integrate the Saccharomyces cerevisiae FLO1 (ScFLO1) gene under control of various regulatory elements, including methanol-inducible and derepressible promoters. Flocculation strength could be enhanced by implementing transcriptional positive feedback circuits based on the methanol-inducible AOX1 promoter. To address methanol-free production requirements, we developed alternative systems to retrofit PAOX1-based ScFLO1 expression and exploited the derepressible PDF promoter, offering broader compatibility with biopharmaceutical manufacturing facilities. Flocculating cells cultivated in a bioreactor demonstrated significantly improved sedimentation behavior, with considerably lower supernatant turbidity after short low-speed centrifugation or gravity sedimentation compared to non-flocculating controls. Crucially, cell flocculation had no negative impact on product amount and quality when expressing a multivalent NANOBODY® VHH molecule with pharmaceutical relevance. Thus, this work establishes the first genetically engineered flocculation system in K. phaffii compatible with recombinant protein production, providing the basis for an innovative approach to streamline harvest operations in biopharmaceutical processes.

Flocculation

Retinal hypoxia reversal with PLGA-oxygen nanobubbles.

Pathologies associated with retinal hypoxia, including diabetic retinopathy, central/branch retinal artery occlusion (CRAO/BRAO), central/branch retinal vein occlusion (CRVO/BRVO), retinopathy of prematurity, sickle cell retinopathy, etc., have limited effective therapeutic intervention strategies. To address this shortcoming, herein we propose a biocompatible and biodegradable poly (lactic-co-glycolic acid) shell-based oxygen nanobubbles (PLGA-ONBs) platform, formulated with PLGA, polyvinyl alcohol (PVA), and NaHCO3. The formulation of a novel PLGA-ONBs was proposed, and the synthesis process was optimized with respect to dependent (sonication power, PVA, and NaHCO3 concentrations) and response (hydrodynamic diameter and oxygen capacity) variables. The optimized formulation has a concentration of (13.8 ± 0.01) × 1010 particles per ml with a hydrodynamic diameter of 142.83 ± 11.46 nm, and oxygen loading capacity of 47.2 ± 2.4 mg L-1. After 4 weeks of storage, the ONBs were found to have an oxygen concentration of 38.9 ± 2.9 mg L-1, indicating excellent oxygen retention capability. The PLGA-ONBs tested in vitro in Muller and R28 retinal cell lines demonstrated excellent biocompatibility and potential to mitigate hypoxia. In addition, the PLGA-ONBs treatment on hypoxic cells demonstrated restoration of mRNA expression of three key hypoxic genes (HIF-1α, PAI-1, and VEGF-A) to normoxic states, indicating hypoxia reversal potential. Biosafety of the PLGA-ONBs was demonstrated in a rabbit model, demonstrating promise in clinical translation. The PLGA-ONBs developed exhibited excellent oxygen loading and retention, potential in hypoxia mitigation, and a safety profile that could be a promising route to treating ischemic diseases of the eye.

Polylactic Acid-Polyglycolic Acid Copolymer

Predicting coarse-grained representations of biogeochemical cycles from metabarcoding data.

MOTIVATION: Taxonomic analysis of environmental microbial communities is now routinely performed thanks to advances in DNA sequencing. Determining the role of these communities in global biogeochemical cycles requires the identification of their metabolic functions, such as hydrogen oxidation, sulfur reduction, and carbon fixation. These functions can be directly inferred from metagenomics data, but in many environmental applications metabarcoding is still the method of choice. The reconstruction of metabolic functions from metabarcoding data and their integration into coarse-grained representations of biogeochemical cycles remains a difficult bioinformatics problem today. RESULTS: We developed a pipeline, called Tabigecy, which exploits taxonomic affiliations to predict metabolic functions constituting biogeochemical cycles. In a first step, Tabigecy uses the tool EsMeCaTa to predict consensus proteomes from input affiliations. To optimize this process, we generated a precomputed database containing information about 2404 taxa from UniProt. The consensus proteomes are searched using bigecyhmm, a newly developed Python package relying on Hidden Markov Models to identify key enzymes involved in metabolic function of biogeochemical cycles. The metabolic functions are then projected on coarse-grained representation of the cycles. We applied Tabigecy to two salt cavern datasets and validated its predictions with microbial activity and hydrochemistry measurements performed on the samples. The results highlight the utility of the approach to investigate the impact of microbial communities on biogeochemical processes. AVAILABILITY AND IMPLEMENTATION: The Tabigecy pipeline is available at https://github.com/ArnaudBelcour/tabigecy. The Python package bigecyhmm and the precomputed EsMeCaTa database are also separately available at https://github.com/ArnaudBelcour/bigecyhmm and https://doi.org/10.5281/zenodo.13354073, respectively.

Metagenomics

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning