PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Metadata”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Reporting and representation of population descriptors in public RNA-seq databases.

Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.

Humans↗

Inclusion of Multi-Omic Biomarkers Improves Prediction Accuracy of Response, Relapse, and Overall Survival in Acute Myeloid Leukemia Patients Receiving High-Intensity Induction Chemotherapy.

BACKGROUND: Despite advancements in genetic markers for acute myeloid leukemia (AML) risk stratification, outcome prediction remains challenging due to disease heterogeneity and dynamic genetic changes, highlighting the need for reliable biomarkers to improve AML treatment strategies and patient outcomes. To refine outcome predictions, we investigated the use of microbial-derived biomarkers to predict composite complete remission (CRc), relapse, and survival for patients on high- and low-intensity regimens, and to integrate those variables into the widely clinically utilized European Leukemia Network (ELN-2022) genetic risk classification model for high-intensity-treated patients. METHODS: We first developed machine learning models that integrate baseline fecal metabolomics, 16S rRNA-based stool microbiome features, and clinical metadata (sex, antibiotic administration, AML somatic mutations, and cytogenetics) from two cohorts of AML patients (n = 83) undergoing remission induction chemotherapy. Univariate tests and sparse canonical correlation analysis were employed for variable selection and to explore fecal metabolite-microbe relationships. A robust machine learning approach using XGBoost was employed, with 100 stratified data splits (80% training, 20% testing) and coarse-to-fine hyperparameter optimization. Variable importance was aggregated across all models to select key predictors. RESULTS: For high-intensity-treated patients, XGBoost models achieved aggregated AUROC scores of 0.719, 0.729, and 0.65 for CRc, relapse, and overall survival, respectively. For low-intensity-treated patients, these models achieved aggregate AUROC scores of 0.945, 0.724, and 0.768 for these same outcomes, respectively. Integrating the biomarkers identified in the high-intensity machine-learning models with the current ELN-2022 AML risk stratification system effectively stratified patients into risk categories, which obtained higher concordance indices and likelihood ratios, demonstrating improved prognostic accuracy for each outcome compared to ELN-2022 alone. CONCLUSIONS: The inclusion of microbial-derived biomarkers serves as a robust prognostic tool to improve outcome prediction in AML patients, highlighting the potential of its integration into AML risk assessment and paving the way for personalized treatment strategies and improved patient outcomes.

Humans↗

Spectral imaging perspective on cytomics.

BACKGROUND: Cytomics involves the analysis of cellular morphology and molecular phenotypes, with reference to tissue architecture and to additional metadata. To this end, a variety of imaging and nonimaging technologies need to be integrated. Spectral imaging is proposed as a tool that can simplify and enrich the extraction of morphological and molecular information. Simple-to-use instrumentation is available that mounts on standard microscopes and can generate spectral image datasets with excellent spatial and spectral resolution; these can be exploited by sophisticated analysis tools. METHODS: This report focuses on brightfield microscopy-based approaches. Cytological and histological samples were stained using nonspecific standard stains (Giemsa; hematoxylin and eosin (H&E)) or immunohistochemical (IHC) techniques employing three chromogens plus a hematoxylin counterstain. The samples were imaged using the Nuance system, a commercially available, liquid-crystal tunable-filter-based multispectral imaging platform. The resulting data sets were analyzed using spectral unmixing algorithms and/or learn-by-example classification tools. RESULTS: Spectral unmixing of Giemsa-stained guinea-pig blood films readily classified the major blood elements. Machine-learning classifiers were also successful at the same task, as well in distinguishing normal from malignant regions in a colon-cancer example, and in delineating regions of inflammation in an H&E-stained kidney sample. In an example of a multiplexed ICH sample, brown, red, and blue chromogens were isolated into separate images without crosstalk or interference from the (also blue) hematoxylin counterstain. CONCLUSION: Cytomics requires both accurate architectural segmentation as well as multiplexed molecular imaging to associate molecular phenotypes with relevant cellular and tissue compartments. Multispectral imaging can assist in both these tasks, and conveys new utility to brightfield-based microscopy approaches.

Animals↗

ProteomeGRID: towards a high-throughput proteomics pipeline through opportunistic cluster image computing for two-dimensional gel electrophoresis.

The quest for high-throughput proteomics has revealed a number of critical issues. Whilst improved two-dimensional gel electrophoresis (2-DE) sample preparation, staining and imaging issues are being actively pursued by industry, reliable high-throughput spot matching and quantification remains a significant bottleneck in the bioinformatics pipeline, thus restricting the flow of data to mass spectrometry through robotic spot excision and protein digestion. To this end, it is important to establish a full multi-site Grid infrastructure for the processing, archival, standardisation and retrieval of proteomic data and metadata. Particular emphasis needs to be placed on large-scale image mining and statistical cross-validation for reliable, fully automated differential expression analysis, and the development of a statistical 2-DE object model and ontology that underpins the emerging HUPO PSI GPS (Human Proteome Organization Proteomics Standards Initiative General Proteomics Standards). The first step towards this goal is to overcome the computational and communications burden entailed by the image analysis of 2-DE gels with Grid enabled cluster computing. This paper presents the proTurbo framework as part of the ProteomeGRID, which utilises Condor cluster management combined with CORBA communications and JPEG-LS lossless image compression for task farming. A novel probabilistic eager scheduler has been developed to minimise make-span, where tasks are duplicated in response to the likelihood of the Condor machines' owners evicting them. A 60 gel experiment was pair-wise image registered (3540 tasks) on a 40 machine Linux cluster. Real-world performance and network overhead was gauged, and Poisson distributed worker evictions were simulated. Our results show a 4:1 lossless and 9:1 near lossless image compression ratio and so network overhead did not affect other users. With 40 workers a 32x speed-up was seen (80% resource efficiency), and the eager scheduler reduced the impact of evictions by 58%.

Algorithms↗

Chemical compound navigator: a web-based chem-BLAST, chemical taxonomy-based search engine for browsing compounds.

A novel technique to annotate, query, and analyze chemical compounds has been developed and is illustrated by using the inhibitor data on HIV protease-inhibitor complexes. In this method, all chemical compounds are annotated in terms of standard chemical structural fragments. These standard fragments are defined by using criteria, such as chemical classification; structural, chemical, or functional groups; and commercial, scientific or common names or synonyms. These fragments are then organized into a data tree based on their chemical substructures. Search engines have been developed to use this data tree to enable query on inhibitors of HIV protease (http://xpdb.nist.gov/hivsdb/hivsdb.html). These search engines use a new novel technique, Chemical Block Layered Alignment of Substructure Technique (Chem-BLAST) to search on the fragments of an inhibitor to look for its chemical structural neighbors. This novel technique to annotate and query compounds lays the foundation for the use of the Semantic Web concept on chemical compounds to allow end users to group, sort, and search structural neighbors accurately and efficiently. During annotation, it enables the attachment of "meaning" (i.e., semantics) to data in a manner that far exceeds the current practice of associating "metadata" with data by creating a knowledge base (or ontology) associated with compounds. Intended users of the technique are the research community and pharmaceutical industry, for which it will provide a new tool to better identify novel chemical structural neighbors to aid drug discovery.

Computational Biology↗

The BioImage Database Project: organizing multidimensional biological images in an object-relational database.

The BioImage Database Project collects and structures multidimensional data sets recorded by various microscopic techniques relevant to modern life sciences. It provides, as precisely as possible, the circumstances in which the sample was prepared and the data were recorded. It grants access to the actual data and maintains links between related data sets. In order to promote the interdisciplinary approach of modern science, it offers a large set of key words, which covers essentially all aspects of microscopy. Nonspecialists can, therefore, access and retrieve significant information recorded and submitted by specialists in other areas. A key issue of the undertaking is to exploit the available technology and to provide a well-defined yet flexible structure for dealing with data. Its pivotal element is, therefore, a modern object relational database that structures the metadata and ameliorates the provision of a complete service. The BioImage database can be accessed through the Internet.

Copyright↗

Radiology considerations for the PREMIUM study: a multicenter randomized controlled trial of abbreviated MRI versus ultrasound for liver cancer screening in cirrhosis.

This paper describes the rationale and radiology considerations in the implementation of the Preventing Liver Cancer Mortality through Imaging with Ultrasound versus MRI (PREMIUM) study. PREMIUM is a multicenter, randomized controlled trial sponsored by the Department of Veterans Affairs comparing dynamic contrast-enhanced (DCE) abbreviated MRI (aMRI) plus serum AFP versus ultrasound (US) plus serum AFP for hepatocellular carcinoma (HCC) screening in patients with cirrhosis. PREMIUM aims to randomize 4,700 participants across over 47 Veterans Affairs Medical Centers to semiannual surveillance for up to eight years, with HCC-related mortality as the primary endpoint. To date, 35 sites have been activated with 1,085 patients randomized. To ensure uniform implementation and reporting of per-protocol screening, the PREMIUM Radiology Workgroup developed standardized imaging protocols, structured LI-RADS-based reporting templates, and a centralized training program for radiologists and technologists. They also perform ongoing quality control on both scans and reports. The aMRI protocols utilize multiphasic post-contrast imaging to allow LI-RADS scoring. A non-contrast-enhanced aMRI protocol is available for participants who develop renal impairment or contrast allergy during the study. US protocols conform to US LI-RADS standards. Structured reporting promotes consistency in documentation of findings, visualization scores, and follow-up recommendations. A centralized Image Repository was established, incorporating advanced de-identification methods to remove metadata and pixel-embedded protected health information from imaging files. More than 20,000 curated liver MRI and US exams are anticipated, supporting both trial outcomes and future radiomics and artificial intelligence research. PREMIUM aims to determine whether screening for HCC with a DCE aMRI protocol reduces HCC-related mortality and also facilitates ancillary studies utilizing the Image Repository.

Abbreviated MRI↗

ICASP: an intensive-care acquisition and signal processing integrated framework.

This paper presents an intensive-care acquisition and signal processing integrated framework in the area of intensive care units. The framework includes nearly all monitored biosignals in the intensive care, along with metadata and processing results. It is structured on two basic applications, i.e., the acquisition and the database one, running in two different PCs that are connected through a local area network, facilitating real-time data exchange between them. The analytical rundown shows that the proposed framework is a serious effort to give a complete clinical condition of a patient and a form of a diagnostic analysis implement in the intensive care by taking in real-time processing.

Hospital Information Systems↗

Integrated ¹H-NMR Metabolomics and Growth Kinetics Uncover Three Distinct Metabolic Scenarios in Lactiplantibacillus pentosus P7 Fermentation of Plant-Derived Prebiotics.

Lactic acid bacteria (LAB) drive a broad range of food and biotechnological fermentations, the outcomes of which depend not only on the bacterial genotype but also on the chemical composition of the fermentation substrate. To resolve how a single strain reorganises chemically distinct plant matrices, we profiled fermentations of Lactiplantibacillus pentosus P7 (GenBank JBLMKZ000000000) on garlic, onion, and kiwifruit extracts prepared in water and 70% ethanol, using growth kinetics combined with solvent-suppressed 500-MHz proton nuclear magnetic resonance metabolomics over 48 h, and integrated the data with whole-genome pathway annotations. Three substrate-specific metabolic scenarios emerged. On garlic, P7 grew vigorously, with the water extract exceeding the de Man-Rogosa-Sharpe reference medium at every time point (peak ΔOD₆₀₀ of 9.38 versus 8.52 at 24 h) and accumulating sorbose, rhamnose, and the aromatic amino acids phenylalanine and tryptophan (3.04- to 3.70-fold increases), providing first metabolic evidence consistent with the strain's four-copy aroE shikimate-dehydrogenase expansion. On onion, the lowest cell density coincided with the highest lactate output of the dataset (5.21-fold rise at 48 h), transient 5-hydroxymethylfurfural reduction, and accumulation of acetoin and 1,3-propanediol, mapping onto a redundant set of pyridine-nucleotide-dependent oxidoreductases and a pdu-independent diol pathway. On kiwifruit, citrate accumulated 8.9-fold at 16 h and then declined, consistent with an intact citCDEFG citrate-lyase operon paired with absence of canonical oxidative tricarboxylic acid enzymes. The optimal extraction solvent was substrate-dependent, water for garlic and ethanol for onion and kiwifruit. Overall, these results show that substrate chemistry, rather than strain identity, dictates which genome-encoded pathways P7 engages, establishing P7 as a versatile, substrate-tunable platform for the functional fermentation and biorefining of furanic-rich substrate streams. Raw NMR data and ISA-Tab metadata are available via MetaboLights with identifier MTBLS14463.

Lactiplantibacillus pentosus↗

Genetic structure correlates with ethnolinguistic diversity in eastern and southern Africa.

African populations are the most diverse in the world yet are sorely underrepresented in medical genetics research. Here, we examine the structure of African populations using genetic and comprehensive multi-generational ethnolinguistic data from the Neuropsychiatric Genetics of African Populations-Psychosis study (NeuroGAP-Psychosis) consisting of 900 individuals from Ethiopia, Kenya, South Africa, and Uganda. We find that self-reported language classifications meaningfully tag underlying genetic variation that would be missed with consideration of geography alone, highlighting the importance of culture in shaping genetic diversity. Leveraging our uniquely rich multi-generational ethnolinguistic metadata, we track language transmission through the pedigree, observing the disappearance of several languages in our cohort as well as notable shifts in frequency over three generations. We find suggestive evidence for the rate of language transmission in matrilineal groups having been higher than that for patrilineal ones. We highlight both the diversity of variation within Africa as well as how within-Africa variation can be informative for broader variant interpretation; many variants that are rare elsewhere are common in parts of Africa. The work presented here improves the understanding of the spectrum of genetic variation in African populations and highlights the enormous and complex genetic and ethnolinguistic diversity across Africa.

Africa, Southern↗

Clinical sequelae of gut microbiome development and disruption in hospitalized preterm infants.

Aberrant preterm infant gut microbiota assembly predisposes to early-life disorders and persistent health problems. Here, we characterize gut microbiome dynamics over the first 3 months of life in 236 preterm infants hospitalized in three neonatal intensive care units using shotgun metagenomics of 2,512 stools and metatranscriptomics of 1,381 stools. Strain tracking, taxonomic and functional profiling, and comprehensive clinical metadata identify Enterobacteriaceae, enterococci, and staphylococci as primarily exploiting available niches to populate the gut microbiome. Clostridioides difficile lineages persist between individuals in single centers, and Staphylococcus epidermidis lineages persist within and, unexpectedly, between centers. Collectively, antibiotic and non-antibiotic medications influence gut microbiome composition to greater extents than maternal or baseline variables. Finally, we identify a persistent low-diversity gut microbiome in neonates who develop necrotizing enterocolitis after day of life 40. Overall, we comprehensively describe gut microbiome dynamics in response to medical interventions in preterm, hospitalized neonates.

Humans↗

Customizable neuroinformatics database system: XooNIps and its application to the pupil platform.

The developing field of neuroinformatics includes technologies for the collection and sharing of neuro-related digital resources. These resources will be of increasing value for understanding the brain. Developing a database system to integrate these disparate resources is necessary to make full use of these resources. This study proposes a base database system termed XooNIps that utilizes the content management system called XOOPS. XooNIps is designed for developing databases in different research fields through customization of the option menu. In a XooNIps-based database, digital resources are stored according to their respective categories, e.g., research articles, experimental data, mathematical models, stimulations, each associated with their related metadata. Several types of user authorization are supported for secure operations. In addition to the directory and keyword searches within a certain database, XooNIps searches simultaneously across other XooNIps-based databases on the Internet. Reviewing systems for user registration and for data submission are incorporated to impose quality control. Furthermore, XOOPS modules containing news, forums schedules, blogs and other information can be combined to enhance XooNIps functionality. These features provide better scalability, extensibility, and customizability to the general neuroinformatics community. The application of this system to data, models, and other information related to human pupils is described here.

Computational Biology↗

Whole genome sequence data set of methicillin-resistant Staphylococcus aureus isolated from a milkman associated with cows with subclinical mastitis in Kiruhura district, Uganda.

The whole-genome sequence data set for methicillin-resistant Staphylococcus aureus, which was isolated from a milkman associated with cows with subclinical mastitis in the Kiruhura district of Uganda, is presented here. The assembled genome size was 2822,509 bp, with a 33% GC, 2 Contigs, a Contig N50 of 2818,424, and 1 Contig L50. You can access the genome sequence and related metadata at https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_056782255.1/. This dataset can be used again for resistance gene mapping, comparing genomic analysis, and comprehending genetic diversity among MRSA isolates from Ugandan milkmen.

Antimicrobial-resistant genes Staphylococcus aureu↗

Global reach and sustained engagement of a structured digital education program in medical mycology: an observational analysis of the 2025 ESCMID-EFISG webinar series.

OBJECTIVES: Evaluate the 2025 European Society of Clinical Microbiology and Infectious Diseases-European Fungal Infection Study Group webinar series to assess digital education as a scalable, equitable model for global professional development in medical mycology. METHODS: This observational study analyzed Zoom metadata across 17 webinars (January-December 2025). Metrics included registration, unique viewers, peak concurrent views, attendance rate, and duration. RESULTS: The series recorded 4631 registrations and 1372 unique participants. Median live attendance was 199 (interquartile range [IQR] 138-269), with a 39.3% attendance rate (IQR 33.7-47.5%) and peak concurrent viewership of 165 (IQR 106-229). Median session duration was 108 minutes. Webinars engaged a median of 60 countries (range 26-89) simultaneously, spanning 128 countries globally. Faculty comprised 67 unique experts from 23 countries with balanced gender representation (52.8% men, 47.2% women), of whom 16.4% (n = 11/67) were affiliated with institutions in low- and middle-income countries. CONCLUSION: Structured digital programs achieve wide global reach and sustained engagement. Strong participation in long-form sessions supports implementing Continuing Medical Education accreditation and unrestricted on-demand access to enhance global health equity.

Antimicrobial resistance↗

Combining different standards and different approaches for health information retrieval in a quality-controlled gateway.

Internet as source of information is increasing in preeminence in numerous fields, including health. We describe in this paper the CISMeF project (acronym of Catalogue and Index of French-speaking Medical Sites) which has been designed to help the health information consumers and health professionals to find what they are looking for among the numerous health documents available online. The catalogue is founded on two standards: a set of metadata and a terminology based on the MeSH thesaurus which has the same structure and use as an ontology of the medical domain. The structure of the catalogue allows us to place the project at an overlap between the present Web, which is informal, and the forthcoming Semantic Web. Many features of information retrieval and navigation through the catalogue were developed. These features take into account the kind of the end-user (health professional, medical student, patient). The CISMeF-patients catalogue is a sub-catalogue of CISMeF and is dedicated to the patients and the general public. It shares the same model as CISMeF whereas MEDLINE and MedlinePlus do not. We also propose to couple two approaches (morphological processing and data mining) to help the users by correcting and refining their queries.

France↗

Analysing proteomic data.

The rapid growth of proteomics has been made possible by the development of reproducible 2D gels and biological mass spectrometry. However, despite technical improvements 2D gels are still less than perfectly reproducible and gels have to be aligned so spots for identical proteins appear in the same place. Gels can be warped by a variety of techniques to make them concordant. When gels are manipulated to improve registration, information is lost, so direct methods for gel registration which make use of all available data for spot matching are preferable to indirect ones. In order to identify proteins from gel spots a property or combination of properties that are unique to that protein are required. These can then be used to search databases for possible matches. Molecular mass, pI, amino acid composition and short sequence tags can all be used in database searches. Currently the method of choice for protein identification is mass spectrometry. Proteins are eluted from the gels and cleaved with specific endoproteases to produce a series of peptides of different molecular mass. In peptide mass fingerprinting, the peptide profile of the unknown protein is compared with theoretical peptide libraries generated from sequences in the different databases. Tandem mass spectroscopy (MS/MS) generates short amino acid sequence tags for the individual peptides. These partial sequences combined with the original peptide masses are then used for database searching, greatly improving specificity. Increasingly protein identification from MS/MS data is being fully or partially automated. When working with organisms, which do not have sequenced genomes (the case with most helminths), protein identification by database searching becomes problematical. A number of approaches to cross species protein identification have been suggested, but if the organism being studied is only distantly related to any organism with a sequenced genome then the likelihood of protein identification remains small. The dynamic nature of the proteome means that there really is no such thing as a single representative proteome and a complete set of metadata (data about the data) is going to be required if the full potential of database mining is to be realised in the future.

Animals↗

Comparing statistical and semantic approaches for identifying change from land cover datasets.

In this paper, we examine methods for integrating spatial data which apparently should be comparable because they are of the same data type or theme, but which are incompatible or discordant because the classes of that theme are different. For a variety of reasons including changes in methods, in understanding of the resource, and in policy initiatives in the commissioning of the survey, this problem is widespread in the results of natural resources surveys. We present two generic methods: one method is grounded in a statistical approach using discriminant analysis, and the other exploits the knowledge of experts. We use the context of land cover mapping of Great Britain to explore these approaches for integrating discordant data. We demonstrate that the expert-based approach gives very good levels of identification of locations with incompatible classifications at different times, and gives a much better rate of recognition of change. Some conclusions are made about the need to expand current metadata and data quality reporting to include descriptions of:- data conceptualisations, semantics and ontologies;- who decided and defined what the features of interest in a dataset are, and why. If the benefits of spatial data initiatives such as GRID, E-science and INSPIRE are to be fully realised then some method needs to be found to communicate that information most effectively to the potential user of the data.

Conservation of Natural Resources↗

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3·58 × 10-4 to 9·76 × 10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50 000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus↗