PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genomics and bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans

Comparative genomics of carbapenem-resistant Acinetobacter baumannii isolated from pediatric patients in a tertiary care hospital.

Acinetobacter baumannii is a short gram-negative bacillus, notable for its intrinsic multidrug resistance and genomic plasticity, which facilitates the acquisition of additional resistance genes via mobile genetic elements. Due to its increasing carbapenem resistance, the World Health Organization has classified it as a critical priority pathogen. This study performed a comparative genomic analysis of 20 carbapenem-resistant A. baumannii clinical strains isolated from the Hospital Infantil de México Federico Gómez (CRAB-HIMFG), alongside 11 genomes from other Mexican strains. The pangenome was determined to be open, and core genome single-nucleotide polymorphism-based analysis grouped the CRAB-HIMFG strains within CC758/IC5 and CC92/IC2. A novel sequence type (ST) in the MLST-Pasteur scheme was identified, related to STPas156, and in the MLST-Oxford scheme, associated with STOxf758 and STOxf1054. Virulence and resistance genes comprised 0.61% to 2.23% of the pangenome. Oxacillinase genes and efflux pumps primarily mediated carbapenem resistance, while virulence genes included those encoding biofilm and type IV pili. Capsule typing revealed a correlation with established international clones, IC2 and IC5. Plasmids exhibited high diversity, harboring maintenance modules and toxin-antitoxin systems, with the dissemination of resistance genes linked to insertion sequences. Biofilm formation and twitching motility were not always expressed, as they depend on additional environmental factors. Our study shows that comparative genomics is an essential tool to analyze clinically and epidemiologically significant genomes, providing critical insights into gene distribution, genomic architecture, and horizontal gene transfer mechanisms in microbial populations.IMPORTANCEIn recent years, a reported increase in the mortality rate associated with infections caused by A. baumannii, along with a rise in carbapenem resistance, poses a serious clinical challenge. The WHO considered this microorganism critical for research into alternative therapies and epidemiological surveillance. Despite advances in bioinformatics, genomic studies have yet to fully elucidate the structural rearrangements and secretion systems of A. baumannii. This knowledge gap hinders our understanding of its remarkable genomic plasticity and its ability to acquire and spread resistance and virulence genes through horizontal gene transfer.

Acinetobacter baumannii

Dual RNA isolation from blood: an optimized protocol for host and bacterial RNA purification for dual RNA-sequencing analysis in whole blood sepsis samples.

Dual RNA-sequencing (dual RNA-seq) holds significant promise for deciphering bacterial virulence mechanisms during systemic infections. However, its application in sepsis research is hindered by technical challenges, including a low bacterial burden in blood and limited sample volumes and RNA yield from vulnerable populations, such as neonates. We developed an optimized protocol [dual RNA isolation from blood (DRIB)] for simultaneous stabilization, isolation and purification of high-quality host leukocyte and bacterial RNA from low-volume whole blood samples (0.5 ml). This protocol is compatible with clinical sample collection workflows and high-throughput RNA sequencing. The feasibility of DRIB for dual RNA-seq was validated using a pilot cohort of clinical adult sepsis samples, enabling the investigation of host-bacterial gene expression during sepsis. The DRIB protocol yielded 2.10-6.91 µg of total RNA per clinical sample in our pilot cohort. Dual-species ribosomal RNA (rRNA) depletion and RNA-seq generated 16.6-24.8 million filtered reads per sample, with 63±7% of reads uniquely mapped to host or bacterial sequences. Host genes accounted for 51-68% (8.4-10.9 million) reads, while 0.5-6.7% (79,496-789,808 reads) mapped to bacterial genomes. Bioinformatic analysis revealed that both shared and individual transcriptional patterns were identified in host and bacterial responses, including pathways related to immune metabolism and metal-ion binding. Our optimized DRIB protocol and RNA-seq pipeline effectively captured both host and bacterial RNA transcription in clinical sepsis samples. Expanding this approach to larger cohorts and varying disease timepoints will provide crucial new insights into host-bacterial gene co-expression dynamics in sepsis progression and outcomes.

Humans

Identifying potential novel biomarkers for varicocele: A bioinformatics approach to genomics analysis.

Introduction: Varicocele, characterized by the enlargement of scrotal veins, is a common contributor to male infertility, but its genetic underpinnings remain largely unknown. Aim: The goal of this study is to identify potential biomarkers associated with varicocele in order to better understand its molecular mechanisms. Materials and methods: Using the three primary databases, NCBI, DisGeNET, and OpenTarget, we analyzed gene variants and found 79 pertinent genes associated with varicocele. Protein-protein interaction analysis was performed using STRING and visualized with Cytoscape. Molecular Complex Detection (MCODE) and CytoHubba tools helped identify significant protein clusters. Results: The gene ontology analysis shows that there are 79 proteins involved in the inflammatory process, the regulation of gene expression, and cellular components that play a role in oxidative stress and angiogenesis. Our results revealed three key biomarkers: Interleukin-1 beta (IL1B), B-cell lymphoma 2 (BCL2), and matrix metalloproteinase-9 (MMP-9). These proteins are involved in critical processes, such as inflammation, oxidative stress, angiogenesis, and vascular damage, that are central to the pathophysiology of varicocele. Conclusion: The identification of IL1B, BCL2, and MMP-9 offers new insights into varicocele’s molecular mechanisms and suggests potential targets for diagnostic and therapeutic strategies, advancing personalized treatment approaches for fertility restoration.

Humans

Bioinformatics for the Structural Genomics of Poxviruses.

Poxviruses are large, complex viruses, and their host species are widespread across the tree of life. As a result, the bioinformatics analysis of their genomes can be complex. Here we show how a few helpful tools and strategies can be used to inform the analysis, leading to a better understanding of the structural properties of poxvirus genomes and to a more accurate quality control of, or comparison between, assembled sequences.

Poxviridae

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software

Comparative analysis of chloroplast genomes in ten holly (Ilex) species: insights into phylogenetics and genome evolution.

In order to clarify the chloroplast genomes and structural features of ten Ilex species and provide insights into the phylogeny and genome evolution of the genus Ilex, we conducted a comparative analysis of chloroplast genomes using bioinformatics methods. The chloroplast genomes of ten Ilex species were obtained, and their structural features and variations were compared. The results indicated that all chloroplast genomes in the genus Ilex exhibit a double-stranded circular structure, with sizes ranging from 157,356 to 158,018 bp, showing minimal differences in size. The chloroplast genomes of the ten Ilex species have a relatively conservative gene count, with a total of 134 to 135 genes, including 88 or 89 protein-coding genes, and a conserved number of 8 rRNA genes. Each chloroplast genome contains 3 to 123 SSR (Simple Sequence Repeat) sites, predominantly composed of mononucleotide and trinucleotide repeats, with no detection of pentanucleotide or hexanucleotide repeats. The variation in dispersed repeat sequences among Ilex species is minimal, with a total repeat sequence number ranging from 1 to 14, concentrated in the length range of 30 to 42 base pairs. The expansion and contraction of chloroplast genome boundaries among Ilex species are relatively stable, with only minor variations observed in individual species. Variations in non-coding regions are more pronounced than those in coding regions, with the variability in the Large Single Copy region (LSC) being the highest, while the variability in the Inverted Repeat region A (IRa) is the lowest. The divergence time among Ilex species was estimated using the MCMC-tree module, revealing the evolutionary relationships among these species, their common ancestors, and their differentiation throughout the evolutionary process. The research findings provide a valuable reference for the systematic study and molecular marker development of Ilex plants.

Genome, Chloroplast

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics

pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats.

Animal forensic genetics plays a critical role in criminal investigations by providing crucial evidence through domestic animal individualization and wildlife species identification. While human forensic genetics benefits from standardized short tandem repeats (STR) genotyping systems, animal forensic applications encounter significant challenges, including the limited availability of validated STR markers, the prevalence of error-prone dinucleotide STRs (di-STRs), and insufficient integration of population data. To address these challenges, we developed pSTRminer, an integrated bioinformatic tool that automates genome-wide STR mining and polymorphism evaluation. By applying pSTRminer to domestic cattle (Bos taurus), we identified 775,444 STRs de novo from the reference genome and genotyped them using whole-genome sequencing data from 60 Chinese and 111 African cattle to evaluate polymorphism across diverse genetic backgrounds. This led to the development of the cattle STR database (CSDB), comprising loci with a genotyping success rate&#x2009;&#x2265;&#x2009;40% and polymorphism information content (PIC)&#x2009;&#x2265;&#x2009;0.5. Experimental validation of 30 randomly selected tetranucleotide STRs (tetra-STRs) and 33 di-STRs via next-generation sequencing in a local Chinese cattle population (n&#x2009;=&#x2009;145) confirmed marker reliability. Although tetra-STRs had lower average polymorphism levels, they exhibited significantly lower stutter ratios (p&#x2009;<&#x2009;0.05), providing a viable path for identifying discriminative markers with fewer artifacts. Systematic screening revealed that certain tetra-STRs could surpass di-STRs in polymorphism. In conclusion, pSTRminer provides a scalable framework for developing standardized STR panels, facilitating the identification of robust and informative markers in forensic applications.

Bioinformatic software

Detection of short tandem repeats in the cattle genome: a comparison of bioinformatic tools.

BACKGROUND: Short tandem repeats (STRs) are repetitive DNA sequences with 1&#x2013;6 nucleotide repeat units, exhibiting high polymorphism due to varying repeat counts. STRs are more variable than SNPs and can cause genetic disorders. With population-scale cattle whole-genome sequencing data available, whole-genome STR identification has attracted new interest, but challenges remain due to the lack of standardized methods, sequencing data limitations, and the diversity of STR-calling tools. This study compared six STR-calling tools: HipSTR, GangSTR, and ExpansionHunter for short-read data, and Straglr, RepeatHMM, and LongTR for Oxford Nanopore (ONT) long-read data&#x2014;using sequences from five Holstein cattle (two parent&#x2013;offspring trios with a shared sire). This is the first cattle study to evaluate short- and long-read STR callers using both data types from the same animals. RESULTS: In short-read data, ExpansionHunter identified the highest number of polymorphic STRs (pSTRs) (327,690), followed by HipSTR (205,900) and GangSTR (110,680), with 93,023 loci detected by all three tools. In long-read data, LongTR detected 470,250 pSTRs, RepeatHMM 224,185, and Straglr 90,275, with only 33,253 loci shared among them. Mendelian consistency of STR genotypes in the trio offspring was high (>&#x2009;0.8) for all short-read tools, with HipSTR and GangSTR highest at 0.98. LongTR was the only long-read tool with high consistency (0.88). Short-read tools also showed higher concordance in STR genotypes among themselves than was observed among long-read tools. However, long-read tools had a clear advantage in detecting large STRs. Relative to computational efficiency, HipSTR and GangSTR (short-reads), and LongTR (long-reads) required less memory and shorter runtimes than the other tools. CONCLUSIONS: Tool selection is critical for accurate whole-genome STR identification in cattle. For short-read data, HipSTR showed relatively high Mendelian consistency and concordance compared to the other tools, while ExpansionHunter was able to detect longer STRs but with lower Mendelian consistency. For long-read data, LongTR demonstrated higher consistency and computational efficiency relative to the other tools. Based on these results, HipSTR and LongTR are suggested as preferred options for short-read and ONT long-read datasets, respectively, in cattle STR analysis. These recommendations are based on the metrics observed in this study, and confirmatory analyses across additional breeds, larger sample sizes, and validated truth sets are encouraged.

Animals

Genome-Wide Identification and Bioinformatics Analysis of the FAD Gene Family in Walnut (Juglans regia L.).

Fatty acid desaturase (FAD) is a core catalytic enzyme in plants for the synthesis of unsaturated fatty acids, profoundly affecting plant growth, development, and adaptability to various environmental stresses. The walnut (Juglans regia L.) is an important woody oil tree species, and its kernel is rich in unsaturated fatty acids. Systematic identification of the walnut FAD gene family and analysis of its function are of great significance for revealing the molecular mechanisms underlying unsaturated fatty acid metabolism in the walnut. Based on walnut whole-genome data, this study used homology alignment and hidden Markov model search methods to identify the JrFAD gene family members. Subsequently, a variety of bioinformatics tools were used to systematically analyze their structural characteristics, evolutionary expansion mechanism, expression regulation, and function. A total of 21 JrFAD gene family members were identified and classified into five subfamilies. The family genes were unevenly distributed on nine chromosomes. WGD/segmental duplication was the main expansion method, and the duplicated gene pairs experienced strong purification selection. The family gene promoter sequence is rich in regulatory elements that respond to light, plant hormones, and various stresses. The expression pattern analysis showed that JrFAD3.1 and JrFAD2.3 showed high expression specifically during the rapid accumulation of walnut kernel oil. This study clarified the composition and evolutionary characteristics of the FAD gene family in the walnut, which provides useful information for in-depth analyses of its functional mechanism in the regulation of lipid metabolism, and also identified potential candidate gene resources for the genetic improvement of walnut varieties with high amounts of unsaturated fatty acids.

Juglans

Functional Annotation Routines Used by ABRF Bioinformatics Core Facilities - Observations, Comparisons, and Considerations.

The functional annotation of gene lists is a common analysis routine required for most genomics experiments, and bioinformatics core facilities must support these analyses. In contrast to methods such as the quantitation of RNA-Seq reads or differential expression analysis, our research group noted a lack of consensus in our preferred approaches to functional annotation. To investigate this observation, we selected 4 experiments that represent a range of experimental designs encountered by our cores and analyzed those data with 6 tools used by members of the Association of Biomolecular Resource Facilities (ABRF) Genomic Bioinformatics Research Group (GBIRG). To facilitate comparisons between tools, we focused on a single biological result for each experiment. These results were represented by a gene set, and we analyzed these gene sets with each tool considered in our study to map the result to the annotation categories presented by each tool. In most cases, each tool produces data that would facilitate identification of the selected biological result for each experiment. For the exceptions, Fisher's exact test parameters could be adjusted to detect the result. Because Fisher's exact test is used by many functional annotation tools, we investigated input parameters and demonstrate that, while background set size is unlikely to have a significant impact on the results, the numbers of differentially expressed genes in an annotation category and the total number of differentially expressed genes under consideration are both critical parameters that may need to be modified during analyses. In addition, we note that differences in the annotation categories tested by each tool, as well as the composition of those categories, can have a significant impact on results.

Computational Biology

Alternative splicing in ovarian cancer.

Ovarian cancer is the second leading cause of gynecologic cancer death worldwide, with only 20% of cases detected early due to its elusive nature, limiting successful treatment. Most deaths occur from the disease progressing to advanced stages. Despite advances in chemo- and immunotherapy, the 5-year survival remains below 50% due to high recurrence and chemoresistance. Therefore, leveraging new research perspectives to understand molecular signatures and identify novel therapeutic targets is crucial for improving the clinical outcomes of ovarian cancer. Alternative splicing, a fundamental mechanism of post-transcriptional gene regulation, significantly contributes to heightened genomic complexity and protein diversity. Increased awareness has emerged about the multifaceted roles of alternative splicing in ovarian cancer, including cell proliferation, metastasis, apoptosis, immune evasion, and chemoresistance. We begin with an overview of altered splicing machinery, highlighting increased expression of spliceosome components and associated splicing factors like BUD31, SF3B4, and CTNNBL1, and their relationships to ovarian cancer. Next, we summarize the impact of specific variants of CD44, ECM1, and KAI1 on tumorigenesis and drug resistance through diverse mechanisms. Recent genomic and bioinformatics advances have enhanced our understanding. By incorporating data from The Cancer Genome Atlas RNA-seq, along with clinical information, a series of prognostic models have been developed, which provided deeper insights into how the splicing influences prognosis, overall survival, the immune microenvironment, and drug sensitivity and resistance in ovarian cancer patients. Notably, novel splicing events, such as PIGV|1299|AP and FLT3LG|50,941|AP, have been identified in multiple prognostic models and are associated with poorer and improved prognosis, respectively. These novel splicing variants warrant further functional characterization to unlock the underlying molecular mechanisms. Additionally, experimental evidence has underscored the potential therapeutic utility of targeting alternative splicing events, exemplified by the observation that knockdown of splicing factor BUD31 or antisense oligonucleotide-induced BCL2L12 exon skipping promotes apoptosis of ovarian cancer cells. In clinical settings, bevacizumab, a humanized monoclonal antibody that specifically targets the VEGF-A isoform, has demonstrated beneficial effects in the treatment of patients with advanced epithelial ovarian cancer. In conclusion, this review constitutes the first comprehensive and detailed exposition of the intricate interplay between alternative splicing and ovarian cancer, underscoring the significance of alternative splicing events as pivotal determinants in cancer biology and as promising avenues for future diagnostic and therapeutic intervention.

Humans

Bioinformatics in crop research: using genomic data for crop improvement.

Sustainable crop development aims to maintain or increase yields while reducing environmental impact and managing the challenges imposed by climate change. As the global population grows and arable land becomes scarcer, the integration of molecular breeding with bioinformatics has emerged as an effective strategy for long-term crop improvement. Bioinformatics enables researchers to analyze and interpret the vast quantities of genetic data generated by high-throughput sequencing, making it possible to identify molecular markers, candidate genes, and regulatory networks linked to specific agronomic traits, which breeders then translate into focused, ecologically sustainable breeding programs. This approach has enabled major progress across several fronts: the identification of genes conferring resistance to biotic stressors (pests, pathogens) and abiotic stressors (drought, salinity, heat); the development of nutrient-efficient, low-input crop varieties; the improvement of agronomic performance and nutritional quality through identification of yield- and quality-related genes; and the conservation and deployment of genetic diversity to safeguard long-term breeding sustainability. By combining genomic data with precision breeding techniques, researchers are developing crops that are better adapted to a growing population and a changing climate, positioning the integration of molecular breeding and bioinformatics as a central pillar of future global food security.

bioinformatics

Whole mitogenome profile of Pelung and Sentul chickens to reveal potency of Indonesian livestock genetic resources.

Native chickens are essential genetic resources in Indonesia, providing economic, cultural, and nutritional value with strong adaptability to local environments. Among these native chickens, Sentul and Pelung are recognized as national genetic resources due to dual-purpose and ornamental traits, respectively, but their whole mitogenome characterization remains limited. Therefore, this study aimed to assemble and analyze the whole mitogenome of Sentul and Pelung chickens as well as perform a comparison with 35 additional genomes from domestic chickens and jungle fowls across Asia. The experiment was carried out using next-generation sequencing and bioinformatics-based genome assembly and analysis. The results showed that both Sentul and Pelung genomes were 16,784&#xa0;bp in length and contained the typical mitochondrial gene composition, including 13 protein-coding genes (PCGs), 22 transfer ribonucleic acids (tRNAs), two ribosomal ribonucleic acids (rRNAs), and a control region (D-loop). Furthermore, comparative analysis identified five single-nucleotide polymorphisms (SNPs) distinguishing the two breeds, located in ND1, COX1, COX2, ND4, and the D-loop region. Phylogenetic reconstruction based on whole mitochondrial sequences showed that Sentul and Pelung chickens belong to haplogroup D, alongside other native breeds and red jungle fowls from Indonesia and the Philippines.

Genetic diversity

Genomic Screening for Infants and Reproductive Adults.

Recent progress in genomic sequencing, bioinformatics, cloud computation, and artificial intelligence is advancing a more mature understanding of the architecture of childhood genetic diseases. This knowledge and these technologies are enabling expanded genomic screening of infant and reproductive adult populations. With many new disease-modifying and curative therapies in development and approval processes, there exists unparalleled opportunity to identify, treat, and decrease the population burden of genetic disease and transform medical genetics. Broad implementation of genomic population screening, however, requires investments for overcoming remaining evidence gaps and operational challenges, and for delivery in a sustainable manner that is acceptable to parents, prospective parents, and physicians.

Journal Article

A scalable HPC framework for bioinformatics in resource-limited settings: design principles, implementation, and sustainability from the UVRI experience.

MOTIVATION: Building and sustaining High-Performance Computing (HPC) infrastructure for bioinformatics research in resource-limited settings presents significant technical, financial and operational challenges. Institutions in low-and middle-income regions often face constraints such as limited technical expertise, unstable infrastructure and restricted funding which can hinder the deployment of large-scale computational platforms necessary for modern genomics and bioinformatics analyses. RESULTS: We present a scalable and modular HPC framework developed at the Uganda Virus Research Institute (UVRI) to support large-scale genomics and other omics data analyses in resource-limited settings. The framework integrates open-source HPC management tools, infrastructure automation, and reproducible configuration management to enable reliable deployment and maintenance. Optimized storage and networking configurations combined with a phased capacity-building strategy support high-throughput genomic workflows while strengthening local technical expertise. From our implementation experience, we derive ten practical design and operational rules that provide a transferable methodology for establishing and sustaining in-house HPC infrastructure. These rules emphasize strategic investment in human capacity, structured planning, leveraging collaborations, adoption of open-source technologies and service management practices to improve operational resilience and long-term sustainability. AVAILABILITY: The design principles, automation strategies and implementation guidelines described in this work are applicable to institutions seeking to establish sustainable HPC resources for bioinformatics research in resource-constrained environments.

Computational Biology

Q&A with Mich&#xe8;le Ramsay.

Mich&#xe8;le Ramsay, PhD, is Director of the Sydney Brenner Institute for Molecular Bioscience, Professor in Human Genetics, and South African Research Chair in Genomics and Bioinformatics of African Populations at the University of the Witwatersrand, Johannesburg. While promoting research excellence in Africa and contributing to research that accurately represents African populations in global science, she supports capacity strengthening in the fields of genomics and precision medicine. Mich&#xe8;le is a founding member of the Human Heredity and Health in Africa Consortium, co-chair of the International Health Cohorts Consortium, member of the WHO Technical Advisory Group for Genomics (TAG-G), and co-chair of the Lancet Commission on Precision Health. She contributes low- and middle-income countries' perspectives to global genomics, promoting ethical, equitable and fair principles and practices.

Humans