PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Python”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning↗

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus↗

Neuroanatomical affiliation visualization-interface system.

A number of knowledge management systems have been developed to allow users to have access to large quantity of neuroanatomical data. The advent of three-dimensional (3D) visualization techniques allows users to interact with complex 3D object. In order to better understand the structural and functional organization of the brain, we present Neuroanatomical Affiliations Visualization-Interface System (NAVIS) as the original software to see brain structures and neuroanatomical affiliations in 3D. This version of NAVIS has made use of the fifth edition of "The Rat Brain in Stereotaxic coordinates" (Paxinos and Watson, 2005). The NAVIS development environment was based on the scripting language name Python, using visualization toolkit (VTK) as 3D-library and wxPython for the graphic user interface. The following manuscript is focused on the nucleus of the solitary tract (Sol) and the set of affiliated structures in the brain to illustrate the functionality of NAVIS. The nucleus of the Sol is the primary relay center of visceral and taste information, and consists of 14 distinct subnuclei that differ in cytoarchitecture, chemoarchitecture, connections, and function. In the present study, neuroanatomical projection data of the rat Sol were collected from selected literature in PubMed since 1975. Forty-nine identified projection data of Sol were inserted in NAVIS. The standard XML format used as an input for affiliation data allows NAVIS to update data online and/or allows users to manually change or update affiliation data. NAVIS can be extended to nuclei other than Sol.

Animals↗

REAPER: a project-centric workflow layer for comparative repeatome analysis.

INTRODUCTION: Repeatome characterization from short-read sequencing data is widely performed using RepeatExplorer2/TAREAN. However, long-lived multisample projects and explicit comparative designs are often executed as ad hoc command sequences that are hard to version, rerun, and monitor on shared compute environments - a gap that motivates a project-centric workflow layer for repeatome analysis. METHODS: We present REAPER (Repeatome Extended Analysis Pipeline-Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses. REAPER does not implement a new repeat-discovery algorithm; it is an orchestration layer, and biological accuracy for clustering and satellite calling depends on the underlying RepeatExplorer2/TAREAN and satMiner methods it coordinates. REAPER standardizes: Read QC Deterministic subsampling and preparation RepeatExplorer2/TAREAN execution via seqclust, with satMiner-inspired iterative assembly Post-TAREAN BLAST-based annotation against curated repeat collections (optionally including taxon-scoped NCBI-derived resources with freshness checks) Optional graph-based comparative reports The pipeline makes comparative read allocation, prefix policy, and analysis-ready tables explicit; caching supports incremental reruns and structured logs support monitoring. Performance was assessed using a Triticeae short-read dataset (five samples), with rule-level logging of runtime and memory across pipeline stages. RESULTS: Rule-level performance logs show that graph-based clustering dominates runtime and memory, while QC and preparation steps are lightweight by comparison. Graph-report annotations for the Triticeae project additionally link high-ranking clusters to established repeat markers - including pTa794- and pSc119-class entries in curated databases. DISCUSSION: These findings illustrate biologically interpretable outputs (recovery of known Triticeae repeat markers) alongside quantitative performance metrics (identification of graph-based clustering as the dominant computational cost). By making comparative read allocation, prefix policy, and analysis-ready tables explicit - and by supporting caching and structured logging - REAPER supports reproducible comparative repeatome analysis in evolving multisample projects. As an orchestration layer rather than a discovery algorithm, REAPER's contribution lies in reproducibility, monitorability, and comparative-analysis infrastructure, with biological accuracy remaining contingent on the underlying RepeatExplorer2/TAREAN and satMiner methods.

TAREAN↗

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps↗

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics↗

Can Australians identify snakes?

A study of the ability of Australians to identify snakes was undertaken, in which 558 volunteers (primary and secondary schoolchildren, doctors and university science and medical students) took part. Over all, subjects correctly identified an average of 19% of snakes; 28% of subjects could identify a taipan, 59% could identify a death adder, 18% a tiger snake, 23% an eastern (or common) brown snake, and 0.5% a rough-scaled snake. Eighty-six per cent of subjects who grew up in rural areas could identify a death adder; only 4% of those who grew up in an Australian capital city could identify a nonvenomous python. Male subjects identified snakes more accurately than did female subjects. Doctors and medical students correctly identified an average of 25% of snakes. The ability to identify medically significant Australian snakes was classified according to the observer's background, education, sex, and according to the individual snake species. Australians need to be better educated about snakes indigenous to this country.

Adolescent↗

POISE: Spectral Inference of Parent-of-Origin Effects in Unlabeled Genomic Data.

MOTIVATION: Parent of Origin Effects (POEs), where the effect of an an allele on a phenotype differs based on maternal or paternal inheritance implicated in growth, metabolism, and neurodevelopment. Traditional tests for POEs require family data to determine parental origins of transmitted alleles. Given that such studies are expensive and time consuming compared to genome-wide association studies (GWAS), tests that function absent inheritance information are highly desirable. We develop a method, based on community detection from machine learning, that infers POEs via a spectral decomposition, obtains confidence intervals via a non-parametric bootstrap, and safeguards against confounding by non POE sources of variation. We refer to our method as Parent of Origin Inference via Spectral Estimation (POISE). RESULTS: We demonstrate that POISE is well-calibrated under both Gaussian and heavy-tailed noise in simulation studies, with improved robustness to true POEs compared to existing covariance-based tests. POISE provides per-trait effect estimates with bias-corrected bootstrap confidence intervals and incorporates an information-theoretic minimum detectable effect size that filters unreliable estimates, conferring robustness to covariance-deflating variance QTL. We then apply POISE to GWAS data from the UK Biobank using BMI, LDL cholesterol, and HDL cholesterol. POISE recovers established POE loci and identifies 134 additional variants at genes implicated in lipid metabolism, immune regulation, and growth. AVAILABILITY AND IMPLEMENTATION: The code for this method in Python is available at https://github.com/bystrogenomics/POISE.

Community Detection↗

Phasis: a software tool for register-resolved discovery of plant phased small RNA loci.

Plant PHAS locus discovery remains challenging because phasiRNA-producing loci must be distinguished from other sRNA-producing regions with high abundance or apparent periodicity. This problem is especially acute for reproductive 24-PHAS loci, which occur within genomes that also produce abundant 24-nt siRNAs from nonPHAS regions. We present Phasis, an open-source Python software tool for plant PHAS-locus discovery from small RNA sequencing data. Phasis combines statistical evidence for phased accumulation with locus-level features and a Register-Resolved Locus Interpretation Layer that evaluates whether candidate loci show coherent phased architecture. Across diverse plant datasets, Phasis recovered validated or annotated 21- and 24-PHAS loci with a strong balance between call-level precision and reference-locus recall, and generally outperformed PhaseTank and ShortStack in matched benchmark analyses. The register-resolved interpretation layer reduced unsupported calls by separating coherent phased loci from ambiguous sRNA-producing regions. In maize dcl5 mutant libraries, Phasis showed strong depletion of 24-PHAS recovery, supporting DCL5-dependent recovery of reproductive 24-PHAS signal. Together, these results support Phasis as a biologically interpretable tool for large-scale discovery of plant DCL-dependent phasiRNA loci.

bioinformatics↗

Hematopoiesis in snakes (Ophidia).

Locations of the hematopoietic tissue have been described in the following ophidian species: Bothrops jararaca, Bothrops jararacusu, Waglerophis merremii, Elaphe teniura teniura, Boa constrictor, and Python reticulatus. Studies were carried out on perfusion fixed vertebrae, ribs, spleen, liver, thymus, and kidney. Routine histological technique was applied using both light and electron microscopy. Hematopoietic tissue was found in the following locations of the vertebrae: neural spine, neural arch, postzygophysis processes, hypapophysis, vertebral centre. Moreover, intense hematopoiesis was found inside the ribs. In the spleen and thymus, only lymphopoiesis was found. Hematopoietic islets in the spleen were sporadically found only in young specimens. No hematopoiesis was observed in the liver and kidney. In the studied species, there were no differences in the location of hematopoietic tissue. A new model of mature and immature blood cell release to the lumen of marrow sinuses different from that known to operate in higher vertebrates is proposed.

Animals↗

SNP analysis and presentation in the Pharmacogenetics of Membrane Transporters Project.

The multidisciplinary UCSF Pharmacogenetics of Membrane Transporters project seeks to systematically identify sequence variants in transporters and to determine the functional significance of these variants through evaluation of relevant cellular and clinical phenotypes. The project is structured around four interacting cores: genomics, cellular phenotyping, clinical phenotyping, and bioinformatics. The bioinformatics core is responsible for collecting, storing, and analyzing the information obtained by the other cores and for presenting the results, in particular, for the genomic data. Most of this process is automated using locally developed software written in Python, an open source language well suited for rapid, modular development that meets requirements that are themselves constantly evolving. Here we present the details of transforming ABI trace file data into useful information for project investigators and a description of the types of data analysis and display that we have developed.

Amino Acid Sequence↗

Non-random base composition in codons of mitochondrial cytochrome b gene in vertebrates.

Cytochrome b is the central catalytic subunit of the quinol:cytochrome c oxidoreductase of complex III of the mitochondrial oxidative phosphorylation system and is essential to the viability of most eukaryotic cells. Partial cytochrome b gene sequences of 14 species representing mammals, birds, reptiles and amphibians are presented here including some species typical for Poland. For the analysed species a comparative analysis of the natural variation in the gene was performed. This information has been used to discuss some aspects of gene sequence - protein function relationships. Review of relevant literature indicates that similar comparisons have been made only for basic mammalian species. Moreover, there is little information about the Polish-specific species. We observed that there is a strong non-random distribution of nucleotides in the cytochrome b sequence in all tested species with the highest differences at the third codon position. This is also the codon position of the strongest compositional bias. Some tested species, representing distant systematic groups, showed unique base composition differing from the others. The quail, frog, python and elk prefer C over A in the light DNA strand. Species belonging to the artiodactyls stand out from the remaining ones and contain fewer pyrimidines. The observed overall rate of amino acid identity is about 61%. The region covering Q(o) center as well as histidines 82 and 96 (heme ligands) are totally conserved in all tested species. Additionally, the applied method and the sequences can also be used for diagnostic species identification by veterinary and conservation agencies.

Animals↗

Obstacles to Implementing an Execution Engine for Clinical Guidelines Formalized in GLIF.

This article is on obstacles we faced when developing an executable representation of guidelines formalized the Guideline Interchange Format (GLIF). The GLIF does not fully specify the representation of guidelines at the implementation level as it is focused mainly on the description of guideline's logical structure. Our effort was to develop an executable representation of guidelines formalized in GLIF and to implement a pilot engine, which will be able to process such guidelines. The engine has been designed as a component of the MUltimedia Distributed Record system version 2 (MUDR(2)). When developing executable representation of guidelines we paid special attention to utilisation of existing technologies to achieve the highest reusability.Main implementation areas, which are not fully covered by GLIF, are a data model and an execution language. Concerning the data model we have decided to use MUDR(2)'s native data model for this moment and to keep watching the standardisation of a virtual medical record to implement it in execution engine in the near future. When developing the execution language, first of all we have specified necessities, which the execution language ought to meet. Then we have considered some of the most suitable candidates: Guideline Execution Language (GEL), GELLO, Java and Python. Finally we have chosen GELLO although it does not completely cover all required areas. The main GELLO's advantage is that it is a proposed HL7 standard. In this paper we show some of the most important disadvantages of GELLO as an executable language and how we have solved them.

Decision Making, Computer-Assisted↗

Retrospective prevalence of snakebites from Hospital Kuala Lumpur (HKL) (1999-2003).

A hospital based retrospective study of the prevalence of snakebite cases at Hospital Kuala Lumpur was carried out over a five-year period from 1999 to 2003. A total of 126 snakebite cases were recorded. The highest admission for snakebites was recorded in 2001 (29 cases). The majority of cases were admitted for three days or less (79%). Most of the snakebite cases were reported in the 11-30 years age group (52%). The male:female ratio was 3:1. The majority of cases were Malaysians (80%, 101 cases). Of the non-Malaysians, Indonesians constituted the most (56%, 14 cases). Bites occurred most commonly on the lower limbs (49%), followed by upper limbs (45%) and on other parts of the body (6%). No fatal cases were detected and complications were scarce. In 60% (70 cases) the snake could not be identified. Of the four species of snakes that were identified, cobra (both suspected and confirmed) constituted the largest group (25%), followed by viper (10%), python (4%) and sea snake (1%). The most common clinical presentations were pain and swelling, 92% (116 cases). All patients were put on snakebite charts and their vital signs were monitored. Of the snakebite cases, 48% (61 cases) were treated with cloxacillin and 25% (32 cases) were given polyvalent snake antivenom.

Adolescent↗

Review of sarcocystosis in Malaysia.

Sarcocystis is a tissue coccidian with an obligatory two-host life cycle. The sexual generations of gametogony and sporogony occur in the lamina propria of the small intestine of definitive hosts which shed infective sporocysts in their stools and present with intestinal sarcocystosis. Asexual multiplication occurs in the skeletal and cardiac muscles of intermediate hosts which harbor Sarcocystis cysts in their muscles and present with muscular sarcocystosis. In Malaysia, Sarcocystis cysts have been reported from many domestic and wild animals, including domestic and field rats, moonrats, bandicoots, slow loris, buffalo, and monkey, and man. The known definitive hosts for some species of Sarcocystis are the domestic cat, dog and the reticulated python. Human muscular sarcocystosis in Malaysia is a zoonotic infection acquired by contamination of food or drink with sporocysts shed by definitive hosts. The cysts reported in human muscle resembled those seen in the moonrat, Echinosorex gymnurus, and the long-tailed monkey, Macaca fascicularis. While human intestinal sarcocystosis has not been reported in Malaysia so far, it can be assumed that such cases may not be infrequent in view of the occurrence of Sarcocystis cysts in meat animals, such as buffalo. The overall seroprevalence of 19.8% reported among the main racial groups in Malaysia indicates that sarcocystosis (both the intestinal and muscular forms) may be emerging as a significant food-borne zoonotic infection in the country.

Animals↗

Sero-epidemiologic investigations on brucellosis in the states of Uttar Pradesh (U.P.) and Delhi (India).

Sero-prevalence of brucellosis in man and animals was studied during the years 1976 and 1977. Samples were collected from Hospitals/slaughter houses/livestock farms located in Delhi and different districts of Uttar Pradesh (U.P.). The sera samples tested were from 1685 men, 1607 goats, 438 sheep, 244 pigs, 361 cattle, 551 buffalos, 50 dogs, 318 equine and 43 free living animals. The percentage of seropositivity, excluding doubtful ones, was recorded as: man 0.89, goat 5.53, sheep 3.42, pigs 15.98, cattle 6.37 buffalo 4.9 and equine 12.89. Additionally an evidence of agglutinins was also detected in a python serum sample. It was observed that occupation, age, sex and season had a bearing on the prevalence of the disease.

Animals↗

Host specificity and host range of the genus Sarcocystis in three snake-rodent life cycles.

Three Sarcocystis species with snake-rodent life cycles were studied for their host range and host specificity in systematically related intermediate and definitive hosts. While S. singaporensis and S. villivillosi developed only in murids closely related to the genus Rattus, a third Sarcocystis species (natural definitive host: Bitis nasicornis - syn. Isospora dirumpens ?) showed a broad intermediate host range. This species was found to use several Bitis species as definitive hosts. S. singaporensis and S. villivillosi in contrary produced sporocysts only in three Python species and in Aspidites melanocephalus .

Animals↗