PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Python”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Cleaver: software for identifying taxon specific restriction endonuclease recognition sites.

UNLABELLED: Cleaver is an application for identifying restriction endonuclease recognition sites that occur in some taxa but not in others. Differences in DNA fragment restriction patterns among taxa are the basis for many diagnostic assays for taxonomic identification and are used in procedures for removing the DNA of some taxa from pools of DNA from mixed sources. Cleaver analyses restriction digestion of groups of orthologous DNA sequences simultaneously to allow identification of differences in restriction pattern among the fragments derived from different taxa. AVAILABILITY: Cleaver is freely available without registration from its website (http://cleaver.sourceforge.net/) and can be copied, modified and re-distributed under the terms of the GNU general public licence version2 (http://www.gnu.org/licences/gpl). The program can be run as a script for computers that have Python 2.3 and necessary extra modules installed. This allows it to run on Gnu/Linux, Unix, MacOSX and Windows platforms. Stand-alone executable versions for Windows and MacOSX operating systems are available.

Algorithms↗

SCAssign: a sparky extension for the NMR resonance assignment of aliphatic side-chains of uniformly 13C,15N-labeled large proteins.

UNLABELLED: SCAssign (side-chain assignment) is a Sparky extension written in Python to assist the NMR resonance assignment of aliphatic side-chains of uniformly (13)C,(15)N-labeled large proteins. It is based on a general strategy recently developed in our laboratory that makes use of 4D (13)C,(15)N-edited NOESY, 3D MQ-(H)CC(m)H(m)-TOCSY, and prior backbone assignments. The program runs on all operating systems for which Sparky is available, and is easy to install, setup and use. Not only can it accelerate the assignment process, it also allows assignments of weak NOEs in 4D NOESY, which used to be very difficult with manual approach. AVAILABILITY: The program, in the form of source code, is provided as free download at http://yangdw.science.nus.edu.sg/SCAssign. The website also contains installation guide, user manual and demonstrations recorded in Flash.

Algorithms↗

Global maintenance of histone post-translational modifications during the transition into anoxia in embryos of the annual killifish Austrofundulus limnaeus.

Many organisms have adapted to survive anoxic or hypoxic environments, but the epigenetic responses involved in this successful stress response are not well described in most species. Embryos of the annual killifish Austrofundulus limnaeus have the greatest tolerance to anoxia of all vertebrates, making them a powerful model to study the cellular mechanisms necessary for anoxia tolerance. However, the global histone landscape of this species has never been quantified or explored in relation to stress tolerance. Liquid chromatography-mass spectrometry and a Python bioinformatics workflow were used to identify histones and their post-translational modifications. This pipeline resulted in the detection of 252 unique biologically relevant histone post-translational modifications (hPTMs) (unimod + residue). These PTMs represent 16 types of biologically relevant hPTMs present during both anoxia and normoxia in Wourms' stage 36 embryos. This hPTM library presents an exciting opportunity to study histone modifications across development and in response to environmental stressors. No significant changes in PTM or histone abundance were observed between anoxic and normoxic embryos, suggesting that 24 h of anoxia is not sufficient to induce epigenetic or histone isoform changes at the organismal level. This result is inconsistent with data presented for similar stresses in mammalian cells and thus stabilization of the hPTM landscape may be an adaptation that supports anoxia tolerance.

anoxia↗

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis↗

MACS3: A Peak-calling Platform for Bulk and Single-cell Regulatory Genomics.

Since the original publication of Model-based Analysis for ChIP-Seq (MACS), the software has been widely used to identify enriched genomic regions in ChIP-seq, ATAC-seq, CUT&RUN, DNase-seq, and related regulatory genomics assays. Over the years, MACS has evolved substantially, with MACS version 3 (MACS3) now serving as the actively maintained implementation. MACS3 preserves the core MACS framework for fragment pileup, dynamic local background noise, statistical enrichment testing, and peak refinement, while adding functionality needed for contemporary bulk and single-cell workflows. It supports conventional bulk peak calling, paired-end and fragment-based file formats, modular signal processing, direct analysis of single-cell ATAC-seq fragment files, barcode-restricted pseudobulk and cluster-level peak calling, specialized ATAC-seq and variant-calling modules, as well as command-line and programmatic interfaces. MACS3 is distributed through standard software channels and supported by continuous testing across operating systems, Python versions, and CPU architectures. Here we describe the architecture, current capabilities, and recommended use of MACS3, providing an updated reference for applying the MACS framework in contemporary bulk and single-cell regulatory genomics workflows. MACS3 is open-source software available at https://github.com/macs3-project/MACS.

Bioinformatics software↗

CoDIAC: A comprehensive approach for interaction analysis reveals novel insights into SH2 domain function and regulation.

Protein domains are conserved structural and functional units that serve as building blocks of proteins. Through evolutionary expansion, domain families are represented by multiple members in diverse configurations with other domains, evolving new specificities for their interacting partners. Here, we develop a structure-based interface analysis to comprehensively map domain interfaces from experimental and predicted structures, including interfaces with macromolecules and intraprotein interfaces. We hypothesized that comprehensive contact mapping of domains could yield new insights into domain selectivity, conservation of domain-domain interfaces across proteins, and identify conserved post-translational modifications (PTMs), relative to interaction interfaces, allowing for the inference of specific effects due to PTMs or mutations. We applied this approach to the human SH2 domain family, a modular unit central to phosphotyrosine-mediated signaling, identifying a novel approach to understanding binding selectivity and evidence of coordinated regulation of SH2 domain binding interfaces by tyrosine and serine/threonine phosphorylation and acetylation. These findings suggest multiple signaling systems can regulate protein activity and SH2 domain interactions in a coordinated manner. We provide the extensive features of the human SH2 domain family and this modular approach as an open source Python package for COmprehensive Domain Interface Analysis of Contacts (CoDIAC).

SH2 domains↗

Fast and flexible minimizer digestion with digest.

Minimizer digestion is an increasingly common component of bioinformatics tools, including tools for De Bruijn-Graph assembly and sequence classification. We describe a new open source tool and library to facilitate efficient digestion of genomic sequences. It can produce digests based on the related ideas of minimizers, modimizers or syncmers. Digest uses efficient data structures, scales well to many threads, and produces digests with expected spacings between digested elements. Digest is implemented in C++17 with a Python API, and is available open-source at https://github.com/VeryAmazed/digest.

digestion↗

Inference and visualization of complex genotype-phenotype maps with gpmap-tools.

Understanding how biological sequences give rise to observable traits, that is, how genotype maps to phenotype, is a central goal in biology. Yet our knowledge of genotype-phenotype maps in natural systems is limited due to the high dimensionality of sequence space and the context-dependent effects of mutations. The emergence of Multiplex assays of variant effect (MAVEs), along with large collections of natural sequences, offer new opportunities to empirically characterize these maps at an unprecedented scale. However, tools for statistical and exploratory analysis of these high-dimensional data are still needed. To address this gap, we developed gpmap-tools (https://github.com/cmarti/gpmap-tools), a python library that integrates a series of models for inference, phenotypic imputation, and error estimation from MAVE data or collections of natural sequences in the presence of genetic interactions of every possible order. gpmap-tools also provides methods for summarizing patterns of epistasis and visualization of genotype-phenotype maps containing up to millions of genotypes. To demonstrate its utility, we used gpmap-tools to infer genotype-phenotype maps containing 262,144 variants of the Shine-Dalgarno sequence from both genomic 5'UTR sequences and experimental MAVE data. Visualization of the inferred landscapes consistently revealed high-fitness ridges that link core motifs at different distances from the start codon. In summary, gpmap-tools provides a flexible, interpretable framework for studying complex genotype-phenotype maps, opening new avenues for understanding the architecture of genetic interactions and their evolutionary consequences.

Gaussian process↗

DeNoFo: a file format and toolkit for standardised, comparable de novo gene annotation.

MOTIVATION: De novo genes emerge from previously non-coding regions of the genome, challenging the traditional view that new genes primarily arise through duplication and adaptation of existing ones. Characterised by their rapid evolution and their novel structural properties or functional roles, de novo genes represent a young area of research. Therefore, the field currently lacks established standards and methodologies, leading to inconsistent terminology and challenges in comparing and reproducing results. RESULTS: This work presents a standardised annotation format to document the methodology of de novo gene datasets in a reproducible way. We developed DeNoFo, a toolkit to provide easy access to this format that simplifies annotation of datasets and facilitates comparison across studies. Unifying the different protocols and methods in one standardised format, while providing integration into established file formats, such as fasta or gff, ensures comparability of studies and advances new insights in this rapidly evolving field. AVAILABILITY AND IMPLEMENTATION: DeNoFo is available through the official Python Package Index (PyPI) and at https://github.com/EDohmen/denofo . All tools have a graphical user interface and a command line interface. The toolkit is implemented in Python3, available for all major platforms and installable with pip and uv.

Journal Article↗

Mapping Allosteric Communication in the Nucleosome with Conditional Activity.

The nucleosome core particle (NCP) regulates genome accessibility through dynamic allosteric communication between histone proteins and DNA. Building on the concept of conditional activity introduced by Lin (2016), we use molecular dynamics simulations and develop an open-source Python library, CONDACT (CONDitional ACTivity), to quantify time-resolved kinetic correlations in nucleosome systems. We analyze long-time simulations of the nucleosome core particle, including two different DNA sequences, the Widom-601 and ASP (alpha-satellite palindromic) sequences. By tracking dihedral angle transitions, we identify residues with high dynamical memory and map inter-residue communication pathways across histone subunits and DNA. Our analysis reveals kinetically connected domains involving post-translational modification sites, oncogenic mutation sites, and DNA contact regions, with dynamic coupling observed over distances up to 7.5 nm. These findings offer new insight into the long-range allosteric behavior of the nucleosome and its potential role in regulating chromatin accessibility. Quantifying this allosteric behavior potentially identifies targetable residues and domains for therapeutic intervention.

Journal Article↗

A Systematic Review of Spatial Epidemiological Modeling Approaches Applied During the COVID-19 Pandemic.

BACKGROUND: A wide range of epidemiological modeling approaches have been applied to the SARS-CoV-2 pandemic, which presents an opportunity to assess common approaches applied to specific research questions. Spatial models interrogate how heterogeneities and host movement dynamics influence local and regional patterns of disease, issues that were of great interest for understanding and controlling SARS-CoV-2. OBJECTIVE: Here we present a systematic review of spatial epidemiological modeling approaches of SARS-CoV-2. We describe common themes and highlight unique strategies, providing a foundation for researchers to devise spatial models most appropriate for future pathogens and epidemics. Our review also categorizes the research questions that were addressed with spatial models, highlights parameter estimation techniques, and describes the cyber infrastructure used for model development. METHODS: We conducted a systematic review using Web of Science and a standardized set of keywords, followed by thorough examination of abstracts and full texts to determine which studies met our inclusion criteria. To guide our description and comparisons of models, we developed a Geography, Population, Movement (GPM) framework that conceptualizes the interactions between three distinct subcomponents of any spatial model. The geographic model represents the physical arena in which the model is implemented, the intra-population model describes the transmission and disease processes that occur within distinct spatial units of the geography, and the movement model describes the algorithms that dictate how hosts move among spatial units within the geography. RESULTS: The search identified a total of 193 articles, of which 109 were included in our review. The most abundant intra-population modeling methods were agent-based (47.7%) and compartmental modeling (29.4%) approaches. Movement models ranged in complexity, with the most complex models implementing commuter movement among many points of interest in the geographic arena, which were sometimes parameterized by fine-scale mobility data. Geographic models ranged from describing microcosms, such as single classrooms, all the way up to multi-country models. Of the 63.3% of models studies that specified the programming language used, we detected ten different languages, with Matlab and Python being the most frequent, although only 30.6% of studies provided open-access code for their models. We also described eight specialized software systems that were used to construct agent-based or compartment models of COVID-19. CONCLUSIONS: Our review identified and characterized a variety of spatial modeling strategies and software that were usefully employed to address many relevant epidemiological questions for COVID-19. Future research is needed to quantitatively assess which modeling approaches are most appropriate in specific situations, to answer specific questions, or to apply to certain disease systems. Moreover, future cyberinfrastructure could help to modularize and standardize modeling approaches, which would increase transparency and reproducibility, and which would facilitate a detailed examination of which model attributes relate to model performance in a variety of contexts.

COVID-19↗

The Bioperl toolkit: Perl modules for the life sciences.

The Bioperl project is an international open-source collaboration of biologists, bioinformaticians, and computer scientists that has evolved over the past 7 yr into the most comprehensive library of Perl modules available for managing and manipulating life-science information. Bioperl provides an easy-to-use, stable, and consistent programming interface for bioinformatics application programmers. The Bioperl modules have been successfully and repeatedly used to reduce otherwise complex tasks to only a few lines of code. The Bioperl object model has been proven to be flexible enough to support enterprise-level applications such as EnsEMBL, while maintaining an easy learning curve for novice Perl programmers. Bioperl is capable of executing analyses and processing results from programs such as BLAST, ClustalW, or the EMBOSS suite. Interoperation with modules written in Python and Java is supported through the evolving BioCORBA bridge. Bioperl provides access to data stores such as GenBank and SwissProt via a flexible series of sequence input/output modules, and to the emerging common sequence data storage format of the Open Bioinformatics Database Access project. This study describes the overall architecture of the toolkit, the problem domains that it addresses, and gives specific examples of how the toolkit can be used to solve common life-sciences problems. We conclude with a discussion of how the open-source nature of the project has contributed to the development effort.

Algorithms↗

PHENIX: building new software for automated crystallographic structure determination.

Structural genomics seeks to expand rapidly the number of protein structures in order to extract the maximum amount of information from genomic sequence databases. The advent of several large-scale projects worldwide leads to many new challenges in the field of crystallographic macromolecular structure determination. A novel software package called PHENIX (Python-based Hierarchical ENvironment for Integrated Xtallography) is therefore being developed. This new software will provide the necessary algorithms to proceed from reduced intensity data to a refined molecular model and to facilitate structure solution for both the novice and expert crystallographer.

Algorithms↗

Untangle, a tool for filtering overlapping diffraction patterns from multicrystals.

Standard crystallographic data-processing protocols are based on single-crystal models; data from aggregates of multiple crystals with different orientations are difficult to process. In certain cases, it is possible to separately index the diffraction patterns from the dominant crystals in the aggregate. Untangle is a program designed to identify and eliminate overlapping spots from such patterns in order to improve data quality. The program has a Python core with a simple and highly portable graphical user interface, permitting visual verification of the process and interactive modification of the overlap threshold. The software is available under an open-source license.

Bacterial Proteins↗

IFEFFIT: interactive XAFS analysis and FEFF fitting.

IFEFFIT, an interactive program and scriptable library of XAFS algorithms is presented. The core algorithms of AUTOBK and FEFFIT have been combined with general data manipulation and interactive graphics into a single package. IFEFFIT comes with a command-line program that can be run either interactively or in batch-mode. It also provides a library of functions that can be used easily from C or Fortran, as well as high level scripting languages such as Tcl, Perl and Python. Using this library, a Graphical User Interface for rapid 'online' data analysis is demonstrated. IFEFFIT is freely available with an Open Source license. Outside use, development, and contributions are encouraged.

Journal Article↗

Recent developments in the PHENIX software for automated crystallographic structure determination.

A new software system called PHENIX (Python-based Hierarchical ENvironment for Integrated Xtallography) is being developed for the automation of crystallographic structure solution. This will provide the necessary algorithms to proceed from reduced intensity data to a refined molecular model, and facilitate structure solution for both the novice and expert crystallographer. Here, the features of PHENIXare reviewed and the recent advances in infrastructure and algorithms are briefly described.

Algorithms↗

Shed snake skin and hairless mouse skin as model membranes for human skin during permeation studies.

Difficulties in obtaining and using human skin have tempted many workers to employ animal membranes for percutaneous absorption studies. We have investigated the suitability of two species of snake (Elaphe obsoleta, Python molurus) for this purpose and compared our in vitro experimental results for human skin and for hairless mouse, a currently popular model. The effects of long-term hydration on the membranes were investigated over 8 d using tritiated water as a model permeant. The initial permeability coefficients of all the membranes were similar (0.74-2.2 X 10(-3) cm 2h-1). Although the human and squamate skins did not change significantly over the test period, the permeability of hairless mouse skin increased 37 times. The actions of typical enhancers on the permeabilities of the membranes to a model penetrant 5-fluorouracil (5-FU) were tested using 3% Azone in Tween 20/saline, propylene glycol (PG), 2% Azone in PG, and 5% oleic acid in PG. While the data from snake membranes tended to underestimate the enhancer effects, those from hairless mouse skin greatly overestimated the changes. None of the membranes was a completely reliable model for assessing human percutaneous absorption as modified by accelerants. Pretreatment with acetone did not significantly change the permeability of human or squamate skins to 5-FU, although that of hairless mouse increased twentyfold. An overall conclusion is that, wherever possible, human skin should be used in absorption studies and not hairless mouse or snake skin; otherwise, misleading results may be obtained.

Acetone↗

Human infestation by Ophionyssus natricis snake mite.

A family presented with a papular vesiculo-bullous eruption of the skin, found to be caused by the snake mite, Ophionyssus natricis (Cervais, 1844). A pet python was the primary host. Treatment of the animal and its environment led to clearance of the human skin lesions.

Adult↗