PubMed Health⌕ Search

Biomedical subjects

Alexander Goesmann

Publications and source records attributed to Alexander Goesmann.

At least 19 recordsLinked to original sources

Pithos - a scalable and secure data container for FAIR-compliant research data management in life sciences.

Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.

Biological Science Disciplines↗

LegionProfiler: a computational tool for the identification of virulence factors and classification of Legionella pneumophila serogroup 1 isolates.

SUMMARY: Legionella pneumophila has significantly contributed to multiple cases of pneumonia with a high rate of mortality globally. Its ability to exploit host mechanisms through several expressed virulence factors poses challenges for diagnosis, treatment, and outbreak control. To address this, we developed LegionProfiler, a computational tool that swiftly identifies virulence factor protein domains within genome assemblies of Legionella pneumophila serogroup 1 isolates and classifies them into high- or low-virulence groups. LegionProfiler automates the probing of genome assemblies for virulence-associated protein domains and determines the isolate's potential to cause severe pneumonia infection. The LegionProfiler workflow is made available through a user-friendly interface to enhance technical control of infectious sources and adds important insights to the general epidemiology of clinical isolates. It could also support the development of targeted therapeutic strategies that will improve patient treatment. AVAILABILITY AND IMPLEMENTATION: LegionProfiler is freely accessible as a web service at https://legionprofiler.uni-muenster.de, and can also be run locally in a Docker container. The source code can be found at https://imigitlab.uni-muenster.de/heiderlab/legionprofiler or at Zenodo (DOI: 10.5281/zenodo.15592325).

Legionella pneumophila↗

Development of bioinformatic tools to support EST-sequencing, in silico- and microarray-based transcriptome profiling in mycorrhizal symbioses.

The great majority of terrestrial plants enters a beneficial arbuscular mycorrhiza (AM) or ectomycorrhiza (ECM) symbiosis with soil fungi. In the SPP 1084 "MolMyk: Molecular Basics of Mycorrhizal Symbioses", high-throughput EST-sequencing was performed to obtain snapshots of the plant and fungal transcriptome in mycorrhizal roots and in extraradical hyphae. To focus activities, the interactions between Medicago truncatula and Glomus intraradices as well as Populus tremula and Amanita muscaria were selected as models for AM and ECM symbioses, respectively. Together, almost, 20.000 expressed sequence tags (ESTs) were generated from different random and suppressive subtractive hybridization (SSH) cDNA libraries, providing a comprehensive overview of the mycorrhizal transcriptome. To automatically cluster and annotate EST-sequences, the BioMake and SAMS software tools were developed. In connection with the eNorthern software SteN, plant genes with a predicted mycorrhiza-induced expression were identified. To support experimental transcriptome profiling, macro- and microarray tools have been constructed for the two model mycorrhizae, based either on PCR-amplified cDNAs or 70mer oligonucleotides. These arrays were used to profile the transcriptome of AM and ECM roots under different conditions, and the data obtained were uploaded to the ArrayLIMS and EMMA databases that are designed to store and evaluate expression profiles from DNA arrays. Together, the EST- and transcriptome databases can be mined to identify candidate genes for targeted functional studies.

Computational Biology↗

Complete genome of the mutualistic, N2-fixing grass endophyte Azoarcus sp. strain BH72.

Azoarcus sp. strain BH72, a mutualistic endophyte of rice and other grasses, is of agrobiotechnological interest because it supplies biologically fixed nitrogen to its host and colonizes plants in remarkably high numbers without eliciting disease symptoms. The complete genome sequence is 4,376,040-bp long and contains 3,992 predicted protein-coding sequences. Genome comparison with the Azoarcus-related soil bacterium strain EbN1 revealed a surprisingly low degree of synteny. Coding sequences involved in the synthesis of surface components potentially important for plant-microbe interactions were more closely related to those of plant-associated bacteria. Strain BH72 appears to be 'disarmed' compared to plant pathogens, having only a few enzymes that degrade plant cell walls; it lacks type III and IV secretion systems, related toxins and an N-acyl homoserine lactones-based communication system. The genome contains remarkably few mobile elements, indicating a low rate of recent gene transfer that is presumably due to adaptation to a stable, low-stress microenvironment.

Azoarcus↗

Whole-genome sequence of Listeria welshimeri reveals common steps in genome reduction with Listeria innocua as compared to Listeria monocytogenes.

We present the complete genome sequence of Listeria welshimeri, a nonpathogenic member of the genus Listeria. Listeria welshimeri harbors a circular chromosome of 2,814,130 bp with 2,780 open reading frames. Comparative genomic analysis of chromosomal regions between L. welshimeri, Listeria innocua, and Listeria monocytogenes shows strong overall conservation of synteny, with the exception of the translocation of an F(o)F(1) ATP synthase. The smaller size of the L. welshimeri genome is the result of deletions in all of the genes involved in virulence and of "fitness" genes required for intracellular survival, transcription factors, and LPXTG- and LRR-containing proteins as well as 55 genes involved in carbohydrate transport and metabolism. In total, 482 genes are absent from L. welshimeri relative to L. monocytogenes. Of these, 249 deletions are commonly absent in both L. welshimeri and L. innocua, suggesting similar genome evolutionary paths from an ancestor. We also identified 311 genes specific to L. welshimeri that are absent in the other two species, indicating gene expansion in L. welshimeri, including horizontal gene transfer. The species L. welshimeri appears to have been derived from early evolutionary events and an ancestor more compact than L. monocytogenes that led to the emergence of nonpathogenic Listeria spp.

Chromosomes, Bacterial↗

Genome sequence of the ubiquitous hydrocarbon-degrading marine bacterium Alcanivorax borkumensis.

Alcanivorax borkumensis is a cosmopolitan marine bacterium that uses oil hydrocarbons as its exclusive source of carbon and energy. Although barely detectable in unpolluted environments, A. borkumensis becomes the dominant microbe in oil-polluted waters. A. borkumensis SK2 has a streamlined genome with a paucity of mobile genetic elements and energy generation-related genes, but with a plethora of genes accounting for its wide hydrocarbon substrate range and efficient oil-degradation capabilities. The genome further specifies systems for scavenging of nutrients, particularly organic and inorganic nitrogen and oligo-elements, biofilm formation at the oil-water interface, biosurfactant production and niche-specific stress responses. The unique combination of these features provides A. borkumensis SK2 with a competitive edge in oil-polluted environments. This genome sequence provides the basis for the future design of strategies to mitigate the ecological damage caused by oil spills.

Base Sequence↗

The subsystems approach to genome annotation and its use in the project to annotate 1000 genomes.

The release of the 1000th complete microbial genome will occur in the next two to three years. In anticipation of this milestone, the Fellowship for Interpretation of Genomes (FIG) launched the Project to Annotate 1000 Genomes. The project is built around the principle that the key to improved accuracy in high-throughput annotation technology is to have experts annotate single subsystems over the complete collection of genomes, rather than having an annotation expert attempt to annotate all of the genes in a single genome. Using the subsystems approach, all of the genes implementing the subsystem are analyzed by an expert in that subsystem. An annotation environment was created where populated subsystems are curated and projected to new genomes. A portable notion of a populated subsystem was defined, and tools developed for exchanging and curating these objects. Tools were also developed to resolve conflicts between populated subsystems. The SEED is the first annotation environment that supports this model of annotation. Here, we describe the subsystem approach, and offer the first release of our growing library of populated subsystems. The initial release of data includes 180 177 distinct proteins with 2133 distinct functional roles. This data comes from 173 subsystems and 383 different organisms.

Acyl Coenzyme A↗

Complete genome sequence and analysis of the multiresistant nosocomial pathogen Corynebacterium jeikeium K411, a lipid-requiring bacterium of the human skin flora.

Corynebacterium jeikeium is a "lipophilic" and multidrug-resistant bacterial species of the human skin flora that has been recognized with increasing frequency as a serious nosocomial pathogen. Here we report the genome sequence of the clinical isolate C. jeikeium K411, which was initially recovered from the axilla of a bone marrow transplant patient. The genome of C. jeikeium K411 consists of a circular chromosome of 2,462,499 bp and the 14,323-bp bacteriocin-producing plasmid pKW4. The chromosome of C. jeikeium K411 contains 2,104 predicted coding sequences, 52% of which were considered to be orthologous with genes in the Corynebacterium glutamicum, Corynebacterium efficiens, and Corynebacterium diphtheriae genomes. These genes apparently represent the chromosomal backbone that is conserved between the four corynebacteria. Among the genes that lack an ortholog in the known corynebacterial genomes, many are located close to transposable elements or revealed an atypical G+C content, indicating that horizontal gene transfer played an important role in the acquisition of genes involved in iron and manganese homeostasis, in multidrug resistance, in bacterium-host interaction, and in virulence. Metabolic analyses of the genome sequence indicated that the "lipophilic" phenotype of C. jeikeium most likely originates from the absence of fatty acid synthase and thus represents a fatty acid auxotrophy. Accordingly, both the complete gene repertoire and the deduced lifestyle of C. jeikeium K411 largely reflect the strict dependence of growth on the presence of exogenous fatty acids. The predicted virulence factors of C. jeikeium K411 are apparently involved in ensuring the availability of exogenous fatty acids by damaging the host tissue.

Anti-Bacterial Agents↗

Insights into genome plasticity and pathogenicity of the plant pathogenic bacterium Xanthomonas campestris pv. vesicatoria revealed by the complete genome sequence.

The gram-negative plant-pathogenic bacterium Xanthomonas campestris pv. vesicatoria is the causative agent of bacterial spot disease in pepper and tomato plants, which leads to economically important yield losses. This pathosystem has become a well-established model for studying bacterial infection strategies. Here, we present the whole-genome sequence of the pepper-pathogenic Xanthomonas campestris pv. vesicatoria strain 85-10, which comprises a 5.17-Mb circular chromosome and four plasmids. The genome has a high G+C content (64.75%) and signatures of extensive genome plasticity. Whole-genome comparisons revealed a gene order similar to both Xanthomonas axonopodis pv. citri and Xanthomonas campestris pv. campestris and a structure completely different from Xanthomonas oryzae pv. oryzae. A total of 548 coding sequences (12.2%) are unique to X. campestris pv. vesicatoria. In addition to a type III secretion system, which is essential for pathogenicity, the genome of strain 85-10 encodes all other types of protein secretion systems described so far in gram-negative bacteria. Remarkably, one of the putative type IV secretion systems encoded on the largest plasmid is similar to the Icm/Dot systems of the human pathogens Legionella pneumophila and Coxiella burnetii. Comparisons with other completely sequenced plant pathogens predicted six novel type III effector proteins and several other virulence factors, including adhesins, cell wall-degrading enzymes, and extracellular polysaccharides.

Adhesins, Bacterial↗

BACCardI--a tool for the validation of genomic assemblies, assisting genome finishing and intergenome comparison.

SUMMARY: We provide the graphical tool BACCardI for the construction of virtual clone maps from standard assembler output files or BLAST based sequence comparisons. This new tool has been applied to numerous genome projects to solve various problems including (a) validation of whole genome shotgun assemblies, (b) support for contig ordering in the finishing phase of a genome project, and (c) intergenome comparison between related strains when only one of the strains has been sequenced and a large insert library is available for the other. The BACCardI software can seamlessly interact with various sequence assembly packages. MOTIVATION: Genomic assemblies generated from sequence information need to be validated by independent methods such as physical maps. The time-consuming task of building physical maps can be circumvented by virtual clone maps derived from read pair information of large insert libraries.

Algorithms↗

Development of joint application strategies for two microbial gene finders.

MOTIVATION: As a starting point in annotation of bacterial genomes, gene finding programs are used for the prediction of functional elements in the DNA sequence. Due to the faster pace and increasing number of genome projects currently underway, it is becoming especially important to have performant methods for this task. RESULTS: This study describes the development of joint application strategies that combine the strengths of two microbial gene finders to improve the overall gene finding performance. Critica is very specific in the detection of similarity-supported genes as it uses a comparative sequence analysis-based approach. Glimmer employs a very sophisticated model of genomic sequence properties and is sensitive also in the detection of organism-specific genes. Based on a data set of 113 microbial genome sequences, we optimized a combined application approach using different parameters with relevance to the gene finding problem. This results in a significant improvement in specificity while there is similarity in sensitivity to Glimmer. The improvement is especially pronounced for GC rich genomes. The method is currently being applied for the annotation of several microbial genomes. AVAILABILITY: The methods described have been implemented within the gene prediction component of the GenDB genome annotation system.

Algorithms↗

A predator unmasked: life cycle of Bdellovibrio bacteriovorus from a genomic perspective.

Predatory bacteria remain molecularly enigmatic, despite their presence in many microbial communities. Here we report the complete genome of Bdellovibrio bacteriovorus HD100, a predatory Gram-negative bacterium that invades and consumes other Gram-negative bacteria. Its surprisingly large genome shows no evidence of recent gene transfer from its prey. A plethora of paralogous gene families coding for enzymes, such as hydrolases and transporters, are used throughout the life cycle of B. bacteriovorus for prey entry, prey killing, and the uptake of complex molecules.

Adenosine Triphosphate↗

Antibiotic multiresistance plasmid pRSB101 isolated from a wastewater treatment plant is related to plasmids residing in phytopathogenic bacteria and carries eight different resistance determinants including a multidrug transport system.

Ten different antibiotic resistance plasmids conferring high-level erythromycin resistance were isolated from an activated sludge bacterial community of a wastewater treatment plant by applying a transformation-based approach. One of these plasmids, designated pRSB101, mediates resistance to tetracycline, erythromycin, roxythromycin, sulfonamides, cephalosporins, spectinomycin, streptomycin, trimethoprim, nalidixic acid and low concentrations of norfloxacin. Plasmid pRSB101 was completely sequenced and annotated. Its size is 47 829 bp. Conserved synteny exists between the pRSB101 replication/partition (rep/par) module and the pXAC33-replicon from the phytopathogen Xanthomonas axonopodis pv. citri. The second pRSB101 backbone module encodes a three-Mob-protein type mobilization (mob) system with homology to that of IncQ-like plasmids. Plasmid pRSB101 is mobilizable with the help of the IncP-1alpha plasmid RP4 providing transfer functions in trans. A 20 kb resistance region on pRSB101 is located within an integron-containing Tn402-like transposon. The variable region of the class 1 integron carries the genes dhfr1 for a dihydrofolate reductase, aadA2 for a spectinomycin/streptomycin adenylyltransferase and bla(TLA-2) for a so far unknown Ambler class A extended spectrum beta-lactamase. The integron-specific 3'-segment (qacEDelta1-sul1-orf5Delta) is connected to a macrolide resistance operon consisting of the genes mph(A) (macrolide 2'-phosphotransferase I), mrx (hydrophobic protein of unknown function) and mphR(A) (regulatory protein). Finally, a putative mobile element with the tetracycline resistance genes tetA (tetracycline efflux pump) and tetR was identified upstream of the Tn402-specific transposase gene tniA. The second 'genetic load' region on pRSB101 harbours four distinct mobile genetic elements, another integron belonging to a new class and footprints of two more transposable elements. A tripartite multidrug (MDR) transporter consisting of an ATP-binding-cassette (ABC)-type ATPase and permease, and an efflux membrane fusion protein (MFP) of the RND-family is encoded between the replication/partition and the mobilization module. Homologues of the macrolide resistance genes mph(A), mrx and mphR(A) were detected on eight other erythromycin resistance-plasmids isolated from activated sludge bacteria. Plasmid pRSB101-like repA amplicons were also obtained from plasmid-DNA preparations of the final effluents of the wastewater treatment plant indicating that pRSB101-like plasmids are released with the final effluents into the environment.

ATP Binding Cassette Transporter, Subfamily B↗

The genome sequence of Mycoplasma mycoides subsp. mycoides SC type strain PG1T, the causative agent of contagious bovine pleuropneumonia (CBPP).

Mycoplasma mycoides subsp. mycoidesSC (MmymySC)is the etiological agent of contagious bovine pleuropneumonia (CBPP), a highly contagious respiratory disease in cattle. The genome of Mmymy SC type strain PG1(T) has been sequenced to map all the genes and to facilitate further studies regarding the cell function of the organism and CBPP. The genome is characterized by a single circular chromosome of 1211703 bp with the lowest G+C content (24 mole%)and the highest density of insertion sequences (13% of the genome size)of all sequenced bacterial genomes. The genome contains 985 putative genes, of which 72 are part of insertion sequences and encode transposases. Anomalies in the GC-skew pattern and the presence of large repetitive sequences indicate a high genomic plasticity. A variety of potential virulence factors was identified, including genes encoding putative variable surface proteins and enzymes and transport proteins responsible for the production of hydrogen peroxide and the capsule, which is believed to have toxic effects on the animal.

Animals↗

Building a BRIDGE for the integration of heterogeneous data from functional genomics into a platform for systems biology.

The flood of data acquired from the increasing number of publicly available genomes has led to new demands for bioinformatics software. With the growing amount of information resulting from high throughput experiments new questions arise that often focus on the comparison of genes, genomes, and their expression profiles. Inferring new knowledge by combining different kinds of "post-genomics" data obviously necessitates the development of new approaches that allow the integration of variable data sources into a flexible framework. In this paper, we describe our concept for the integration of heterogeneous data into a platform for systems biology. We have implemented a Bioinformatics Resource for the Integration of heterogeneous Data from Genomic Explorations (BRIDGE) and illustrate the usability of our approach as a platform for systems biology for two sample applications.

Algorithms↗

Whole genome shotgun sequencing guided by bioinformatics pipelines--an optimized approach for an established technique.

While the sequencing of bacterial genomes has become a routine procedure at major sequencing centers, there are still a number of genome projects at small- or medium-size facilities. For these facilities a maximum of control over sequencing, assembling and finishing is essential. At the same time, facilities have to be able to co-operate at minimum costs for the overall project. We have established a pipeline for the distributed sequencing of Alcanivorax borkumensis SK2, Azoarcus sp. BH72, Clavibacter michiganensis subsp. michiganensis NCPPB382, Sorangium cellulosum So ce56 and Xanthomonas campestris pv. vesicatoria 85-10. Our pipeline relies on standard tools (e.g. PHRED/PHRAP, CAP3 and Consed/Autofinish) wherever possible, supplementing them with new tools (BioMake and BACCardI) to achieve the aims described above.

Algorithms↗

Bioinformatics support for high-throughput proteomics.

In the "post-genome" era, mass spectrometry (MS) has become an important method for the analysis of proteome data. The rapid advancement of this technique in combination with other methods used in proteomics results in an increasing number of high-throughput projects. This leads to an increasing amount of data that needs to be archived and analyzed. To cope with the need for automated data conversion, storage, and analysis in the field of proteomics, the open source system ProDB was developed. The system handles data conversion from different mass spectrometer software, automates data analysis, and allows the annotation of MS spectra (e.g. assign gene names, store data on protein modifications). The system is based on an extensible relational database to store the mass spectra together with the experimental setup. It also provides a graphical user interface (GUI) for managing the experimental steps which led to the MS data. Furthermore, it allows the integration of genome and proteome data. Data from an ongoing experiment was used to compare manual and automated analysis. First tests showed that the automation resulted in a significant saving of time. Furthermore, the quality and interpretability of the results was improved in all cases.

Algorithms↗

EMMA: a platform for consistent storage and efficient analysis of microarray data.

As a high throughput technique, microarray experiments produce large data sets, consisting of measured data, laboratory protocols, and experimental settings. We have implemented the open source platform EMMA to store and analyze these data. The system provides automated pipelines for data processing and has a modular architecture that can be easily extended. EMMA features detailed reports about spots and their corresponding measurements. In addition to routine data analysis algorithms, the system can be integrated with other components that contain additional data sources (e.g. genome annotation systems).

Algorithms↗