PubMed Health⌕ Search

Biomedical subjects

Clemens Kreutz

Publications and source records attributed to Clemens Kreutz.

6 recordsLinked to original sources

mmContext: an open framework for multimodal contrastive learning of omics and text data.

SUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493.

Computational Biology↗

Genome-wide analysis of DNA copy number changes and LOH in CLL using high-density SNP arrays.

Recurrent genomic aberrations are important prognostic parameters in chronic lymphocytic leukemia (CLL). High-resolution 10k and 50k Affymetrix SNP arrays were evaluated as a diagnostic tool for CLL and revealed chromosomal imbalances in 65.6% and 81.5% of 70 consecutive cases, respectively. Among the prognostically important aberrations, the del13q14 was present in 36 (51.4%), trisomy 12 in 9 (12.8%), del11q22 in 9 (12.8%), and del17p13 in 4 cases (5.7%). A prominent clustering of breakpoints on both sides of the MIRN15A/MIRN16-1 genes indicated the presence of recombination hot spots in the 13q14 region. Patients with a monoallelic del13q14 had slower lymphocyte growth kinetics (P=.002) than patients with biallelic deletions. In 4 CLL cases with unmutated VH genes, a common minimal 3.5-Mb gain of 2p16 spanning the REL and BCL11A oncogenes was identified, implicating these genes in the pathogenesis of CLL. Twenty-four large (>10 Mb) copy-neutral regions with loss of heterozygosity were identified in 14 cases. These regions with loss of heterozygosity are not detectable by alternative methods and may harbor novel imprinted genes or loss-of-function alleles that may be important for the pathogenesis of CLL. Genomic profiling with SNP arrays is a convenient and efficient screening method for simultaneous genome-wide detection of chromosomal aberrations.

Adult↗

Gene profiling of polycystic kidneys.

BACKGROUND: While the genetic basis of autosomal dominant polycystic kidney disease (ADPKD) has been clearly established, the pathogenesis of renal failure in ADPKD remains elusive. Cyst formation originates from proliferating renal tubular epithelial cells that de-differentiate. Fluid secretion with cyst expansion and reactive changes in the extracellular matrix composition combined with increased apoptosis and proliferation rates have been implicated in cystogenesis. METHODS: To identify genes that characterize pathogenical changes in ADPKD, we compared the expression profiles of 12 ADPKD kidneys, 13 kidneys with chronic transplant nephropathy and 16 normal kidneys using a 7 k cDNA microarray. RT-PCR and immunohistochemical techniques were used to confirm the microarray data. RESULTS: Hierarchical clustering revealed that the gene expression profiles of normal, ADPKD and rejected kidneys were clearly distinct. A total of 87 genes were specifically regulated in ADPKD; 26 of these 87 genes were typical for smooth muscle, suggesting epithelial-to-myofibroblast transition (EMT) as a pathogenetic factor in ADPKD. Immunohistology revealed that smooth muscle actin, a typical marker for myofibroblast transition, and caldesmon were mainly expressed in the interstitium of ADPKD kidneys. In contrast, up-regulated keratin 19 and fibulin-1 were confined to cystic epithelia. CONCLUSION: Our results show that the end stage of ADPKD is associated with increased markers of EMT, suggesting that EMT contributes to the progressive loss of renal function in ADPKD.

Cluster Analysis↗

Host cell responses induced by hepatitis C virus binding.

Initiation of hepatitis C virus (HCV) infection is mediated by docking of the viral envelope to the hepatocyte cell surface membrane followed by entry of the virus into the host cell. Aiming to elucidate the impact of this interaction on host cell biology, we performed a genomic analysis of the host cell response following binding of HCV to cell surface proteins. As ligands for HCV-host cell surface interaction, we used recombinant envelope glycoproteins and HCV-like particles (HCV-LPs) recently shown to bind or enter hepatocytes and human hepatoma cells. Gene expression profiling of HepG2 hepatoma cells following binding of E1/E2, HCV-LPs, and liver tissue samples from HCV-infected individuals was performed using a 7.5-kd human cDNA microarray. Cellular binding of HCV-LPs to hepatoma cells resulted in differential expression of 565 out of 7,419 host cell genes. Examination of transcriptional changes revealed a broad and complex transcriptional program induced by ligand binding to target cells. Expression of several genes important for innate immune responses and lipid metabolism was significantly modulated by ligand-cell surface interaction. To assess the functional relevance and biological significance of these findings for viral infection in vivo, transcriptional changes were compared with gene expression profiles in liver tissue samples from HCV-infected patients or controls. Side-by-side analysis revealed that the expression of 27 genes was similarly altered following HCV-LP binding in hepatoma cells and viral infection in vivo. In conclusion, HCV binding results in a cascade of intracellular signals modulating target gene expression and contributing to host cell responses in vivo. Reprogramming of cellular gene expression induced by HCV-cell surface interaction may be part of the viral strategy to condition viral entry and replication and escape from innate host cell responses.

Antigens, Viral↗

Gene expression profiling in polycythaemia vera: overexpression of transcription factor NF-E2.

Summary The molecular aetiology of polycythaemia vera (PV) remains unknown and the differential diagnosis between PV and secondary erythrocytosis (SE) can be challenging. Gene expression profiling can identify candidates involved in the pathophysiology of PV and generate a molecular signature to aid in diagnosis. We thus performed cDNA microarray analysis on 40 PV and 12 SE patients. Two independent data sets were obtained: using a two-step training/validation design, a set of 64 genes (class predictors) was determined, which correctly discriminated PV from SE patients. Separately 253 genes were identified to be upregulated and 391 downregulated more than 1.5-fold in PV compared with healthy controls (P < 0.01). Of the genes overexpressed in PV, 27 contained Sp1 sites: we therefore propose that altered activity of Sp1-like transcription factors may contribute to the molecular aetiology of PV. One Sp1 target, the transcription factor NF-E2 [nuclear factor (erythroid-derived 2)], is overexpressed 2- to 40-fold in PV patients. In PV bone marrow, NF-E2 is overexpressed in megakaryocytes, erythroid and granulocytic precursors. It has been shown that overexpression of NF-E2 leads to the development of erythropoietin-independent erythroid colonies and that ectopic NF-E2 expression can reprogram monocytic cells towards erythroid and megakaryocytic differentiation. Transcription factor concentration may thus control lineage commitment. We therefore propose that elevated concentrations of NF-E2 in PV patients lead to an overproduction of erythroid and, in some patients, megakaryocytic cells/platelets. In this model, the level of NF-E2 overexpression determines both the severity of erythrocytosis and the concurrent presence or absence of thrombocytosis.

Blotting, Northern↗

Computational processing and error reduction strategies for standardized quantitative data in biological networks.

High-quality quantitative data generated under standardized conditions is critical for understanding dynamic cellular processes. We report strategies for error reduction, and algorithms for automated data processing and for establishing the widely used techniques of immunoprecipitation and immunoblotting as highly precise methods for the quantification of protein levels and modifications. To determine the stoichiometry of cellular components and to ensure comparability of experiments, relative signals are converted to absolute values. A major source for errors in blotting techniques are inhomogeneities of the gel and the transfer procedure leading to correlated errors. These correlations are prevented by randomized gel loading, which significantly reduces standard deviations. Further error reduction is achieved by using housekeeping proteins as normalizers or by adding purified proteins in immunoprecipitations as calibrators in combination with criteria-based normalization. Additionally, we developed a computational tool for automated normalization, validation and integration of data derived from multiple immunoblots. In this way, large sets of quantitative data for dynamic pathway modeling can be generated, enabling the identification of systems properties and the prediction of targets for efficient intervention.

Algorithms↗