PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

FusionTarget: Computational framework for drug repurposing against modeled fusion protein structures from genomic breakpoints.

Many fusion genes have been recognized as biomarkers and therapeutic targets. However, the lack of knowledge on protein structures and targeting approaches made it challenging to develop effective targeting therapeutics. To fill this, we developed a computational pipeline, FusionTarget, which annotates the genomic DNA breakage to RNA and protein sequences, predicts the 3D structures of fusion proteins, and performs comparative virtual screening, comparative molecular dynamics simulation, and quantitative analyses to identify the fusion protein-selective small molecules by selecting drugs with consistent high-fold binding affinity between fusion and wild-type proteins in multiple isoforms. We applied our pipeline to EWSR1::FLI1 in Ewing sarcoma and KMT2A::AFF1 in infant acute lymphoblastic leukemia. Further cell assay experiments confirmed that cells expressing individual fusion genes were more sensitive to the suggested drugs, and the key downstream genes were affected by our drugs. FusionTarget provides a unique foundation for developing therapeutics targeting fusion proteins.

applied computing in medical science↗

The bioinformatics approach to identifying pathogenic variants for colorectal cancer (CRC).

Colorectal cancer (CRC) is the third most prevalent cancer globally, accounting for 9.6% of newly diagnosed cases and 9.3% of cancer-related deaths. It develops from the uncontrolled proliferation of glandular cells in the colon and rectum and is categorized into three primary types: sporadic, hereditary, and colitis-associated. While genetic susceptibility is a key factor in CRC pathogenesis, identifying high-impact pathogenic variants remains a significant challenge. This study integrates bioinformatics and population genetics approaches to identify CRC-associated single-nucleotide polymorphisms (SNPs) with potential clinical significance. CRC-associated SNPs were extracted from the Genome-Wide Association Studies (GWAS) Catalog, functionally annotated via HaploReg, and validated via Ensembl. In addition, expression quantitative trait locus (eQTL) data from the GTEx database were used to assess the effects of these variants on gene expression across human tissues. Our analysis identified three high-priority SNPs (rs9379084, rs3184504, and rs11557154) associated with the RREB1, ATXN2, SH2B3, and DCAF12 genes, which exhibited marked allele frequency differences among populations. These findings suggest potential biomarkers for CRC risk assessment and highlight the importance of genetic screening across diverse populations.

Bioinformatics↗

Genome-wide protein interaction maps using two-hybrid systems.

Automated sequence technology has rendered functional biology amenable to genomic scale analysis. Among genome-wide exploratory approaches, the two-hybrid system in yeast (Y2H) has outranked other techniques because it is the system of choice to detect protein-protein interactions. Deciphering the cascade of binding events in a whole cell helps define signal transduction and metabolic pathways or enzymatic complexes. The function of proteins is eventually attributed through whole cell protein interaction maps where totally unknown proteins are partnered with fully annotated proteins belonging to the same functional category. Since its first description in the late 1980's, several versions of the Y2H have been developed in order to overcome the major limitations of the system, namely false positives and false negatives. Optimized versions have been recently applied at multi-molecular and genomic scale. These genome-wide surveys can be methodologically divided into two types of approaches: one either tests combinations of predefined polypeptides (the so-called matrix approach) using various short-cuts to speed up the process, or one screens with a given polypeptide (bait) for potential partners (preys) present in complex libraries of genomic or complementary DNA (library screening). In the former strategy, one tests what one knows, for example pair-wise interactions between full-length open reading frames from recently sequenced and annotated genomes. Although based on a one-by-one scheme, this method is reported to be amenable to large-scale genomics thanks to multicloning strategies and to the use of small robotics workstations. In the latter, highly complex cDNA or genomic libraries of protein domains can be screened to saturation with high-throughput screening systems allowing the discovery of yet unidentified proteins. Both approaches have strengths and drawbacks that will be discussed here. None yields a full proteome-wide screening since certain proteins (e.g. some transcription factors) are not usable in Y2H. Novel two-hybrid assays have been recently described in bacteria. Applications of these time- and cost-effective assays to genomic screening will be discussed and compared to the Y2H technology.

Animals↗

The competence gene, comF, from Synechocystis sp. strain PCC 6803 is involved in natural transformation, phototactic motility and piliation.

The gene slr0388 was previously annotated to encode a hypothetical protein in Synechocystis sp. strain PCC 6803. When a positively phototactic strain of this cyanobacterium was insertionally inactivated at slr0388, the mutants were not transformable, and appeared to aggregate as a result of increased bundling of type IV pili. Also, these mutants were rendered non-phototactic compared to the wild-type. Quantitative real-time PCR revealed a 3.5-fold increase in pilA1 transcript levels in the mutant over wild-type cells, while there were no changes in the level of pilT1 and comA transcripts. Supernatant from mutant liquid culture contained more PilA1 protein, confirmed by mass spectrometric analysis, compared to the wild-type cells, which corresponded to the increase in pilA1 transcripts. The increase in PilA1 subunits may contribute to the bundling morphology of pili that was observed, which in turn may act to retard DNA uptake by hindering the retraction of pili. This gene is therefore proposed to be designated comF, as it possesses a phosphoribosyltransferase domain, a distinguishing feature of other ComF proteins of naturally transformable heterotrophic bacteria. This report is the second of a competence-related gene from Synechocystis sp. strain PCC 6803, the product of which does not show homology to other well-studied type IV pili proteins.

Amino Acid Sequence↗

Gene array analysis of bone morphogenetic protein type I receptor-induced osteoblast differentiation.

UNLABELLED: The genomic response to BMP was investigated by ectopic expression of activated BMP type I receptors in C2C12 myoblast using cDNA microarrays. Novel BMP receptor target genes with possible roles in inhibition of myoblast differentiation and stimulation of osteoblast differentiation were identified. INTRODUCTION: Bone morphogenetic proteins (BMPs) have an important role in controlling mesenchymal cell fate and mediate these effects by regulating gene expression. BMPs signal through three distinct specific BMP type I receptors (also termed activin receptor-like kinases) and their downstream nuclear effectors, termed Smads. The critical target genes by which activated BMP receptors mediate change cell fate are poorly characterized. MATERIALS AND METHODS: We performed transcriptional profiling of C2C12 myoblasts differentiation into osteoblast-like cells by ectopic expression of three distinct constitutively active (ca)BMP type I receptors using adenoviral gene transfer. Cells were harvested 48 h after infection, which allowed detection of both early and late response genes. Expression analysis was performed using the mouse GEM1 microarray, which is comprised of approximately 8700 unique sequences. Hybridizations were performed in duplicate with a reverse fluor labeling. Genes were considered to be significantly regulated if the p value for differential expression was less than 0.01 and inverted expression ratios per duplicate successful reciprocal hybridizations differed by less than 25%. RESULTS AND CONCLUSIONS: Each of the three caBMP type I receptors stimulated equal levels of R-Smad phosphorylation and alkaline phosphatase activity, an early marker for osteoblast differentiation. Interestingly, all three type I receptors induced identical transcriptional profiles; 97 genes were significantly upregulated and 103 genes were downregulated. Many extracellular matrix genes were upregulated, muscle-related genes downregulated, and transcription factors/signaling components modulated. In addition to 41 expressed sequence tags without known function and a number of known BMP target genes, including PPAR-gamma and fibromodulin, a large number of novel BMP target genes with an annotated function were identified, including transcription factors HesR1, ITF-2, and ICSBP, apoptosis mediators DRP-1 death kinase and ZIP kinase, IkappaB alpha, Edg-2, ZO-1, and E3 ligase Dactylin. These target genes, some of them unexpected, offer new insights into how BMPs elicit biological effects, in particular into the mechanism of inhibition of myoblast differentiation and stimulation of osteoblast differentiation.

Animals↗

Identification of psl, a locus encoding a potential exopolysaccharide that is essential for Pseudomonas aeruginosa PAO1 biofilm formation.

Bacteria inhabiting biofilms usually produce one or more polysaccharides that provide a hydrated scaffolding to stabilize and reinforce the structure of the biofilm, mediate cell-cell and cell-surface interactions, and provide protection from biocides and antimicrobial agents. Historically, alginate has been considered the major exopolysaccharide of the Pseudomonas aeruginosa biofilm matrix, with minimal regard to the different functions polysaccharides execute. Recent chemical and genetic studies have demonstrated that alginate is not involved in the initiation of biofilm formation in P. aeruginosa strains PAO1 and PA14. We hypothesized that there is at least one other polysaccharide gene cluster involved in biofilm development. Two separate clusters of genes with homology to exopolysaccharide biosynthetic functions were identified from the annotated PAO1 genome. Reverse genetics was employed to generate mutations in genes from these clusters. We discovered that one group of genes, designated psl, are important for biofilm initiation. A PAO1 strain with a disruption of the first two genes of the psl cluster (PA2231 and PA2232) was severely compromised in biofilm initiation, as confirmed by static microtiter and continuous culture flow cell and tubing biofilm assays. This impaired biofilm phenotype could be complemented with the wild-type psl sequences and was not due to defects in motility or lipopolysaccharide biosynthesis. These results implicate an as yet unknown exopolysaccharide as being required for the formation of the biofilm matrix. Understanding psl-encoded exopolysaccharide expression and protection in biofilms will provide insight into the pathogenesis of P. aeruginosa in cystic fibrosis and other infections involving biofilms.

Bacterial Proteins↗

Phylogeny of Na+/Ca2+ exchanger (NCX) genes from genomic data identifies new gene duplications and a new family member in fish species.

The Na+/Ca2+ exchanger (NCX) is a member of the cation/Ca2+ antiporter (CaCA) family and plays a key role in maintaining cellular Ca2+ homeostasis in a variety of cell types. NCX is present in a diverse group of organisms and exhibits high overall identity across species. To date, three separate genes, i.e., NCX1, NCX2, and NCX3, have been identified in mammals. However, phylogenetic analysis of the exchanger has been hindered by the lack of nonmammalian NCX sequences. In this study, we expand and diversify the list of NCX sequences by identifying NCX homologs from whole-genome sequences accessible through the Ensembl Genome Browser. We identified and annotated 13 new NCX sequences, including 4 from zebrafish, 4 from Japanese pufferfish, 2 from chicken, and 1 each from honeybee, mosquito, and chimpanzee. Examination of NCX gene structure, together with construction of phylogenetic trees, provided novel insights into the molecular evolution of NCX and allowed us to more accurately annotate NCX gene names. For the first time, we report the existence of NCX2 and NCX3 in organisms other than mammals, yielding the hypothesis that two serial NCX gene duplications occurred around the time vertebrates and invertebrates diverged. In addition, we have found a putative new NCX protein, named NCX4, that is related to NCX1 but has been observed only in fish species genomes. These findings present a stronger foundation for our understanding of the molecular evolution of the NCX gene family and provide a framework for further NCX phylogenetic and molecular studies.

Amino Acid Sequence↗

Transcription profiling of renal cell carcinoma.

AIMS: Our aim was to prepare a comprehensive catalogue of the changes in gene expression accompanying the development and progression of renal cell carcinoma, and to correlate these with histo-pathological, cytogenetic and clinical findings. METHODS: mRNA samples from paired neoplastic and non-cancerous human kidney tissue were labeled and hybridized in duplicate against high-density cDNA arrays. Two array technologies were used: 31,500-element transcriptome-wide nylon arrays for hybridization with 37 radioactively labelled sample pairs, and 4200-element kidney- and cancer-specific glass microarrays for hybridization with 19 fluorescently labelled sample pairs. RESULTS: We identified more than 1700 cDNA clones that show differential transcription levels in kidney tumor tissue compared to normal kidney tissue. The functional classification of 389 annotated genes provided views of the changes in the activities of specific biological processes in renal cancer. Among the biological processes with a large proportion of up-regulated genes we found cell adhesion, signal transduction, and nucleotide metabolism. Down-regulated processes included small molecule transport, ion homeostasis, and oxygen and radical metabolism. Furthermore, we explored the feasibility of molecular diagnosis for renal cell tumors using cDNA microarrays on glass slides, investigating the association of transcription levels with tumor type, progression, and a putative prognostic variable. The experimental data is available from the GEO gene expression database (http://www.ncbi.nlm.nih.gov/geo; accession no. GSE3), and a comprehensive presentation of the results is available in the web supplement (http://www.dkfz-heidelberg.de/abt0840/whuber/rcc). CONCLUSION: Transcription profiling using high-density cDNA arrays is a powerful method with the potential to improve cancer diagnosis and prognosis. The identification and classification of differentially transcribed genes, as described in our study, is the beginning of a more complete understanding of kidney cancer.

Carcinoma, Renal Cell↗

Gene expression profiling and analysis of signaling pathways involved in priming and differentiation of human neural stem cells.

Human neural stem cells have the ability to differentiate into all three major cell types in the CNS including neurons, astrocytes and oligodendrocytes. The multipotency of human neural stem cells shed a light on the possibility of using stem cells as a therapeutic tool for various neurological disorders including neurodegenerative diseases and neurotrauma that involve a loss of functional neurons. We have discovered previously a priming procedure to direct primarily cultured human neural stem cells to differentiate into almost pure neurons when grafted into adult CNS. However, the molecular mechanism underlying this phenomenon is still unknown. To unravel transcriptional changes of human neural stem cells upon priming, cDNA microarray was used to study temporal changes in human neural stem cell gene expression profile during priming and differentiation. As a result, transcriptional levels of 520 annotated genes were detected changed in at least at two time points during the priming process. In addition, transcription levels of more than 3000 hypothetical protein encoding genes and EST genes were modulated during the priming and differentiation processes of human neural stem cells. We further analyzed the named genes and grouped them into 14 functional categories. Of particular interest, key cell signal transduction pathways, including the G-protein-mediated signaling pathways (heterotrimeric and small monomeric GTPase pathways), the Wnt signaling pathway and the TGF-beta pathway, are modulated by the neural stem cell priming, suggesting important roles of these key signaling pathways in priming and differentiation of human neural stem cells.

Bone Morphogenetic Proteins↗

Low-frequency Fourier spectrum for predicting membrane protein types.

Cell membranes are vitally important to living cells. Although the infrastructure of biological membrane is provided by the lipid bilayer, membrane proteins perform most of the specific functions. Knowledge of membrane protein types often provides crucial hints toward determining the function of an uncharacterized membrane protein. With the avalanche of new protein sequences generated in the post-genomic era, it is highly demanded to develop a high throughput tool in identifying the type of newly found membrane proteins according to their primary sequences, so as to timely annotate them for reference usage in both basic research and drug discovery. To realize this, the key is to establish a powerful identifier that can catch their characteristic sequence patterns for different membrane protein types. However, it is not easy because they are buried in a pile of long and complicated sequences. In this paper, based on the concept of the pseudo-amino acid composition [K.C. Chou, PROTEINS: Struct., Funct., Genet. 43 (2001) 246-255], the low-frequency Fourier spectrum analysis is introduced. The merits by doing so are that the sequence pattern information can be more effectively incorporated into a set of discrete components, and that all the existing prediction algorithms can be straightforwardly used on such a formulation for protein samples. High success rates were observed by the re-substitution test, jackknife test, and independent dataset test, indicating that the low-frequency Fourier spectrum approach may become a very useful tool for membrane protein type prediction. The novel approach also holds a high potential for predicting many other attributes of proteins.

Fourier Analysis↗

Three-dimensional database of subcortical electrophysiology for image-guided stereotactic functional neurosurgery.

We present a method of constructing a database of intraoperatively observed human subcortical electrophysiology. In this approach, patient electrophysiological data are standardized using a multiparameter coding system, annotated to their respective magnetic resonance images (MRIs), and nonlinearly registered to a high-resolution MRI reference brain. Once registered, we are able to demonstrate clustering of like interpatient physiologic responses within the thalamus, globus pallidus, subthalamic nucleus, and adjacent structures. These data may in turn be registered to a three-dimensional patient MRI within our image-guided visualization program enabling prior to surgery the delineation of surgical targets, anatomy with high probability of containing specific cell types, and functional borders. The functional data were obtained from 88 patients (106 procedures) via microelectrode recording and electrical stimulation performed during stereotactic neurosurgery at the London Health Sciences Centre. Advantages of this method include the use of nonlinear registration to accommodate for interpatient anatomical variability and the avoidance of digitized versions of printed atlases of anatomy as a common database coordinate system. The resulting database is expandable, easily searched using a graphical user interface, and provides a visual representation of functional organization within the deep brain.

Brain↗

Assembly and annotation of human chromosome 2q33 sequence containing the CD28, CTLA4, and ICOS gene cluster: analysis by computational, comparative, and microarray approaches.

Human chromosome 2q33 is an immunologically important region based on the linkage of numerous autoimmune diseases to the CTLA4 locus. Here, we sequenced and assembled 2q33 bacterial artificial chromosome (BAC) clones, resulting in 381,403 bp of contiguous sequence containing genes encoding a NADH: ubiquinone oxidoreductase, the costimulatory receptors CD28, CTLA4, and ICOS, and a HERV-H type endogenous retrovirus located 366 bp downstream of ICOS in the reverse orientation. Genomic microarray expression analysis using differentially activated T-cell RNA against a subcloned CTLA4/ICOS BAC library revealed upregulation of CTLA4 and ICOS sequences, plus antisense ICOS transcripts generated by the HERV-H, suggesting a potential mechanism for ICOS regulation. We identified four nonlinked, polymorphic, simple repetitive sequence elements in this region, which may be used to delineate genetic effects of ICOS and CTLA4 in disease populations. Comparative genomic analysis of mouse genomic Icos sequences revealed 60% sequence identity in the 5' UTR and regions between exon 2 and the 3' UTR, suggesting the importance of ICOS gene function.

Abatacept↗

Multi-season analysis reveals hundreds of drought-responsive genes in sorghum.

Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3 years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.

Sorghum↗

Transcriptome profiling of adult zebrafish at the late stage of chronic tuberculosis due to Mycobacterium marinum infection.

The Mycobacterium marinum-zebrafish infection model was used in this study for analysis of a host transcriptome response to mycobacterium infection at the organismal level. RNA isolated from adult zebrafish that showed typical signs of fish tuberculosis due to a chronic progressive infection with M. marinum was compared with RNA from healthy fish in microarray analyses. Spotted oligonucleotide sets (designed by Sigma-Compugen and MWG) and Affymetrix GeneChips were used, in total comprising 45,465 zebrafish transcript annotations. Based on a detailed comparative analysis and quantitative reverse transcriptase-PCR analysis, we present a validated reference set of 159 genes whose regulation is strongly affected by mycobacterial infection in the three types of microarrays analyzed. Furthermore, we analyzed the separate datasets of the microarrays with special emphasis on the expression profiles of immune-related genes. Upregulated genes include many known components of the inflammatory response and several genes that have previously been implicated in the response to mycobacterial infections in cell cultures of other organisms. Different marker genes of the myeloid lineage that have been characterized in zebrafish also showed increased expression. Furthermore, the zebrafish homologs of many signal transduction genes with relationship to the immune response were induced by M. marinum infection. Future functional analysis of these genes may contribute to understanding the mechanisms of mycobacterial pathogenesis. Since a large group of genes linked to immune responses did not show altered expression in the infected animals, these results suggest specific responses in mycobacterium-induced disease.

Animals↗

Conventional and Shared Genetic Association Analysis Between Diabetes Mellitus and Sensorineural Hearing Loss.

PURPOSE: This study aims to investigate the epidemiological and genetic associations between diabetes mellitus (DM) and sensorineural hearing loss (SNHL) across different subtypes. METHODS: We analyzed 502,490 participants from the UK Biobank using multivariate logistic regression to examine the association between DM and SNHL, considering gender, age, and HbA1c levels. Genetic correlations and causality were examined by linkage disequilibrium score regression and bidirectional Mendelian randomization. Cross-trait meta-analyses identified shared loci between DM and SNHL, followed by gene annotation, functional analysis, and drug candidate exploration for the shared traits. RESULTS: Observational analysis revealed significant associations between DM and SNHL, consistent in subgroups based on age, sex, and certain HbA1c levels. A positive genetic correlation was found between type 2 diabetes mellitus (T2D) and SNHL (Rg = 0.0982, p = 0.0095) between T2D and SNHL, and four loci were identified, with ARHGEF28 and TCF7L2 prioritized as credible pleiotropic genes. Enrichment was indicated in glucose metabolism and organogenesis, with shared heritability in metabolic tissues and outer hair cells. Metformin was identified as potential drug candidates for the T2D-SNHL comorbidity. CONCLUSION: These findings progress our understanding of the epidemiological association, shared genetic basis, and potential therapeutic targets between T2D and SNHL, which might contribute to the management of their comorbidity.

Humans↗

GPX-Macrophage Expression Atlas: a database for expression profiles of macrophages challenged with a variety of pro-inflammatory, anti-inflammatory, benign and pathogen insults.

BACKGROUND: Macrophages play an integral role in the host immune system, bridging innate and adaptive immunity. As such, they are finely attuned to extracellular and intracellular stimuli and respond by rapidly initiating multiple signalling cascades with diverse effector functions. The macrophage cell is therefore an experimentally and clinically amenable biological system for the mapping of biological pathways. The goal of the macrophage expression atlas is to systematically investigate the pathway biology and interaction network of macrophages challenged with a variety of insults, in particular via infection and activation with key inflammatory mediators. As an important first step towards this we present a single searchable database resource containing high-throughput macrophage gene expression studies. DESCRIPTION: The GPX Macrophage Expression Atlas (GPX-MEA) is an online resource for gene expression based studies of a range of macrophage cell types following treatment with pathogens and immune modulators. GPX-MEA follows the MIAME standard and includes an objective quality score with each experiment. It places special emphasis on rigorously capturing the experimental design and enables the searching of expression data from different microarray experiments. Studies may be queried on the basis of experimental parameters, sample information and quality assessment score. The ability to compare the expression values of individual genes across multiple experiments is provided. In addition, the database offers access to experimental annotation and analysis files and includes experiments and raw data previously unavailable to the research community. CONCLUSION: GPX-MEA is the first example of a quality scored gene expression database focussed on a macrophage cellular system that allows efficient identification of transcriptional patterns. The resource will provide novel insights into the phenotypic response of macrophages to a variety of benign, inflammatory, and pathogen insults. GPX-MEA is available through the GPX website at http://www.gti.ed.ac.uk/GPX.

Animals↗

SpliceHarmonization: an integrated method for identifying RNA splicing events in therapeutics for splicing modulation.

MOTIVATION: Splicing, a critical co-transcriptional process in eukaryotes, enhances transcriptome diversity by generating isoforms specific to cell types, tissues, or developmental stages. Recent advancements in splicing modulators have opened new avenues for targeting previously undruggable genes by inducing significant perturbations in splicing events. These developments underscore the need for comprehensive methods to accurately identify and compare splicing events. While several tools have been developed to detect local splice variants, inconsistencies across methods remain a significant challenge. To address this, we present SpliceHarmonization, an integrated approach that combines the strengths of rMATS, LeafCutter, and MAJIQ, enabling robust and reliable splicing analysis with event type annotations. RESULTS: In a comprehensive evaluation using diverse simulated datasets, SpliceHarmonization streamlined and standardized the outputs from three detection methods into a unified format, thereby improving splicing detection with event type annotation and outperforming individual methods. By integrating the outputs from rMATS, LeafCutter, and MAJIQ, our approach not only enhanced identification of a wide range of splicing events but also effectively mitigated method-specific discrepancies. This integration led to an accuracy exceeding 0.8 and a recall of up to 0.5, with an observed increase in AUC of up to 10%. Furthermore, SpliceHarmonization demonstrated high sensitivity in detecting low-abundance and complex splicing events, providing annotations including genomic coordinates and event type. AVAILABILITY AND IMPLEMENTATION: SpliceHarmonization is available at https://github.com/interactivereport/SpliceHarmonization.

RNA Splicing↗

Progress in rickettsial genome analysis from pioneering of Rickettsia prowazekii to the recent Rickettsia typhi.

Three rickettsial genomes have been sequenced and annotated. Rickettsia prowazekii and R. typhi have similar gene order and content. The few differences between R. prowazekii and R. typhi include a 12-kb insertion in R. prowazekii, a large inversion close to the origin of replication in R. typhi, and loss of the complete cytochrome c oxidase system by R. typhi. R. prowazekii, R. typhi, and R. conorii have 13, 24, and 560 unique genes, respectively, and share 775 genes, most likely their essential genes. The small genomes contain many pseudogenes and much noncoding DNA, reflecting the process of genome decay. R. typhi contains the largest number of pseudogenes (41), and R. conorii the fewest, in accordance with its larger number of genes and smaller proportion of noncoding DNA. Conversely, typhus rickettsiae contain fewer repetitive sequences. These genomes portray the key themes of rickettsial intracellular survival: lack of enzymes for sugar metabolism, lipid biosynthesis, nucleotide synthesis, and amino acid metabolism, suggesting that rickettsiae depend on the host for nutrition and building blocks; enzymes for the complete TCA cycle and several copies of ATP/ADP translocase genes, suggesting independent synthesis of ATP and acquisition of host ATP; and type IV secretion system. All rickettsiae share two outer membrane proteins (OmpB and Sca 4) and LPS biosynthesis machinery. RickA, unique to spotted fever rickettsiae, plays a role in induction of actin polymerization in R. conorii, but not in R. prowazekii or R. typhi. The genome of R. typhi contains four potentially membranolytic genes (tlyA, tlyC, pldA, and pat-1) and five autotransporter genes, sca 1, sca 2, sca 3, ompA, and ompB. The presence of six 50-amino acid repeat units in Sca 2 suggests function as an adhesin. The high laboratory passage of the sequenced strains raises the issue of the occurrence of laboratory mutations in genes not required for growth in cell culture or eggs. Resequencing revealed that eight annotated pseudogenes of E strain are actually intact genes. Comparative genomics of virulent and avirulent strains of rickettsial species may reveal their virulence factors.

Genome, Bacterial↗