PubMed Health⌕ Search

Biomedical subjects

James J Cai

Publications and source records attributed to James J Cai.

7 recordsLinked to original sources

scPOEM: robust co-embedding of peaks and genes revealing peak-gene regulation.

MOTIVATION: Identifying regulatory elements in various chromosomal regions that influence gene expression is a fundamental challenge in epigenomics, with profound implications for understanding gene regulation and disease mechanisms. The advent of paired single-cell RNA sequencing and single-cell ATAC sequencing has created unprecedented opportunities to address this challenge by enabling simultaneous profiling of gene expression and chromatin accessibility at single-cell resolution. However, the inherent signals between them are weak due to the highly sparse and noisy nature of data. RESULTS: This article proposes single-cell meta-Path based Omics Embedding (scPOEM), a novel embedding method that jointly projects chromatin accessibility peaks and expressed genes into a shared low-dimensional space. By integrating the relationships among peak-peak, peak-gene, and gene-gene interactions, scPOEM assigns closer representations in the embedding space to related peak-gene pairs. Our experiments demonstrate that scPOEM generates stable representations of peaks and genes, outperforms existing methods in recovering biologically meaningful peak-gene regulatory relationships and enables new insights in subgroup and differential analysis of gene regulation. These results highlight its potential to uncover gene regulatory mechanisms and enhance the understanding of transcriptional regulation at single-cell resolution. AVAILABILITY AND IMPLEMENTATION: The source code of scPOEM is available at https://github.com/Houyt23/scPOEM. The datasets can be obtained from the 10× Genomics (https://www.10xgenomics.com/datasets/pbmc-from-a-healthy-donor-granulocytes-removed-through-cell-sorting-10-k-1-standard-1-0-0) and GEO database under access codes GSE194122 and GSE239916.

Gene Expression Regulation↗

Accelerated evolutionary rate may be responsible for the emergence of lineage-specific genes in ascomycota.

The evolutionary origin of "orphan" genes, genes that lack sequence similarity to any known gene, remains a mystery. One suggestion has been that most orphan genes evolve rapidly so that similarity to other genes cannot be traced after a certain evolutionary distance. This can be tested by examining the divergence rates of genes with different degrees of lineage specificity. Here the lineage specificity (LS) of a gene describes the phylogenetic distribution of that gene's orthologues in related species. Highly lineage-specific genes will be distributed in fewer species in a phylogeny. In this study, we have used the complete genomes of seven ascomycotan fungi and two animals to define several levels of LS, such as Eukaryotes-core, Ascomycota-core, Euascomycetes-specific, Hemiascomycetes-specific, Aspergillus-specific, and Saccharomyces-specific. We compare the rates of gene evolution in groups of higher LS to those in groups with lower LS. Molecular evolutionary analyses indicate an increase in nonsynonymous nucleotide substitution rates in genes with higher LS. Several analyses suggest that LS is correlated with the evolutionary rate of the gene. This correlation is stronger than those of a number of other factors that have been proposed as predictors of a gene's evolutionary rate, including the expression level of genes, gene essentiality or dispensability, and the number of protein-protein interactions. The accelerated evolutionary rates of genes with higher LS may reflect the influence of selection and adaptive divergence during the emergence of orphan genes. These analyses suggest that accelerated rates of gene evolution may be responsible for the emergence of apparently orphan genes.

Ascomycota↗

Genomic and experimental evidence for a potential sexual cycle in the pathogenic thermal dimorphic fungus Penicillium marneffei.

All meiotic genes (except HOP1) and genes encoding putative pheromone processing enzymes, pheromone receptors and pheromone response pathways proteins in Aspergillus fumigatus and Aspergillus nidulans and a putative MAT-1 alpha box mating-type gene were present in the Penicillium marneffei genome. A putative MAT-2 high-mobility group mating-type gene was amplified from a MAT-1 alpha box mating-type gene-negative P. marneffei strain. Among 37 P. marneffei patient strains, MAT-1 alpha box and MAT-2 high-mobility group mating-type genes were present in 23 and 14 isolates, respectively. We speculate that P. marneffei can potentially be a heterothallic fungus that does not switch mating type.

Amino Acid Sequence↗

MBEToolbox: a MATLAB toolbox for sequence data analysis in molecular biology and evolution.

BACKGROUND: MATLAB is a high-performance language for technical computing, integrating computation, visualization, and programming in an easy-to-use environment. It has been widely used in many areas, such as mathematics and computation, algorithm development, data acquisition, modeling, simulation, and scientific and engineering graphics. However, few functions are freely available in MATLAB to perform the sequence data analyses specifically required for molecular biology and evolution. RESULTS: We have developed a MATLAB toolbox, called MBEToolbox, aimed at filling this gap by offering efficient implementations of the most needed functions in molecular biology and evolution. It can be used to manipulate aligned sequences, calculate evolutionary distances, estimate synonymous and nonsynonymous substitution rates, and infer phylogenetic trees. Moreover, it provides an extensible, functional framework for users with more specialized requirements to explore and analyze aligned nucleotide or protein sequences from an evolutionary perspective. The full functions in the toolbox are accessible through the command-line for seasoned MATLAB users. A graphical user interface, that may be especially useful for non-specialist end users, is also provided. CONCLUSION: MBEToolbox is a useful tool that can aid in the exploration, interpretation and visualization of data in molecular biology and evolution. The software is publicly available at http://web.hku.hk/~jamescai/mbetoolbox/ and http://bioinformatics.org/project/?group_id=454

Algorithms↗

Characterization and complete genome sequence of a novel coronavirus, coronavirus HKU1, from patients with pneumonia.

Despite extensive laboratory investigations in patients with respiratory tract infections, no microbiological cause can be identified in a significant proportion of patients. In the past 3 years, several novel respiratory viruses, including human metapneumovirus, severe acute respiratory syndrome (SARS) coronavirus (SARS-CoV), and human coronavirus NL63, were discovered. Here we report the discovery of another novel coronavirus, coronavirus HKU1 (CoV-HKU1), from a 71-year-old man with pneumonia who had just returned from Shenzhen, China. Quantitative reverse transcription-PCR showed that the amount of CoV-HKU1 RNA was 8.5 to 9.6 x 10(6) copies per ml in his nasopharyngeal aspirates (NPAs) during the first week of the illness and dropped progressively to undetectable levels in subsequent weeks. He developed increasing serum levels of specific antibodies against the recombinant nucleocapsid protein of CoV-HKU1, with immunoglobulin M (IgM) titers of 1:20, 1:40, and 1:80 and IgG titers of <1:1,000, 1:2,000, and 1:8,000 in the first, second and fourth weeks of the illness, respectively. Isolation of the virus by using various cell lines, mixed neuron-glia culture, and intracerebral inoculation of suckling mice was unsuccessful. The complete genome sequence of CoV-HKU1 is a 29,926-nucleotide, polyadenylated RNA, with G+C content of 32%, the lowest among all known coronaviruses with available genome sequence. Phylogenetic analysis reveals that CoV-HKU1 is a new group 2 coronavirus. Screening of 400 NPAs, negative for SARS-CoV, from patients with respiratory illness during the SARS period identified the presence of CoV-HKU1 RNA in an additional specimen, with a viral load of 1.13 x 10(6) copies per ml, from a 35-year-old woman with pneumonia. Our data support the existence of a novel group 2 coronavirus associated with pneumonia in humans.

Aged↗

The mitochondrial genome of the thermal dimorphic fungus Penicillium marneffei is more closely related to those of molds than yeasts.

We report the complete sequence of the mitochondrial genome of Penicillium marneffei, the first complete mitochondrial DNA sequence of a thermal dimorphic fungus. This 35 kb mitochondrial genome contains the genes encoding ATP synthase subunits 6, 8, and 9 (atp6, atp8, and atp9), cytochrome oxidase subunits I, II, and III (cox1, cox2, and cox3), apocytochrome b (cob), reduced nicotinamide adenine dinucleotide ubiquinone oxireductase subunits (nad1, nad2, nad3, nad4, nad4L, nad5, and nad6), ribosomal protein of the small ribosomal subunit (rps), 28 tRNAs, and small and large ribosomal RNAs. Analysis of gene contents, gene orders, and gene sequences revealed that the mitochondrial genome of P. marneffei is more closely related to those of molds than yeasts.

Base Sequence↗

Exploring the Penicillium marneffei genome.

Penicillium marneffei is a dimorphic fungus that intracellularly infects the reticuloendothelial system of humans and bamboo rats. Endemic in Southeast Asia, it infects 10% of AIDS patients in this region. The absence of a sexual stage and the highly infectious nature of the mould-phase conidia have impaired studies on thermal dimorphic switching and host-microbe interactions. Genomic analysis, therefore, could provide crucial information. Pulsed-field gel electrophoresis of genomic DNA of P. marneffei revealed three or more chromosomes (5.0, 4.0, and 2.2 Mb). Telomeric fingerprinting revealed 6-12 bands, suggesting that there were chromosomes of similar sizes. The genome size of P. marneffei was hence about 17.8-26.2 Mb. G+C content of the genome is 48.8 mol%. Random exploration of the genome of P. marneffei yielded 2303 random sequence tags (RSTs), corresponding to 9% of the genome, with 11.7, 6.3, and 17.4% of the RSTs having sequence similarity to yeast-specific sequences, non-yeast fungus sequences, and both (common sequences), respectively. Analysis of the RSTs revealed genes for information transfer (ribosomal protein genes, tRNA synthetase subunits, translation initiation, and elongation factors), metabolism, and compartmentalization, including several multi-drug-resistance protein genes and homologues of fluconazole-resistance gene. Furthermore, the presence of genes encoding pheromone homologues and ankyrin repeat-containing proteins of other fungi and algae strongly suggests the presence of a sexual stage that presumably exists in the environment.

Base Composition↗