PubMed Health⌕ Search

Biomedical subjects

Qingfa Wu

Publications and source records attributed to Qingfa Wu.

11 recordsLinked to original sources

SAGE detects microRNA precursors.

BACKGROUND: MicroRNAs (miRNAs) have been shown to play important roles in regulating gene expression. Since miRNAs are often evolutionarily conserved and their precursors can be folded into stem-loop hairpins, many miRNAs have been predicted. Yet experimental confirmation is difficult since miRNA expression is often specific to particular tissues and developmental stages. RESULTS: Analysis of 29 human and 230 mouse longSAGE libraries revealed the expression of 22 known and 10 predicted mammalian miRNAs. Most were detected in embryonic tissues. Four SAGE tags detected in human embryonic stem cells specifically match a cluster of four human miRNAs (mir-302a, b, c&d) known to be expressed in embryonic stem cells. LongSAGE data also suggest the existence of a mouse homolog of human and rat mir-493. CONCLUSION: The observation that some orphan longSAGE tags uniquely match miRNA precursors provides information about the expression of some known and predicted miRNAs.

Animals↗

A large quantity of novel human antisense transcripts detected by LongSAGE.

MOTIVATION: Taking advantage of the high sensitivity and specificity of LongSAGE tag for transcript detection and genome mapping, we analyzed the 632 813 unique human LongSAGE tags deposited in public databases to identify novel human antisense transcripts. RESULTS: Our study identified 45 321 tags that match the antisense strand of 9804 known mRNA sequences, 6606 of which contain antisense ESTs and 3198 are mapped only by SAGE tags. Quantitative analysis showed that the detected antisense transcripts are present at levels lower than their counterpart sense transcripts. Experimental results confirmed the presence of antisense transcripts detected by the antisense tags. We also constructed an antisense tag database that can be used to identify the antisense SAGE tags originated from the antisense strand of known mRNA sequences included in the RefSeq database. CONCLUSIONS: Our study highlights the benefits of exploring SAGE data for comprehensive identification of human antisense transcripts and demonstrates the prevalence of antisense transcripts in the human genome.

Base Sequence↗

A novel primate specific gene, CEI, is located in the homeobox gene IRXA2 promoter in Homo sapiens.

The Iroquois (IRX) homeobox gene family consists of six highly conserved transcription factors that are of importance for normal embryonic development. They are organized in two gene clusters in human, one on 5p15.33 and the other one on 16q12.2, respectively, and both the organization and the structure of the genes are highly conserved. An open reading frame coding for an unknown protein is identified in the promoter of IRXA2 on chromosome 5p. This new gene is composed of four exons and it is orientated in a head-to-head manner to IRXA2. Only a short 851 bp segment separates the two translation start codons and the two genes may share a bi-directional promoter. This bi-directional promoter is embedded in a large CpG-island, that also continues into both genes. RT-PCR analysis of the new gene reveals two alternative mRNA transcripts and a third mRNA transcript can be predicted from EST clones. The expression profile of the gene analysed in 9 different human tissues reveals that it is expressed in a coordinated fashion with IRXA2, which has led to the name CEI (Coordinated Expression to IRXA2). The CEI protein lacks homology to any known protein or protein domain in public databases, and a putative amino terminal signal peptide suggests the protein is secreted or ER compartment located. The gene is only found in the human and the chimpanzee genome, but not in the mouse or the rat genome, which suggests that CEI is unique for higher primates. As the identified bi-directional promoter not being a relic of an ancient compact genome, CEI may play an important role in the evolution of higher primates in coordination with the IRX genes.

Amino Acid Sequence↗

Annotating nonspecific SAGE tags with microarray data.

SAGE (serial analysis of gene expression) detects transcripts by extracting short tags from the transcripts. Because of the limited length, many SAGE tags are shared by transcripts from different genes. Relying on sequence information in the general gene expression database has limited power to solve this problem due to the highly heterogeneous nature of the deposited sequences. Considering that the complexity of gene expression at a single tissue level should be much simpler than that in the general expression database, we reasoned that by restricting gene expression to tissue level, the accuracy of gene annotation for the nonspecific SAGE tags should be significantly improved. To test the idea, we developed a tissue-specific SAGE annotation database based on microarray data (). This database contains microarray expression information represented as UniGene clusters for 73 normal human tissues and 18 cancer tissues and cell lines. The nonspecific SAGE tag is first matched to the database by the same tissue type used by both SAGE and microarray analysis; then the multiple UniGene clusters assigned to the nonspecific SAGE tag are searched in the database under the matched tissue type. The UniGene cluster presented solely or at higher expression levels in the database is annotated to represent the specific gene for the nonspecific SAGE tags. The accuracy of gene annotation by this database was largely confirmed by experimental data. Our study shows that microarray data provide a useful source for annotating the nonspecific SAGE tags.

Cell Line↗

The Genomes of Oryza sativa: a history of duplications.

We report improved whole-genome shotgun sequences for the genomes of indica and japonica rice, both with multimegabase contiguity, or almost 1,000-fold improvement over the drafts of 2002. Tested against a nonredundant collection of 19,079 full-length cDNAs, 97.7% of the genes are aligned, without fragmentation, to the mapped super-scaffolds of one or the other genome. We introduce a gene identification procedure for plants that does not rely on similarity to known genes to remove erroneous predictions resulting from transposable elements. Using the available EST data to adjust for residual errors in the predictions, the estimated gene count is at least 38,000-40,000. Only 2%-3% of the genes are unique to any one subspecies, comparable to the amount of sequence that might still be missing. Despite this lack of variation in gene content, there is enormous variation in the intergenic regions. At least a quarter of the two sequences could not be aligned, and where they could be aligned, single nucleotide polymorphism (SNP) rates varied from as little as 3.0 SNP/kb in the coding regions to 27.6 SNP/kb in the transposable elements. A more inclusive new approach for analyzing duplication history is introduced here. It reveals an ancient whole-genome duplication, a recent segmental duplication on Chromosomes 11 and 12, and massive ongoing individual gene duplications. We find 18 distinct pairs of duplicated segments that cover 65.7% of the genome; 17 of these pairs date back to a common time before the divergence of the grasses. More important, ongoing individual gene duplications provide a never-ending source of raw material for gene genesis and are major contributors to the differences between members of the grass family.

Base Sequence↗

Determination of the 'critical region' for cat-like cry of Cri-du-chat syndrome and analysis of candidate genes by quantitative PCR.

Cri-du-chat (CDC, OMIM 123450) is a chromosomal syndrome that results from partial deletions on the short arm of chromosome 5. The clinical features of CDC normally include high-pitched cat-like cry, mental retardation, microcephaly, hypertelorism and epicanthic folds. The cat-like cry is the most prominent clinical characteristic in newborn children and is usually considered as diagnostic for the CDC syndrome. Using a strategy of 'phenotype dissection', the critical region for cat-like cry was mapped to the chromosomal segment 5p15.3-5p15.2 in previous reports. In this study, the distal breakpoints of two interstitial deletions in two clinical distinctive CDC patients are analysed, one with and one without the cat-like cry. Using PCR, the critical region for the cat-like cry is mapped to a short 640 kbp region on chromosome 5p. Genome analysis of this critical region reveals a gene-rich sequence containing five known genes, five putative genes and three spliced EST sequences, altogether 71 predicted exons. Three genes, FLJ25076, a homolog to a ubiquitin-conjugating enzyme UBC-E2, FLJ20303, a nucleolar protein NOP2, which may play a role in the regulation of the cell cycle and MGC5309, a protein with similarity to Nut2, a Drosophila transcriptional coactivator, have been characterized and expression profiles determined by quantitative PCR. These results suggest that one candidate gene, FLJ25076, encodes a ubiquitin-conjugated enzyme E2 type, which is locally expressed in thoracic and scalp tissues. The other two genes are expressed uniformly in all tissues tested, which suggest that they are housekeeping genes.

Animals↗

A draft sequence for the genome of the domesticated silkworm (Bombyx mori).

We report a draft sequence for the genome of the domesticated silkworm (Bombyx mori), covering 90.9% of all known silkworm genes. Our estimated gene count is 18,510, which exceeds the 13,379 genes reported for Drosophila melanogaster. Comparative analyses to fruitfly, mosquito, spider, and butterfly reveal both similarities and differences in gene content.

Algorithms↗

Development of Taqman RT-nested PCR system for clinical SARS-CoV detection.

Severe acute respiratory syndrome (SARS) is an acute newly emerged infectious respiratory illness. The etiologic agent of SARS was named 'SARS-associated coronavirus' (SARS-CoV) that can be detected with reverse transcription-polymerase chain reaction (RT-PCR) assays. In this study, 12 sets of nested primers covering the SARS-CoV genome have been screened and showed sufficient sensitivity to detect SARS-CoV in RNA isolated from virus cultured in Vero 6 cells. To optimize further the reaction condition of those nested primers sets, seven sets of nested primers have been chosen to compare their reverse transcribed efficiency with specific and random primers, which is useful to combine RT with the first round of PCR into a one-step RT-PCR. Based on the sensitivity and simplicity of results, the no. 73 primer set was chosen as the candidate primer set for clinical diagnoses. To specify the amplicon to minimize false positive results, a Taqman RT-nested PCR system of no. 73 nested primer set was developed. Through investigations on a test panel of whole blood obtained from 30 SARS patients and 9 control persons, the specificity and sensitivity of the Taqman RT-nested PCR system was found to be 100 and 83%, respectively, which suggests that the method is a promising one to diagnose SARS in early stages.

Adolescent↗

A genome sequence of novel SARS-CoV isolates: the genotype, GD-Ins29, leads to a hypothesis of viral transmission in South China.

We report a complete genomic sequence of rare isolates (minor genotype) of the SARS-CoV from SARS patients in Guangdong, China, where the first few cases emerged. The most striking discovery from the isolate is an extra 29-nucleotide sequence located at the nucleotide positions between 27,863 and 27,864 (referred to the complete sequence of BJ01) within an overlapped region composed of BGI-PUP5 (BGI-postulated uncharacterized protein 5) and BGI-PUP6 upstream of the N (nucleocapsid) protein. The discovery of this minor genotype, GD-Ins29, suggests a significant genetic event and differentiates it from the previously reported genotype, the dominant form among all sequenced SARS-CoV isolates. A 17-nt segment of this extra sequence is identical to a segment of the same size in two human mRNA sequences that may interfere with viral replication and transcription in the cytosol of the infected cells. It provides a new avenue for the exploration of the virus-host interaction in viral evolution, host pathogenesis, and vaccine development.

Base Sequence↗

The E protein is a multifunctional membrane protein of SARS-CoV.

The E (envelope) protein is the smallest structural protein in all coronaviruses and is the only viral structural protein in which no variation has been detected. We conducted genome sequencing and phylogenetic analyses of SARS-CoV. Based on genome sequencing, we predicted the E protein is a transmembrane (TM) protein characterized by a TM region with strong hydrophobicity and alpha-helix conformation. We identified a segment (NH2-_L-Cys-A-Y-Cys-Cys-N_-COOH) in the carboxyl-terminal region of the E protein that appears to form three disulfide bonds with another segment of corresponding cysteines in the carboxyl-terminus of the S (spike) protein. These bonds point to a possible structural association between the E and S proteins. Our phylogenetic analyses of the E protein sequences in all published coronaviruses place SARS-CoV in an independent group in Coronaviridae and suggest a non-human animal origin.

Amino Acid Sequence↗

Complete genome sequences of the SARS-CoV: the BJ Group (Isolates BJ01-BJ04).

Beijing has been one of the epicenters attacked most severely by the SARS-CoV (severe acute respiratory syndrome-associated coronavirus) since the first patient was diagnosed in one of the city's hospitals. We now report complete genome sequences of the BJ Group, including four isolates (Isolates BJ01, BJ02, BJ03, and BJ04) of the SARS-CoV. It is remarkable that all members of the BJ Group share a common haplotype, consisting of seven loci that differentiate the group from other isolates published to date. Among 42 substitutions uniquely identified from the BJ group, 32 are non-synonymous changes at the amino acid level. Rooted phylogenetic trees, proposed on the basis of haplotypes and other sequence variations of SARS-CoV isolates from Canada, USA, Singapore, and China, gave rise to different paradigms but positioned the BJ Group, together with the newly discovered GD01 (GD-Ins29) in the same clade, followed by the H-U Group (from Hong Kong to USA) and the H-T Group (from Hong Kong to Toronto), leaving the SP Group (Singapore) more distant. This result appears to suggest a possible transmission path from Guangdong to Beijing/Hong Kong, then to other countries and regions.

Genome, Viral↗