PubMed Health⌕ Search

Biomedical subjects

Andrew J G Simpson

Publications and source records attributed to Andrew J G Simpson.

15 recordsLinked to original sources

Identification of 9 novel transcripts and two RGSL genes within the hereditary prostate cancer region (HPC1) at 1q25.

We applied a systematic bioinformatics approach, followed by careful manual inspection and experimental validation to identify additional expressed sequences located at the Hereditary Prostate Cancer Region (HPC1) between D1S2818 and D1S1642 on chromosome 1q25. All transcripts already described for the 1q25 region were identified and we were able to define 11 additional expressed sequences within this region (three full-length cDNA clone sequences and eight ESTs), increasing the total number of gene count in this region by 38%. Five out of the 11 expressed sequences identified were shown to be expressed in prostate tissue and thus represent novel disease gene candidates for the HPC1 region. Here, we report a detailed characterization of these five novel disease gene candidates, their expression pattern in various tissues, their genomic organization and functional annotation. Two candidates (RGSL1 and RGSL2) correspond to novel members of the RGS family, which is involved in the regulation of G-protein signaling. RGSL1 and RGLS2 expression was detected by real-time polymerase chain reaction in normal prostate tissue, but could not be detected in prostate tumor cell lines, suggesting they might have a role in prostate cancer.

Chromosome Mapping↗

Comprehensive sampling of gene expression in human cell lines with massively parallel signature sequencing.

Whereas information is rapidly accumulating about the structure and position of genes encoded in the human genome, less is known about the complexity and relative abundance of their expression in individual human cells and tissues. Here, we describe the characteristics of the transcriptomes of two cultured cell lines, HB4a (normal breast epithelium) and HCT-116 (colon adenocarcinoma), using massively parallel signature sequencing (MPSS). We generated in excess of 10(7) short signature sequences per cell line, thus providing a comprehensive snapshot of gene expression, within the technical limitations of the method. The number of genes expressed at one copy per cell or more in either of the lines was estimated to be between 10,000 and 15,000. The vast majority of the transcripts found in these cells can be mapped to known genes and their polyadenylation variants. Among the genes that could be identified from their signature sequences, approximately 8,500 were expressed by both cell lines, whereas 6,000 showed cellular specificity. Taking into account sequence tags that map uniquely to the genome but not to known transcripts, overall the data are consistent with an upper limit of 17,000 for the total number of genes expressed at more than one copy per cell in one or both of the two cell lines examined.

Adenocarcinoma↗

The common mitochondrial DNA deletion deltamtDNA(4977): shedding new light to the concept of a tumor suppressor mutation.

We propose that the age-related accumulation of deltamtDNA(4977) mutations may serve a protective function against tumor-promoting effects of other somatic mutations. The evidence discussed here is consistent with the concept that deltamtDNA(4977) plays a tumor-suppressor role, thus shedding new light to the concept of a tumor suppressor mutation. This concept may help understand how a tumor-promoting mutation may be able to cause malignant transformation in cells lacking a tumor-suppressor mutation, while the same tumor-promoting mutation can be present in cells that carry a tumor-suppressor mutation, without causing cancer.

DNA, Mitochondrial↗

Sequence-based cancer genomics: progress, lessons and opportunities.

Technologies that provide a genome-wide view offer an unprecedented opportunity to scrutinize the molecular biology of the cancer cell. The information that is derived from these technologies is well suited to the development of public databases of alterations in the cancer genome and its expression. Here, we describe the synergistic efforts of research programmes in Brazil, the United Kingdom and the United States towards building integrated databases that are widely accessible to the research community, to enable basic and applied applications in cancer research.

Brazil↗

Low-stringency single specific primer PCR for identification of Leptospira.

Thirty-five Leptospira serovars from the species Leptospira interrogans, Leptospira borgpetersenii, Leptospira santarosai, Leptospira kirschneri, Leptospira weilii, Leptospira biflexa and Leptospira meyeri were characterized by the low-stringency single specific primer PCR (LSSP-PCR) technique. LSSP-PCR analysis was performed to detect DNA polymorphisms in a 285 bp DNA fragment amplified from genomic DNA with G1 and G2 selected primers. Similar LSSP-PCR profiles were obtained for serovars from the same genomic species, while serovars from non-related species produced distinct multiband patterns. Based on the data from sequence analysis, all genomic fragments amplified with G1 and G2 primers from distinct serovars of Leptospira were 285 bp in length, with nucleotide variation observed most frequently among different genomic species. The simplicity and accuracy of the LSSP-PCR technique were found to be suitable for identification of Leptospira species.

Base Sequence↗

High-throughput SELEX SAGE method for quantitative modeling of transcription-factor binding sites.

The ability to determine the location and relative strength of all transcription-factor binding sites in a genome is important both for a comprehensive understanding of gene regulation and for effective promoter engineering in biotechnological applications. Here we present a bioinformatically driven experimental method to accurately define the DNA-binding sequence specificity of transcription factors. A generalized profile was used as a predictive quantitative model for binding sites, and its parameters were estimated from in vitro-selected ligands using standard hidden Markov model training algorithms. Computer simulations showed that several thousand low- to medium-affinity sequences are required to generate a profile of desired accuracy. To produce data on this scale, we applied high-throughput genomics methods to the biochemical problem addressed here. A method combining systematic evolution of ligands by exponential enrichment (SELEX) and serial analysis of gene expression (SAGE) protocols was coupled to an automated quality-controlled sequence extraction procedure based on Phred quality scores. This allowed the sequencing of a database of more than 10,000 potential DNA ligands for the CTF/NFI transcription factor. The resulting binding-site model defines the sequence specificity of this protein with a high degree of accuracy not achieved earlier and thereby makes it possible to identify previously unknown regulatory sequences in genomic DNA. A covariance analysis of the selected sites revealed non-independent base preferences at different nucleotide positions, providing insight into the binding mechanism.

Base Sequence↗

In silico comparison of the transcriptome derived from purified normal breast cells and breast tumor cell lines reveals candidate upregulated genes in breast tumor cells.

Genes that are differentially expressed in tumor tissues are potential diagnostic markers and drug targets. The DNA sequence information available in the public databases can be used to identify transcripts differentially expressed in cancer. We report here the combined use of the ORESTES sequences generated in the FAPESP/LICR Human Cancer Genome Project and information available in the UniGene and SAGE databases to characterize the transcriptome of normal and breast tumor cells. We have identified 154 genes as candidates for overexpression in breast tumor cells. Among these, 28 genes have been shown by others to be overexpressed in breast or other tumors. Using RT-PCR, we tested 11 candidate genes and found that 9 were indeed overexpressed in breast tumor cells.

Breast↗

Nineteen additional unpredicted transcripts from human chromosome 21.

The identification of all human chromosome 21 (HC21) genes is a necessary step in understanding the molecular pathogenesis of trisomy 21 (Down syndrome). The first analysis of the sequence of 21q included 127 previously characterized genes and predicted an additional 98 novel anonymous genes. Recently we evaluated the quality of this annotation by characterizing a set of HC21 open reading frames (C21orfs) identified by mapping spliced expressed sequence tags (ESTs) and predicted genes (PREDs), identified only in silico. This study underscored the limitations of in silico-only gene prediction, as many PREDs were incorrectly predicted. To refine the HC21 annotation, we have developed a reliable algorithm to extract and stringently map sequences that contain bona fide 3' transcript ends to the genome. We then created a specific 21q graphical display allowing an integrated view of the data that incorporates new ESTs as well as features such as CpG islands, repeats, and gene predictions. Using these tools we identified 27 new putative genes. To validate these, we sequenced previously cloned cDNAs and carried out RT-PCR, 5'- and 3'-RACE procedures, and comparative mapping. These approaches substantiated 19 new transcripts, thus increasing the HC21 gene count by 9.5%. These transcripts were likely not previously identified because they are small and encode small proteins. We also identified four transcriptional units that are spliced but contain no obvious open reading frame. The HC21 data presented here further emphasize that current gene prediction algorithms miss a substantial number of transcripts that nevertheless can be identified using a combination of experimental approaches and multiple refined algorithms.

Chromosomes, Human, Pair 21↗

hMLH1 and hMSH2 gene mutation in Brazilian families with suspected hereditary nonpolyposis colorectal cancer.

BACKGROUND: The aim of this study was to search for mutations in the human mutS homolog 2 (hMSH2) and human mutL homolog 1 (hMLH1) genes in 25 unrelated Brazilian kindreds with suspected hereditary nonpolyposis colorectal cancer (HNPCC). METHODS: The families were grouped according to the following clinical criteria: Amsterdam I or II; familial colorectal cancer (CRC); an early age of onset of CRC in the proband only; or with at least one or two relatives who had HNPCC-related cancers; CRC in the proband only. All patients were studied with direct sequencing. RESULTS: Ten mutations were detected (10 of 25 [40%]); of nine different mutations, seven were novel. The hMLH1 gene had a higher mutation detection rate than hMSH2 (8 of 25 [32%] vs. 2 of 25 [8%]). Only 3 of these 10 families fulfilled the Amsterdam criteria. Two different polymorphisms were detected in the hMLH1 gene and four in the hMSH2 gene. CONCLUSIONS: The hMLH1 gene had a higher mutation detection rate than hMSH2. The physician who deals with CRC must take into consideration the heredity issue with patients who present with an early age of onset or a familial history of CRC- or HNPCC-related cancers, including gastric cancer, even if they do not fulfill the former Amsterdam criteria.

Adaptor Proteins, Signal Transducing↗

Human gene discovery through experimental definition of transcribed regions of the human genome.

The sequencing of the human genome has failed to realize its primary goal: the identification of all human genes. We have learned that genes can only be identified with certainty within this vast and information-sparse structure by comparison with transcript sequences. Significantly more sequence data of this kind is required before we can claim to have deciphered our genetic blueprint.

Algorithms↗

Long-range heterogeneity at the 3' ends of human mRNAs.

The publication of a draft of the human genome and of large collections of transcribed sequences has made it possible to study the complex relationship between the transcriptome and the genome. In the work presented here, we have focused on mapping mRNA 3' ends onto the genome by use of the raw data generated by the expressed sequence tag (EST) sequencing projects. We find that at least half of the human genes encode multiple transcripts whose polyadenylation is driven by multiple signals. The corresponding transcript 3' ends are spread over distances in the kilobase range. This finding has profound implications for our understanding of gene expression regulation and of the diversity of human transcripts, for the design of cDNA microarray probes, and for the interpretation of gene expression profiling experiments.

3' Flanking Region↗

Polymerase chain reaction and restriction fragment length polymorphism of cytocrome oxidase subunit I used for differentiation of Brazilian Biomphalaria species intermediate host of Schistosoma mansoni.

The intermediate hosts of Schistosoma mansoni, in Brazil, Biomphalaria glabrata, B. tenagophila and B. straminea, were identified by restriction fragment length polymorphism analysis of the mitochondrial gene cytochrome oxidase I (COI). We performed digestions with two enzymes (AluI and RsaI), previously selected, based on sequences available in Genbank. The profiles obtained with RsaI showed to be the most informative once they were polymorphic patterns, corroborating with much morphological data. In addition, we performed COI digestion of B. straminea snails from Uruguay and Argentina.

Animals↗