PubMed Health⌕ Search

PubMed · 16773561

Mapping tumor-suppressor genes with multipoint statistics from copy-number-variation data.

Abstract

Array-based comparative genomic hybridization (arrayCGH) is a microarray-based comparative genomic hybridization technique that has been used to compare tumor genomes with normal genomes, thus providing rapid genomic assays of tumor genomes in terms of copy-number variations of those chromosomal segments that have been gained or lost. When properly interpreted, these assays are likely to shed important light on genes and mechanisms involved in the initiation and progression of cancer. Specifically, chromosomal segments, deleted in one or both copies of the diploid genomes of a group of patients with cancer, point to locations of tumor-suppressor genes (TSGs) implicated in the cancer. In this study, we focused on automatic methods for reliable detection of such genes and their locations, and we devised an efficient statistical algorithm to map TSGs, using a novel multipoint statistical score function. The proposed algorithm estimates the location of TSGs by analyzing segmental deletions (hemi- or homozygous) in the genomes of patients with cancer and the spatial relation of the deleted segments to any specific genomic interval. The algorithm assigns, to an interval of consecutive probes, a multipoint score that parsimoniously captures the underlying biology. It also computes a P value for every putative TSG by using concepts from the theory of scan statistics. Furthermore, it can identify smaller sets of predictive probes that can be used as biomarkers for diagnosis and therapeutics. We validated our method using different simulated artificial data sets and one real data set, and we report encouraging results. We discuss how, with suitable modifications to the underlying statistical model, this algorithm can be applied generally to a wider class of problems (e.g., detection of oncogenes).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Iuliana Ionita, Raoul-Sam Daruwala, Bud Mishra. 2006-05-30. Mapping tumor-suppressor genes with multipoint statistics from copy-number-variation data.. https://doi.org/10.1086/504354

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Cre-lox-based system for multiple gene deletions and selectable-marker removal in Lactobacillus plantarum.

The classic strategy to achieve gene deletion variants is based on double-crossover integration of nonreplicating vectors into the genome. In addition, recombination systems such as Cre-lox have been used extensively, mainly for eukaryotic organisms. This study presents the construction of a Cre-lox-based system for multiple gene deletions in Lactobacillus plantarum that could be adapted for use on gram-positive bacteria. First, an effective mutagenesis vector (pNZ5319) was constructed that allows direct cloning of blunt-end PCR products representing homologous recombination target regions. Using this mutagenesis vector, double-crossover gene replacement mutants could be readily selected based on their antibiotic resistance phenotype. In the resulting mutants, the target gene is replaced by a lox66-P(32)-cat-lox71 cassette, where lox66 and lox71 are mutant variants of loxP and P(32)-cat is a chloramphenicol resistance cassette. The lox sites serve as recognition sites for the Cre enzyme, a protein that belongs to the integrase family of site-specific recombinases. Thus, transient Cre recombinase expression in double-crossover mutants leads to recombination of the lox66-P(32)-cat-lox71 cassette into a double-mutant loxP site, called lox72, which displays strongly reduced recognition by Cre. The effectiveness of the Cre-lox-based strategy for multiple gene deletions was demonstrated by construction of both single and double gene deletions at the melA and bsh1 loci on the chromosome of the gram-positive model organism Lactobacillus plantarum WCFS1. Furthermore, the efficiency of the Cre-lox-based system in multiple gene replacements was determined by successive mutagenesis of the genetically closely linked loci melA and lacS2 in L. plantarum WCFS1. The fact that 99.4% of the clones that were analyzed had undergone correct Cre-lox resolution emphasizes the suitability of the system described here for multiple gene replacement and deletion strategies in a single genetic background.

Gene Deletion↗

Role of large sequence polymorphisms (LSPs) in generating genomic diversity among clinical isolates of Mycobacterium tuberculosis and the utility of LSPs in phylogenetic analysis.

Mycobacterium tuberculosis strains contain different genomic insertions or deletions called large sequence polymorphisms (LSPs). Distinguishing between LSPs that occur one time versus ones that occur repeatedly in a genomic region may provide insights into the biological roles of LSPs and identify useful phylogenetic markers. We analyzed 163 clinical M. tuberculosis isolates for 17 LSPs identified in a genomic comparison of M. tuberculosis strains H37Rv and CDC1551. LSPs were mapped onto a single-nucleotide polymorphism (SNP)-based phylogenetic tree created using nine novel SNP markers that were found to reproduce a 212-SNP-based phylogeny. Four LSPs (group A) mapped to a single SNP tree segment. Two LSPs (group B) and 11 LSPs (group C) were inferred to have arisen independently in the same genomic region either two or more than two times, respectively. None of the group A LSPs but one group B LSP and five group C LSPs were flanked by IS6110 sequences in the references strains. Genes encoding members of the proline-glutamic acid or proline-proline-glutamic acid protein families were present only in group B or C LSPs. SNP- versus LSP-based phylogenies were also compared. We classified each isolate into 58 LSP types by using a separate LSP-based phylogenetic analysis and mapped the LSP types onto the SNP tree. LSPs often assigned isolates to the correct phylogenetic lineage; however, significant mistakes occurred for 6/58 (10%) of the LSP types. In conclusion, most LSPs occur in genomic regions that are prone to repeated insertion/deletion events and were responsible for an unexpectedly high degree of genomic variation in clinical M. tuberculosis. Group B and C LSPs may represent polymorphisms that occur due to selective pressure and affect the phenotype of the organism, while group A LSPs are preferable phylogenetic markers.

Gene Deletion↗

The shape of human gene family phylogenies.

BACKGROUND: The shape of phylogenetic trees has been used to make inferences about the evolutionary process by comparing the shapes of actual phylogenies with those expected under simple models of the speciation process. Previous studies have focused on speciation events, but gene duplication is another lineage splitting event, analogous to speciation, and gene loss or deletion is analogous to extinction. Measures of the shape of gene family phylogenies can thus be used to investigate the processes of gene duplication and loss. We make the first systematic attempt to use tree shape to study gene duplication using human gene phylogenies. RESULTS: We find that gene duplication has produced gene family trees significantly less balanced than expected from a simple model of the process, and less balanced than species phylogenies: the opposite to what might be expected under the 2R hypothesis. CONCLUSION: While other explanations are plausible, we suggest that the greater imbalance of gene family trees than species trees is due to the prevalence of tandem duplications over regional duplications during the evolution of the human genome.

Gene Deletion↗