PubMed Health⌕ Search

PubMed · 15615859

GERBIL: Genotype resolution and block identification using likelihood.

Abstract

The abundance of genotype data generated by individual and international efforts carries the promise of revolutionizing disease studies and the association of phenotypes with individual polymorphisms. A key challenge is providing an accurate resolution (phasing) of the genotypes into haplotypes. We present here results on a method for genotype phasing in the presence of recombination. Our analysis is based on a stochastic model for recombination-poor regions ("blocks"), in which haplotypes are generated from a small number of core haplotypes, allowing for mutations, rare recombinations, and errors. We formulate genotype resolution and block partitioning as a maximum-likelihood problem and solve it by an expectation-maximization algorithm. The algorithm was implemented in a software package called GERBIL (genotype resolution and block identification using likelihood), which is efficient and simple to use. We tested GERBIL on four large-scale sets of genotypes. It outperformed two state-of-the-art phasing algorithms. The phase algorithm was slightly more accurate than GERBIL when allowed to run with default parameters, but required two orders of magnitude more time. When using comparable running times, GERBIL was consistently more accurate. For data sets with hundreds of genotypes, the time required by phase becomes prohibitive. We conclude that GERBIL has a clear advantage for studies that include many hundreds of genotypes and, in particular, for large-scale disease studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gad Kimmel, Ron Shamir. 2004-12-22. GERBIL: Genotype resolution and block identification using likelihood.. https://doi.org/10.1073/pnas.0404730102

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Mitotic karyotyping and FISH mapping of the gender-specific locus indicate an advanced XY system in Hippophae rhamnoides.

Hippophae rhamnoides ssp. turkestanica, a subdioecious plant inhabiting the cold desert of the Indian Himalaya, has gained immense recognition for its nutritional and medicinal values. In recent years, the plant species has proven to be a suitable system to understand the evolution of dioecy. Despite its biological significance, the cytogenetics of this dioecious plant is unclear due to various conflicting accounts of its X-Y chromosome system, particularly the length of Y-chromosome. In this study, we resolved these ambiguities through comprehensive cytogenetic analyses across diverse western Himalayan populations. Using morphometric analysis and fluorescence in situ hybridization (FISH) with a gender-specific marker (HRMSSR), we confirmed homomorphic XX chromosomes in females and heteromorphic sex-chromosomes in males with a notably smaller Y-chromosome. The investigation also revealed a predominant somatic chromosome number of 2n = 24, although minor deviations (2n = 18, 20, 22) appeared at the seed level. These findings highlight an evolutionarily advanced sex-chromosome system. This first detailed cytogenetic investigation of Himalayan Seabuckthorn provides critical insights into the chromosomal architecture, laying a crucial foundation for future evolutionary, genomic, and conservation studies in the species.

Chromosome Mapping↗

Subtractive hybridization magnetic bead capture: a new technique for the recovery of full-length ORFs from the metagenome.

A new method for the recovery of full-length open reading frames from metagenomic nucleic acid samples is reported. This technique, based on subtractive hybridization magnetic bead capture technology, has the potential to access multiple gene variants from a single amplification reaction. It is now widely accepted that classical microbiological methods provide only limited access to the true microbial biodiversity (less than 1%). The desire to access a higher proportion of the metagenome has led to the development of efficient environmental nucleic acid extraction technologies and to a range of sequence-dependent and sequence-independent gene discovery techniques. These methods avoid many of the limitations of culture-dependent gene targeting.

Chromosome Mapping↗

The elusive goal of pedigree weights.

Non-parametric linkage analysis methods generally involve calculating an allele-sharing statistic for each pedigree in a data set, then standardizing and summing the statistics over pedigrees. Pedigrees of different sizes can be weighted differently in the sum, though it is perhaps most common to weight all standardized pedigree statistics equally. Most other common weighting schemes are based on the number of affected individuals in the pedigree. It is also possible to derive optimal weights, which maximize power to detect linkage under particular trait models. We started by investigating three different analytical and simulation-based methods to calculate power and derive optimal weights. We found that simulation methods produce noticeably more accurate power calculations than the other methods. However, although the different calculation methods give different "optimal" weights, the power at those weights is very similar. That is, the analytical calculation methods are sufficient for finding good weights even though the simulation methods are most appropriate for calculating power. In comparing optimal weights for different trait models, we found that the weights vary quite a bit with the model, such that optimal weights for one model are not necessarily powerful at all for other models. Finally, we studied the power of a number of general weighting schemes, and of some new ones that incorporate information on how closely the affected individuals are related. We were able to find some schemes that performed well in the sense of giving reasonably powerful weights for most of the trait models and pedigree types we considered.

Chromosome Mapping↗