PubMed Health⌕ Search

Biomedical subjects

Marco Davila

Publications and source records attributed to Marco Davila.

3 recordsLinked to original sources

Computational tools for understanding sequence variability in recombination signals.

The recombination signals (RSs) that guide V(D)J rearrangement are remarkably diverse. In mice, fewer than 16% of RSs carry consensus heptamers and nonamers and none also contain a consensus spacer sequence. It is increasingly clear that this variability regulates recombination: genetic variability in RSs may help enforce allelic exclusion, determine the general nature of antigen receptor repertoires, and mitigate autoreactivity in B lymphocytes. The great diversity of RSs has largely precluded, however, empiric determinations of how RS sequence affects recombination. For example, 4(39) unique 23-RSs are possible or approximately 3 x 10(23) sequences; some 7 x 10(13) unique 23-RSs can be produced just by changes in the spacer. In contrast, the recombination activities of only 100 or so RSs have been measured, and it is unlikely that the activities of even a tiny fraction of extant RSs can be determined. We have addressed the problem of how sequence determines the efficiency of RS templates by generating computational models that describe the correlation structure of mouse RSs. These models successfully predict RS activity and identify functional, cryptic RSs (cRSs). These models permit studies to identify RSs and cRSs for empiric study and constitute a tool useful for understanding RS structure and function.

Animals↗

Prospective estimation of recombination signal efficiency and identification of functional cryptic signals in the genome by statistical modeling.

The recombination signals (RS) that guide V(D)J recombination are phylogenetically conserved but retain a surprising degree of sequence variability, especially in the nonamer and spacer. To characterize RS variability, we computed the position-wise information, a measure correlated with sequence conservation, for each nucleotide position in an RS alignment and demonstrate that most position-wise information is present in the RS heptamers and nonamers. We have previously demonstrated significant correlations between RS positions and here show that statistical models of the correlation structure that underlies RS variability efficiently identify physiologic and cryptic RS and accurately predict the recombination efficiencies of natural and synthetic RS. In scans of mouse and human genomes, these models identify a highly conserved family of repetitive DNA as an unexpected source of frequent, cryptic RS that rearrange both in extrachromosomal substrates and in their genomic context.

Animals↗

Identification and utilization of arbitrary correlations in models of recombination signal sequences.

BACKGROUND: A significant challenge in bioinformatics is to develop methods for detecting and modeling patterns in variable DNA sequence sites, such as protein-binding sites in regulatory DNA. Current approaches sometimes perform poorly when positions in the site do not independently affect protein binding. We developed a statistical technique for modeling the correlation structure in variable DNA sequence sites. The method places no restrictions on the number of correlated positions or on their spatial relationship within the site. No prior empirical evidence for the correlation structure is necessary. RESULTS: We applied our method to the recombination signal sequences (RSS) that direct assembly of B-cell and T-cell antigen-receptor genes via V(D)J recombination. The technique is based on model selection by cross-validation and produces models that allow computation of an information score for any signal-length sequence. We also modeled RSS using order zero and order one Markov chains. The scores from all models are highly correlated with measured recombination efficiencies, but the models arising from our technique are better than the Markov models at discriminating RSS from non-RSS. CONCLUSIONS: Our model-development procedure produces models that estimate well the recombinogenic potential of RSS and are better at RSS recognition than the order zero and order one Markov models. Our models are, therefore, valuable for studying the regulation of both physiologic and aberrant V(D)J recombination. The approach could be equally powerful for the study of promoter and enhancer elements, splice sites, and other DNA regulatory sites that are highly variable at the level of individual nucleotide positions.

Animals↗