PubMed Health⌕ Search

Biomedical subjects

Richa Agarwala

Publications and source records attributed to Richa Agarwala.

18 recordsLinked to original sources

Composition-based statistics and translated nucleotide searches: improving the TBLASTN module of BLAST.

BACKGROUND: TBLASTN is a mode of operation for BLAST that aligns protein sequences to a nucleotide database translated in all six frames. We present the first description of the modern implementation of TBLASTN, focusing on new techniques that were used to implement composition-based statistics for translated nucleotide searches. Composition-based statistics use the composition of the sequences being aligned to generate more accurate E-values, which allows for a more accurate distinction between true and false matches. Until recently, composition-based statistics were available only for protein-protein searches. They are now available as a command line option for recent versions of TBLASTN and as an option for TBLASTN on the NCBI BLAST web server. RESULTS: We evaluate the statistical and retrieval accuracy of the E-values reported by a baseline version of TBLASTN and by two variants that use different types of composition-based statistics. To test the statistical accuracy of TBLASTN, we ran 1000 searches using scrambled proteins from the mouse genome and a database of human chromosomes. To test retrieval accuracy, we modernize and adapt to translated searches a test set previously used to evaluate the retrieval accuracy of protein-protein searches. We show that composition-based statistics greatly improve the statistical accuracy of TBLASTN, at a small cost to the retrieval accuracy. CONCLUSION: TBLASTN is widely used, as it is common to wish to compare proteins to chromosomes or to libraries of mRNAs. Composition-based statistics improve the statistical accuracy, and therefore the reliability, of TBLASTN results. The algorithms used by TBLASTN are not widely known, and some of the most important are reported here. The data used to test TBLASTN are available for download and may be useful in other studies of translated search algorithms.

Algorithms↗

Retrieval accuracy, statistical significance and compositional similarity in protein sequence database searches.

Protein sequence database search programs may be evaluated both for their retrieval accuracy--the ability to separate meaningful from chance similarities--and for the accuracy of their statistical assessments of reported alignments. However, methods for improving statistical accuracy can degrade retrieval accuracy by discarding compositional evidence of sequence relatedness. This evidence may be preserved by combining essentially independent measures of alignment and compositional similarity into a unified measure of sequence similarity. A version of the BLAST protein database search program, modified to employ this new measure, outperforms the baseline program in both retrieval and statistical accuracy on ASTRAL, a SCOP-based test set.

Data Interpretation, Statistical↗

A 1.5-Mb-resolution radiation hybrid map of the cat genome and comparative analysis with the canine and human genomes.

We report the construction of a 1.5-Mb-resolution radiation hybrid map of the domestic cat genome. This new map includes novel microsatellite loci and markers derived from the 2X genome sequence that target previous gaps in the feline-human comparative map. Ninety-six percent of the 1793 cat markers we mapped have identifiable orthologues in the canine and human genome sequences. The updated autosomal and X-chromosome comparative maps identify 152 cat-human and 134 cat-dog homologous synteny blocks. Comparative analysis shows the marked change in chromosomal evolution in the canid lineage relative to the felid lineage since divergence from their carnivoran ancestor. The canid lineage has a 30-fold difference in the number of interchromosomal rearrangements relative to felids, while the felid lineage has primarily undergone intrachromosomal rearrangements. We have also refined the pseudoautosomal region and boundary in the cat and show that it is markedly longer than those of human or mouse. This improved RH comparative map provides a useful tool to facilitate positional cloning studies in the feline model.

Animals↗

High-resolution gene maps of horse chromosomes 14 and 21: additional insights into evolution and rearrangements of HSA5 homologs in mammals.

High-resolution physically ordered gene maps for equine homologs of human chromosome 5 (HSA5), viz., horse chromosomes 14 and 21 (ECA14 and ECA21), were generated by adding 179 new loci (131 gene-specific and 48 microsatellites) to the existing maps of the two chromosomes. The loci were mapped primarily by genotyping on a 5000-rad horse x hamster radiation hybrid panel, of which 28 were mapped by fluorescence in situ hybridization. The approximately fivefold increase in the number of mapped markers on the two chromosomes improves the average resolution of the map to 1 marker/0.9 Mb. The improved resolution is vital for rapid chromosomal localization of traits of interest on these chromosomes and for facilitating candidate gene searches. The comparative gene mapping data on ECA14 and ECA21 finely align the chromosomes to sequence/gene maps of a range of evolutionarily distantly related species. It also demonstrates that compared to ECA14, the ECA21 segment corresponding to HSA5 is a more conserved region because of preserved gene order in a larger number of and more diverse species. Further, comparison of ECA14 and the distal three-quarters region of ECA21 with corresponding chromosomal segments in 50 species belonging to 11 mammalian orders provides a broad overview of the evolution of these segments in individual orders from the putative ancestral chromosomal configuration. Of particular interest is the identification and precise demarcation of equid/Perissodactyl-specific features that for the first time clearly distinguish the origins of ECA14 and ECA21 from similar-looking status in the Cetartiodactyls.

Animals↗

An approximately 140-kb deletion associated with feline spinal muscular atrophy implies an essential LIX1 function for motor neuron survival.

The leading genetic cause of infant mortality is spinal muscular atrophy (SMA), a clinically and genetically heterogeneous group of disorders. Previously we described a domestic cat model of autosomal recessive, juvenile-onset SMA similar to human SMA type III. Here we report results of a whole-genome scan for linkage in the feline SMA pedigree using recently developed species-specific and comparative mapping resources. We identified a novel SMA gene candidate, LIX1, in an approximately140-kb deletion on feline chromosome A1q in a region of conserved synteny to human chromosome 5q15. Though LIX1 function is unknown, the predicted secondary structure is compatible with a role in RNA metabolism. LIX1 expression is largely restricted to the central nervous system, primarily in spinal motor neurons, thus offering explanation of the tissue restriction of pathology in feline SMA. An exon sequence screen of 25 human SMA cases, not otherwise explicable by mutations at the SMN1 locus, failed to identify comparable LIX1 mutations. Nonetheless, a LIX1-associated etiology in feline SMA implicates a previously undetected mechanism of motor neuron maintenance and mandates consideration of LIX1 as a candidate gene in human SMA when SMN1 mutations are not found.

Animals↗

Novel gene acquisition on carnivore Y chromosomes.

Despite its importance in harboring genes critical for spermatogenesis and male-specific functions, the Y chromosome has been largely excluded as a priority in recent mammalian genome sequencing projects. Only the human and chimpanzee Y chromosomes have been well characterized at the sequence level. This is primarily due to the presumed low overall gene content and highly repetitive nature of the Y chromosome and the ensuing difficulties using a shotgun sequence approach for assembly. Here we used direct cDNA selection to isolate and evaluate the extent of novel Y chromosome gene acquisition in the genome of the domestic cat, a species from a different mammalian superorder than human, chimpanzee, and mouse (currently being sequenced). We discovered four novel Y chromosome genes that do not have functional copies in the finished human male-specific region of the Y or on other mammalian Y chromosomes explored thus far. Two genes are derived from putative autosomal progenitors, and the other two have X chromosome homologs from different evolutionary strata. All four genes were shown to be multicopy and expressed predominantly or exclusively in testes, suggesting that their duplication and specialization for testis function were selected for because they enhance spermatogenesis. Two of these genes have testis-expressed, Y-borne copies in the dog genome as well. The absence of the four newly described genes on other characterized mammalian Y chromosomes demonstrates the gene novelty on this chromosome between mammalian orders, suggesting it harbors many lineage-specific genes that may go undetected by traditional comparative genomic approaches. Specific plans to identify the male-specific genes encoded in the Y chromosome of mammals should be a priority.

Animals↗

A fast and symmetric DUST implementation to mask low-complexity DNA sequences.

The DUST module has been used within BLAST for many years to mask low-complexity sequences. In this paper, we present a new implementation of the DUST module that uses the same function to assign a complexity score to a sequence, but uses a different rule by which high-scoring sequences are masked. The new rule masks every nucleotide masked by the old rule and occasionally masks more. The new masking rule corrects two related deficiencies with the old rule. First, the new rule is symmetric with respect to reversing the sequence. Second, the new rule is not context sensitive; the decision to mask a subsequence does not depend on what sequences flank it. The new implementation is at least four times faster than the old on the human genome. We show that both the percentage of additional bases masked and the effect on MegaBLAST outputs are very small.

Genome, Human↗

Does having children extend life span? A genealogical study of parity and longevity in the Amish.

BACKGROUND: The relationship between parity and life span is uncertain, with evidence of both positive and negative relationships being reported previously. We evaluated this issue by using genealogical data from an Old Order Amish community in Lancaster, Pennsylvania, a population characterized by large nuclear families, homogeneous lifestyle, and extensive genealogical records. METHODS: The analysis was restricted to the set of 2,015 individuals who had children, were born between 1749 and 1912, and survived until at least age 50 years. Pedigree structures and birth and death dates were extracted from Amish genealogies, and the relationship between parity and longevity was examined using a variance component framework. RESULTS: Life span of fathers increased in linear fashion with increasing number of children (0.23 years per additional child; p =.01), while life span of mothers increased linearly up to 14 children (0.32 years per additional child; p =.004) but decreased with each additional child beyond 14 (p =.0004). Among women, but not men, a later age at last birth was associated with longer life span (p =.001). Adjusting for age at last birth obliterated the correlation between maternal life span and number of children, except among mothers with ultrahigh (>14 children) parity. CONCLUSIONS: We conclude that high parity among men and later menopause among women may be markers for increased life span. Understanding the biological and/or social factors mediating these relationships may provide insights into mechanisms underlying successful aging.

Aged↗

WindowMasker: window-based masker for sequenced genomes.

MOTIVATION: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes. RESULTS: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis. AVAILABILITY: WM is included in the NCBI C++ toolkit. The source code for the entire toolkit is available at ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools++/CURRENT/. Once the toolkit source is unpacked, the instructions for building WindowMasker application in the UNIX environment can be found in file src/app/winmasker/README.build. SUPPLEMENTARY INFORMATION: Supplementary data are available at ftp://ftp.ncbi.nlm.nih.gov/pub/agarwala/windowmasker/windowmasker_suppl.pdf

Algorithms↗

A high-resolution physical map of equine homologs of HSA19 shows divergent evolution compared with other mammals.

A high-resolution (1 marker/700 kb) physically ordered radiation hybrid (RH) and comparative map of 122 loci on equine homologs of human Chromosome 19 (HSA19) shows a variant evolution of these segments in equids/Perissodactyls compared with other mammals. The segments include parts of both the long and the short arm of horse Chromosome 7 (ECA7), the proximal part of ECA21, and the entire short arm of ECA10. The map includes 93 new markers, of which 89 (64 gene-specific and 25 microsatellite) were genotyped on a 5000-rad horse x hamster RH panel, and 4 were mapped exclusively by FISH. The orientation and alignment of the map was strengthened by 21 new FISH localizations, of which 15 represent genes. The approximately sevenfold-improved map resolution attained in this study will prove extremely useful for candidate gene discovery in the targeted equine chromosomal regions. The highlight of the comparative map is the fine definition of homology between the four equine chromosomal segments and corresponding HSA19 regions specified by physical coordinates (bp) in the human genome sequence. Of particular interest are the regions on ECA7 and ECA21 that correspond to the short arm of HSA19-a genomic rearrangement discovered to date only in equids/Perissodactyls as evidenced through comparative Zoo-FISH analysis of the evolution of ancestral HSA19 segments in eight mammalian orders involving about 50 species.

Animals↗

A rhesus macaque radiation hybrid map and comparative analysis with the human genome.

The genomes of nonhuman primates are powerful references for better understanding the recent evolution of the human genome. Here we compare the order of 802 genomic markers mapped in a rhesus macaque (Macaca mulatta) radiation hybrid panel with the human genome, allowing for nearly complete cross-reference to the human genome at an average resolution of 3.5 Mb. At least 23 large-scale chromosomal rearrangements, mostly inversions, are needed to explain the changes in marker order between human and macaque. Analysis of the breakpoints flanking inverted chromosomal segments and estimation of their duplication divergence dates provide additional evidence implicating segmental duplications as a major mechanism of chromosomal rearrangement in recent primate evolution.

Animals↗

Protein database searches using compositionally adjusted substitution matrices.

Almost all protein database search methods use amino acid substitution matrices for scoring, optimizing, and assessing the statistical significance of sequence alignments. Much care and effort has therefore gone into constructing substitution matrices, and the quality of search results can depend strongly upon the choice of the proper matrix. A long-standing problem has been the comparison of sequences with biased amino acid compositions, for which standard substitution matrices are not optimal. To address this problem, we have recently developed a general procedure for transforming a standard matrix into one appropriate for the comparison of two sequences with arbitrary, and possibly differing compositions. Such adjusted matrices yield, on average, improved alignments and alignment scores when applied to the comparison of proteins with markedly biased compositions. Here we review the application of compositionally adjusted matrices and consider whether they may also be applied fruitfully to general purpose protein sequence database searches, in which related sequence pairs do not necessarily have strong compositional biases. Although it is not advisable to apply compositional adjustment indiscriminately, we describe several simple criteria under which invoking such adjustment is on average beneficial. In a typical database search, at least one of these criteria is satisfied by over half the related sequence pairs. Compositional substitution matrix adjustment is now available in NCBI's protein-protein version of blast.

Algorithms↗

Reduced incidence of hip fracture in the Old Order Amish.

UNLABELLED: The incidence of hip fracture was estimated in a community of Old Order Amish and compared with available data from non-Amish whites. Hip fracture rates were 40% lower in the Amish, and the Amish also experienced higher BMD. INTRODUCTION: Understanding the patterns of fracture risk across populations could reveal insights about bone health and lead to the earlier detection and prevention of osteoporosis. Toward this aim, we compared hip fracture incidence and bone mineral density (BMD) between an Old Order Amish (OOA) community, characterized by a rural and relatively active lifestyle, and non-Amish U.S. whites. MATERIALS AND METHODS: All hospital admissions for hip fracture among OOA individuals in Lancaster County, PA, were identified between 1995 and 1998 from four area hospitals. Hip fracture incidence was calculated by cross-referencing an available Anabaptist genealogy database with communities located within these hospital service areas and compared with non-Amish whites obtained from National Hospital Discharge data. Additionally, BMD at the hip was compared between 287 Amish subjects and non-Amish whites from the National Health and Nutrition Examination Survey III survey. RESULTS AND CONCLUSIONS: OOA experienced 42% fewer hip fractures than would be expected had they experienced the same rate of hip fracture as observed in non-Amish whites (p < 0.01) and a higher mean BMD that was significant in women (p < 0.05) but not men. Further evaluation of lifestyle and/or genetic differences between Amish and non-Amish populations may shed insights into etiologic factors influencing hip fracture risk.

Absorptiometry, Photon↗

Anabaptist genealogy database.

In late 1996 we set out to build a computer-searchable genealogy of the Old Order Amish of Lancaster County, Pennsylvania, for use by geneticists. The goals of the project included: 1) using the genealogy to expedite the mapping of genes mutated in three rare recessive disorders under study at the National Institutes of Health (NIH); 2) building a freely available software package, PedHunter, to answer genetically relevant queries on our database and other similar databases; and 3) providing genealogy assistance to researchers outside NIH. All of these scientific goals had to be accomplished while maintaining the confidentiality of the persons in the database and the confidentiality of preliminary research results. We expanded the project to include complementary data sources that contained many individuals who were Anabaptist, but not Amish, and many individuals who never lived in Lancaster County. For this reason, the project was renamed Anabaptist Genealogy Database (AGDB). All of the initial goals of the project have been accomplished, and we recently marked the 5-year anniversary of answering the first of over 100 queries by researchers outside NIH. Thus, it is an opportune time to review the construction of AGDB, summarize its usage to date, and speculate on future projects it might stimulate and facilitate.

Databases, Genetic↗

A genome-wide scan for primary open-angle glaucoma (POAG): the Barbados Family Study of Open-Angle Glaucoma.

Primary open-angle glaucoma (POAG) is characterized by damage to the optic nerve with associated loss of vision. Six named genetic loci have been identified as contributing to POAG susceptibility by genetic linkage analysis of mostly Caucasian families, and two of the six causative genes have been identified. The Barbados Family Study of Open-Angle Glaucoma (BFSG) was designed to evaluate the genetic component of POAG in a population of African descent. A genome-wide scan was performed on 1327 individuals from 146 families in Barbados, West Indies. Linkage results were based on models and parameter estimates derived from a segregation analysis of these families, and on model-free analyses. Two-point LOD scores >1.0 were identified on chromosomes 1, 2, 9, 10, 11, and 14, with increased multipoint LOD scores being found on chromosomes 2, 10, and 14. Fine mapping was subsequently carried out and indicated that POAG may be linked to intervals on chromosome 2q between D2S2188 and D2S2178 and chromosome 10p between D10S1477 and D10S601. Heterogeneity testing strongly supports linkage for glaucoma to at least one of these regions and suggests possible linkages to both. Although TIGR/myocilin and optineurin mutations have been shown to be causally linked to POAG in other populations, findings from this study do not support either of these as causative genes in an Afro-Caribbean population known to have relatively high rates of POAG.

Barbados↗

Sequence variations in the public human genome data reflect a bottlenecked population history.

Single-nucleotide polymorphisms (SNPs) constitute the great majority of variations in the human genome, and as heritable variable landmarks they are useful markers for disease mapping and resolving population structure. Redundant coverage in overlaps of large-insert genomic clones, sequenced as part of the Human Genome Project, comprises a quarter of the genome, and it is representative in terms of base compositional and functional sequence features. We mined these regions to produce 500,000 high-confidence SNP candidates as a uniform resource for describing nucleotide diversity and its regional variation within the genome. Distributions of marker density observed at different overlap length scales under a model of recombination and population size change show that the history of the population represented by the public genome sequence is one of collapse followed by a recent phase of mild size recovery. The inferred times of collapse and recovery are Upper Paleolithic, in agreement with archaeological evidence of the initial modern human colonization of Europe.

Databases, Nucleic Acid↗

Initial sequencing and comparative analysis of the mouse genome.

The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.

Animals↗

Mutant deoxynucleotide carrier is associated with congenital microcephaly.

The disorder Amish microcephaly (MCPHA) is characterized by severe congenital microcephaly, elevated levels of alpha-ketoglutarate in the urine and premature death. The disorder is inherited in an autosomal recessive pattern and has been observed only in Old Order Amish families whose ancestors lived in Lancaster County, Pennsylvania. Here we show, by using a genealogy database and automated pedigree software, that 23 nuclear families affected with MCPHA are connected to a single ancestral couple. Through a whole-genome scan, fine mapping and haplotype analysis, we localized the gene affected in MCPHA to a region of 3 cM, or 2 Mb, on chromosome 17q25. We constructed a map of contiguous genomic clones spanning this region. One of the genes in this region, SLC25A19, which encodes a nuclear mitochondrial deoxynucleotide carrier (DNC), contains a substitution that segregates with the disease in affected individuals and alters an amino acid that is highly conserved in similar proteins. Functional analysis shows that the mutant DNC protein lacks the normal transport activity, implying that failed deoxynucleotide transport across the inner mitochondrial membrane causes MCPHA. Our data indicate that mitochondrial deoxynucleotide transport may be essential for prenatal brain growth.

Carrier Proteins↗