PubMed Health⌕ Search

Biomedical subjects

John W Keele

Publications and source records attributed to John W Keele.

15 recordsLinked to original sources

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., ≥10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Animals↗

Novel porcine repetitive elements.

BACKGROUND: Repetitive elements comprise approximately 45% of mammalian genomes and are increasingly known to impact genomic function by contributing to the genomic architecture, by direct regulation of gene expression and by affecting genomic size, diversity and evolution. The ubiquity and increasingly understood importance of repetitive elements contribute to the need to identify and annotate them. We set out to identify previously uncharacterized repetitive DNA in the porcine genome. Once found, we characterized the prevalence of these repeats in other mammals. RESULTS: We discovered 27 repetitive elements in 220 BACs covering 1% of the porcine genome (Comparative Vertebrate Sequencing Initiative; CVSI). These repeats varied in length from 55 to 1059 nucleotides. To estimate copy numbers, we went to an independent source of data, the BAC-end sequences (Wellcome Trust Sanger Institute), covering approximately 15% of the porcine genome. Copy numbers in BAC-ends were less than one hundred for 6 repeat elements, between 100 and 1000 for 16 and between 1,000 and 10,000 for 5. Several of the repeat elements were found in the bovine genome and we have identified two with orthologous sites, indicating that these elements were present in their common ancestor. None of the repeat elements were found in primate, rodent or dog genomes. We were unable to identify any of the replication machinery common to active transposable elements in these newly identified repeats. CONCLUSION: The presence of both orthologous and non-orthologous sites indicates that some sites existed prior to speciation and some were generated later. The identification of low to moderate copy number repetitive DNA that is specific to artiodactyls will be critical in the assembly of livestock genomes and studies of comparative genomics.

Animals↗

Prion gene haplotypes of U.S. cattle.

BACKGROUND: Bovine spongiform encephalopathy (BSE) is a fatal neurological disorder characterized by abnormal deposits of a protease-resistant isoform of the prion protein. Characterizing linkage disequilibrium (LD) and haplotype networks within the bovine prion gene (PRNP) is important for 1) testing rare or common PRNP variation for an association with BSE and 2) interpreting any association of PRNP alleles with BSE susceptibility. The objective of this study was to identify polymorphisms and haplotypes within PRNP from the promoter region through the 3'UTR in a diverse sample of U.S. cattle genomes. RESULTS: A 25.2-kb genomic region containing PRNP was sequenced from 192 diverse U.S. beef and dairy cattle. Sequence analyses identified 388 total polymorphisms, of which 287 have not previously been reported. The polymorphism alleles define PRNP by regions of high and low LD. High LD is present between alleles in the promoter region through exon 2 (6.7 kb). PRNP alleles within the majority of intron 2, the entire coding sequence and the untranslated region of exon 3 are in low LD (18.0 kb). Two haplotype networks, one representing the region of high LD and the other the region of low LD yielded nineteen different combinations that represent haplotypes spanning PRNP. The haplotype combinations are tagged by 19 polymorphisms (htSNPS) which characterize variation within and across PRNP. CONCLUSION: The number of polymorphisms in the prion gene region of U.S. cattle is nearly four times greater than previously described. These polymorphisms define PRNP haplotypes that may influence BSE susceptibility in cattle.

3' Untranslated Regions↗

Characterization of 954 bovine full-CDS cDNA sequences.

BACKGROUND: Genome assemblies rely on the existence of transcript sequence to stitch together contigs, verify assembly of whole genome shotgun reads, and annotate genes. Functional genomics studies also rely on transcript sequence to create expression microarrays or interpret digital tag data produced by methods such as Serial Analysis of Gene Expression (SAGE). Transcript sequence can be predicted based on reconstruction from overlapping expressed sequence tags (EST) that are obtained by single-pass sequencing of random cDNA clones, but these reconstructions are prone to errors caused by alternative splice forms, transcripts from gene families with related sequences, and expressed pseudogenes. These errors confound genome assembly and annotation. The most useful transcript sequences are derived by complete insert sequencing of clones containing the entire length, or at least the full protein coding sequence (CDS) portion, of the source mRNA. While the bovine genome sequencing initiative is nearing completion, there is currently a paucity of bovine full-CDS mRNA and protein sequence data to support bovine genome assembly and functional genomics studies. Consequently, the production of high-quality bovine full-CDS cDNA sequences will enhance the bovine genome assembly and functional studies of bovine genes and gene products. The goal of this investigation was to identify and characterize the full-CDS sequences of bovine transcripts from clones identified in non-full-length enriched cDNA libraries. In contrast to several recent full-length cDNA investigations, these full-CDS cDNAs were selected, sequenced, and annotated without the benefit of the target organism's genomic sequence, by using comparison of bovine EST sequence to existing human mRNA to identify likely full-CDS clones for full-length insert cDNA (FLIC) sequencing. RESULTS: The predicted bovine protein lengths, 5' UTR lengths, and Kozak consensus sequences from 954 bovine FLIC sequences (bFLICs; average length 1713 nt, representing 762 distinct loci) are all consistent with previously sequenced mammalian full-length transcripts. CONCLUSION: In most cases, the bFLICs span the entire CDS of the genes, providing the basis for creating predicted bovine protein sequences to support proteomics and comparative evolutionary research as well as functional genomics and genome annotation. The results demonstrate the utility of the comparative approach in obtaining predicted protein sequences in other species.

5' Untranslated Regions↗

Linkage mapping bovine EST-based SNP.

BACKGROUND: Existing linkage maps of the bovine genome primarily contain anonymous microsatellite markers. These maps have proved valuable for mapping quantitative trait loci (QTL) to broad regions of the genome, but more closely spaced markers are needed to fine-map QTL, and markers associated with genes and annotated sequence are needed to identify genes and sequence variation that may explain QTL. RESULTS: Bovine expressed sequence tag (EST) and bacterial artificial chromosome (BAC)sequence data were used to develop 918 single nucleotide polymorphism (SNP) markers to map genes on the bovine linkage map. DNA of sires from the MARC reference population was used to detect SNPs, and progeny and mates of heterozygous sires were genotyped. Chromosome assignments for 861 SNPs were determined by twopoint analysis, and positions for 735 SNPs were established by multipoint analyses. Linkage maps of bovine autosomes with these SNPs represent 4585 markers in 2475 positions spanning 3058 cM. Markers include 3612 microsatellites, 913 SNPs and 60 other markers. Mean separation between marker positions is 1.2 cM. New SNP markers appear in 511 positions, with mean separation of 4.7 cM. Multi-allelic markers, mostly microsatellites, had a mean (maximum) of 216 (366) informative meioses, and a mean 3-lod confidence interval of 3.6 cM Bi-allelic markers, including SNP and other marker types, had a mean (maximum) of 55 (191) informative meioses, and were placed within a mean 8.5 cM 3-lod confidence interval. Homologous human sequences were identified for 1159 markers, including 582 newly developed and mapped SNP. CONCLUSION: Addition of these EST- and BAC-based SNPs to the bovine linkage map not only increases marker density, but provides connections to gene-rich physical maps, including annotated human sequence. The map provides a resource for fine-mapping quantitative trait loci and identification of positional candidate genes, and can be integrated with other data to guide and refine assembly of bovine genome sequence. Even after the bovine genome is completely sequenced, the map will continue to be a useful tool to link observable phenotypes and animal genotypes to underlying genes and molecular mechanisms influencing economically important beef and dairy traits.

Alleles↗

Software agents in molecular computational biology.

Progress made in applying agent systems to molecular computational biology is reviewed and strategies by which to exploit agent technology to greater advantage are investigated. Communities of software agents could play an important role in helping genome scientists design reagents for future research. The advent of genome sequencing in cattle and swine increases the complexity of data analysis required to conduct research in livestock genomics. Databases are always expanding and semantic differences among data are common. Agent platforms have been developed to deal with generic issues such as agent communication, life cycle management and advertisement of services (white and yellow pages). This frees computational biologists from the drudgery of having to re-invent the wheel on these common chores, giving them more time to focus on biology and bioinformatics. Agent platforms that comply with the Foundation for Intelligent Physical Agents (FIPA) standards are able to interoperate. In other words, agents developed on different platforms can communicate and cooperate with one another if domain-specific higher-level communication protocol details are agreed upon between different agent developers. Many software agent platforms are peer-to-peer, which means that even if some of the agents and data repositories are temporarily unavailable, a subset of the goals of the system can still be met. Past use of software agents in bioinformatics indicates that an agent approach should prove fruitful. Examination of current problems in bioinformatics indicates that existing agent platforms should be adaptable to novel situations.

Algorithms↗

Integrating linkage and radiation hybrid mapping data for bovine chromosome 15.

BACKGROUND: Bovine chromosome (BTA) 15 contains a quantitative trait loci (QTL) for meat tenderness, as well as several breaks in synteny with human chromosome (HSA) 11. Both linkage and radiation hybrid (RH) maps of BTA 15 are available, but the linkage map lacks gene-specific markers needed to identify genes underlying the QTL, and the gene-rich RH map lacks associations with marker genotypes needed to define the QTL. Integrating the maps will provide information to further explore the QTL as well as refine the comparative map between BTA 15 and HSA 11. A recently developed approach to integrating linkage and RH maps uses both linkage and RH data to resolve a consensus marker order, rather than aligning independently constructed maps. Automated map construction procedures employing this maximum-likelihood approach were developed to integrate BTA RH and linkage data, and establish comparative positions of BTA 15 markers with HSA 11 homologs. RESULTS: The integrated BTA 15 map represents 145 markers; 42 shared by both data sets, 36 unique to the linkage data and 67 unique to RH data. Sequence alignment yielded comparative positions for 77 bovine markers with homologs on HSA 11. The map covers approximately 32% of HSA 11 sequence in five segments of conserved synteny, another 15% of HSA 11 is shared with BTA 29. Bovine and human order are consistent in portions of the syntenic segments, but some rearrangement is apparent. Comparative positions of gene markers near the meat tenderness QTL indicate the region includes separate segments of HSA 11. The two microsatellite markers flanking the QTL peak are between defined syntenic segments. CONCLUSIONS: Combining data to construct an integrated map not only consolidates information from different sources onto a single map, but information contributed from each data set increases the accuracy of the map. Comparison of bovine maps with well annotated human sequence can provide useful information about genes near mapped bovine markers, but bovine gene order may be different than human. Procedures to connect genetic and physical mapping data, build integrated maps for livestock species, and connect those maps to more fully annotated sequence can be automated, facilitating the maintenance of up-to-date maps, and providing a valuable tool to further explore genetic variation in livestock.

Animals↗

Beta-2-microglobulin haplotypes in U.S. beef cattle and association with failure of passive transfer in newborn calves.

Failure of passive transfer (FPT) is a condition in which neonates do not acquire protective serum levels of maternal antibodies. A principal component of antibody transport is the neonatal receptor for the Fc portion of immunoglobulin, a heterodimer of a MHC-1 alpha-chain homolog ( FCGRT) and beta-2-microglobulin ( B2M). Previously, two FCGRT haplotypes were associated with differences in immunoglobulin G (IgG) passive transfer in cattle (Laegreid et al. (2002) Mamm Genome 13, 704-710). The present study had two objectives: first, to characterize the B2M haplotype structure in a diverse group of U.S. beef cattle, and second, to evaluate those haplotypes for association with either high or low serum IgG levels in newborn calves. Twelve single nucleotide polymorphisms (SNPs), assorted into eight haplotypes, were identified by sequencing regions of B2M exons II and IV in a multi-breed panel of 96 beef cattle. Calves homozygous for one of the eight haplotypes ( B2M 2,2) were at increased risk of FPT (odds ratio = 10.60, CI(95%) 2.07-54.24, p = 0.005). These results indicate that this haplotype is in linkage disequilibrium with genetic risk factors affecting passive transfer of IgG in beef calves, an important determinant of neonatal calf morbidity and mortality.

Animal Husbandry↗

Positional candidate gene selection from livestock EST databases using Gene Ontology.

MOTIVATION: The number of expressed sequence tags (ESTs) in GenBank has now surpassed 200,000 for cattle and 100,000 for swine. The Institute of Genome Research (TIGR) has organized these sequences into approximately 60,000 non-redundant consensus sequences (identified by TIGR Gene Indices) for cattle and 40,000 for swine. Anonymous ESTs are of limited value unless they are connected to function. Functional information is difficult to manage electronically because of heterogeneity of meaning and form among databases. The Gene Ontology (GO) Consortium has produced ontologies for gene function with consistent meaning and form across species. Linking livestock EST to gene function through similarity with sequences from other annotation-rich mammals could accelerate: (1) the discovery of positional candidate genes underlying a livestock quantitative trait locus (QTL) and (2) comparative mapping between livestock and other mammals (e.g. humans, mouse and rat). We initiated this investigation to determine if incorporation of the GO into the annotation process could accelerate livestock positional candidate gene discovery. RESULTS: We have associated livestock ESTs with GO nodes through sequence similarity to the NCBI Reference Sequences (RefSeq). Positional candidate genes are identified within minutes that otherwise required days. The schema described here accommodates queries that return GO nodes from terms familiar to biologists, such as gene name, alternate/alias symbol, and OMIM phenotype. AVAILABILITY: Scripts and schema are available on request from the authors.

Animals↗

Prion gene sequence variation within diverse groups of U.S. sheep, beef cattle, and deer.

Prions are proteins that play a central role in transmissible spongiform encephalopathies in a variety of mammals. Among the most notable prion disorders in ungulates are scrapie in sheep, bovine spongiform encephalopathy in cattle, and chronic wasting disease in deer. Single nucleotide polymorphisms in the sheep prion gene ( PRNP) have been correlated with susceptibility to natural scrapie in some populations. Similar correlations have not been reported in cattle or deer; however, characterization of PRNP nucleotide diversity in those species is incomplete. This report describes nucleotide sequence variation and frequency estimates for the PRNP locus within diverse groups of U.S. sheep, U.S. beef cattle, and free-ranging deer ( Odocoileus virginianus and O. hemionus from Wyoming). DNA segments corresponding to the complete prion coding sequence and a 596-bp portion of the PRNP promoter region were amplified and sequenced from DNA panels with 90 sheep, 96 cattle, and 94 deer. Each panel was designed to contain the most diverse germplasm available from their respective populations to facilitate polymorphism detection. Sequence comparisons identified a total of 86 polymorphisms. Previously unreported polymorphisms were identified in sheep (9), cattle (13), and deer (32). The number of individuals sampled within each population was sufficient to detect more than 95% of all alleles present at a frequency greater than 0.02. The estimation of PRNP allele and genotype frequencies within these diverse groups of sheep, cattle, and deer provides a framework for designing accurate genotype assays for use in genetic epidemiology, allele management, and disease control.

Amyloid↗

Porcine gene discovery by normalized cDNA-library sequencing and EST cluster assembly.

Genetic and environmental factors affect the efficiency of pork production by influencing gene expression during porcine reproduction, tissue development, and growth. The identification and functional analysis of gene products important to these processes would be greatly enhanced by the development of a database of expressed porcine gene sequence. Two normalized porcine cDNA libraries (MARC 1PIG and MARC 2PIG), derived respectively from embryonic and reproductive tissues, were constructed, sequenced, and analyzed. A total of 66,245 clones from these two libraries were 5?-end sequenced and deposited in GenBank. Cluster analysis revealed that within-library redundancy is low, and comparison of all porcine ESTs with the human database suggests that the sequences from these two libraries represent portions of a significant number of independent pig genes. A Porcine Gene Index (PGI), comprising 15,616 tentative consensus sequences and 31,466 singletons, includes all sequences in public repositories and has been developed to facilitate further comparative map development and characterization of porcine genes (http://www.tigr.org/tdb/ssgi/). The clones and sequences from these libraries provide a catalog of expressed porcine genes and a resource for development of high-density hybridization arrays for transcriptional profiling of porcine tissues. In addition, comparison of porcine ESTs with sequences from other species serves as a valuable resource for comparative map development. Both arrayed cDNA libraries are available for unrestricted public use.

Animals↗

Use of bovine EST data and human genomic sequences to map 100 gene-specific bovine markers.

A system to use bovine EST data in conjunction with human genomic sequence to improve the bovine linkage map over the entire genome or on specific chromosomes was evaluated. Bovine EST sequence was used to provide primer sequences corresponding to bovine genes, while human genomic sequence directed primer design to flank introns and produce amplicons of appropriate size for efficient direct sequencing. The sequence tagged sites (STS) produced in this way from the four sires of the MARC reference families were examined for single nucleotide polymorphisms (SNPs) that could be used to map the corresponding genes. With this approach, along with a primer/extension mass spectrometry SNP genotyping assay, 100 ESTs were placed on the bovine genetic linkage map. The first 70 were chosen at random from bovine EST-human genomic comparisons. An additional 30 ESTs were successfully mapped to bovine Chromosome 19 (BTA19), and comparison of the resulting BTA19 map to the position of the corresponding human orthologs on the HSA17 draft sequences revealed differences in the spacing and order of genes. Over 80% of successful amplicons contained SNPs, indicating that this is an efficient approach to generating EST-associated genetic markers. We have demonstrated the feasibility of constructing a linkage map based on SNPs associated with ESTs and the plausibility of utilizing EST, comparative mapping information, and human sequence data to target regions of the bovine genome for SNP marker development.

Animals↗

Selection and use of SNP markers for animal identification and paternity analysis in U.S. beef cattle.

DNA marker technology represents a promising means for determining the genetic identity and kinship of an animal. Compared with other types of DNA markers, single nucleotide polymorphisms (SNPs) are attractive because they are abundant, genetically stable, and amenable to high-throughput automated analysis. In cattle, the challenge has been to identify a minimal set of SNPs with sufficient power for use in a variety of popular breeds and crossbred populations. This report describes a set of 32 highly informative SNP markers distributed among 18 autosomes and both sex chromosomes. Informativity of these SNPs in U.S. beef cattle populations was estimated from the distribution of allele and genotype frequencies in two panels: one consisting of 96 purebred sires representing 17 popular breeds, and another with 154 purebred American Angus from six herds in four Midwestern states. Based on frequency data from these panels, the estimated probability that two randomly selected, unrelated individuals will possess identical genotypes for all 32 loci was 2.0 x 10(-13) for multi-breed composite populations and 1.9 x 10(-10) for purebred Angus populations. The probability that a randomly chosen candidate sire will be excluded from paternity was estimated to be 99.9% and 99.4% for the same respective populations. The DNA immediately surrounding the 32 target SNPs was sequenced in the 96 sires of the multi-breed panel and found to contain an additional 183 polymorphic sites. Knowledge of these additional sites, together with the 32 target SNPs, allows the design of robust, accurate genotype assays on a variety of high-throughput SNP genotyping platforms.

Animals↗

Association of bovine neonatal Fc receptor alpha-chain gene (FCGRT) haplotypes with serum IgG concentration in newborn calves.

This report describes allelic variation in FCGRT (which encodes the alpha-chain of FcRn) and its association with variation of IgG concentration in neonatal calves. Five SNPs were identified by sequencing 1305 bp of FCGRT genomic DNA from a multi-breed panel of 96 cattle and 27 founders of a reference population. These SNPs defined five FCGRT haplotypes that were verified by segregation and used to test association of FCGRT with neonatal IgG concentration in a case-control study. This study established that dams with FCGRT haplotype 3 had a significantly greater risk of failure of passive transfer in their calves (odds ratio [OR] = 3.80, CI95% 1.10-13.18, p = 0.035). Calves with FCGRT haplotype 2 were less likely to have high levels of passively acquired immunoglobulin (OR = 0.18, CI95% 0.05-0.68, p = 0.011). These results indicate that the bovine FCGRT haplotype markers are in linkage disequilibrium with genetic risk factors affecting passive transfer of IgG in beef cattle, an important determinant of neonatal calf morbidity and mortality.

Animals↗

Identification of the single base change causing the callipyge muscle hypertrophy phenotype, the only known example of polar overdominance in mammals.

A small genetic region near the telomere of ovine chromosome 18 was previously shown to carry the mutation causing the callipyge muscle hypertrophy phenotype in sheep. Expression of this phenotype is the only known case in mammals of paternal polar overdominance gene action. A region surrounding two positional candidate genes was sequenced in animals of known genotype. Mutation detection focused on an inbred ram of callipyge phenotype postulated to have inherited chromosome segments identical-by-descent with exception of the mutated position. In support of this hypothesis, this inbred ram was homozygous over 210 Kb of sequence, except for a single heterozygous base position. This single polymorphism was genotyped in multiple families segregating the callipyge locus (CLPG), providing 100% concordance with animals of known CLPG genotype, and was unique to descendants of the founder animal. The mutation lies in a region of high homology among mouse, sheep, cattle, and humans, but not in any previously identified expressed transcript. A substantial open reading frame exists in the sheep sequence surrounding the mutation, although this frame is not conserved among species. Initial functional analysis indicates sequence encompassing the mutation is part of a novel transcript expressed in sheep fetal muscle we have named CLPG1.

Animals↗