PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗

Hot L1s account for the bulk of retrotransposition in the human population.

Although LINE-1 (long interspersed nucleotide element-1, L1) retrotransposons comprise 17% of the human genome, an exhaustive search of the December 2001 "freeze" of the haploid human genome working draft sequence (95% complete) yielded only 90 L1s with intact ORFs. We demonstrate that 38 of 86 (44%) L1s are polymorphic as to their presence in human populations. We cloned 82 (91%) of the 90 L1s and found that 40 of the 82 (49%) are active in a cultured cell retrotransposition assay. From these data, we predict that there are 80-100 retrotransposition-competent L1s in an average human being. Remarkably, 84% of assayed retrotransposition capability was present in six highly active L1s (hot L1s). By comparison, four of five full-length L1s involved in recent human insertions had retrotransposition activity comparable to the six hot L1s in the human genome working draft sequence. Thus, our data indicate that most L1 retrotransposition in the human population stems from hot L1s, with the remaining elements playing a lesser role in genome plasticity.

Alleles↗

Isochore structures in the chicken genome.

The availability of the complete chicken genome sequence provides an unprecedented opportunity to study the global genome organization at the sequence level. Delineating compositionally homogeneous G + C domains in DNA sequences can provide much insight into the understanding of the organization and biological functions of the chicken genome. A new segmentation algorithm, which is simple and fast, has been proposed to partition a given genome or DNA sequence into compositionally distinct domains. By applying the new segmentation algorithm to the draft chicken genome sequence, the mosaic organization of the chicken genome can be confirmed at the sequence level. It is shown herein that the chicken genome is also characterized by a mosaic structure of isochores, long DNA segments that are fairly homogeneous in the G + C content. Consequently, 25 isochores longer than 2 Mb (megabases) have been identified in the chicken genome. These isochores have a fairly homogeneous G + C content and often correspond to meaningful biological units. With the aid of the technique of cumulative GC profile, we proposed an intuitive picture to display the distribution of segmentation points. The relationships between G + C content and the distributions of genes (CpG islands, and other genomic elements) were analyzed in a perceivable manner. The cumulative GC profile, equipped with the new segmentation algorithm, would be an appropriate starting point for analyzing the isochore structures of higher eukaryotic genomes.

Algorithms↗

A draft sequence of the rice genome (Oryza sativa L. ssp. japonica).

The genome of the japonica subspecies of rice, an important cereal and model monocot, was sequenced and assembled by whole-genome shotgun sequencing. The assembled sequence covers 93% of the 420-megabase genome. Gene predictions on the assembled sequence suggest that the genome contains 32,000 to 50,000 genes. Homologs of 98% of the known maize, wheat, and barley proteins are found in rice. Synteny and gene homology between rice and the other cereal genomes are extensive, whereas synteny with Arabidopsis is limited. Assignment of candidate rice orthologs to Arabidopsis genes is possible in many cases. The rice genome sequence provides a foundation for the improvement of cereals, our most important crops.

Arabidopsis↗

A predicted microsatellite map of the passerine genome based on chicken-passerine sequence similarity.

Abstract We present a predicted passerine genome map consisting of 196 microsatellite markers distributed across 25 chromosomes. The map was constructed by assigning chromosomal locations based on the sequence similarity between 550 publicly available passerine microsatellites and the draft chicken genome sequence published by the International Chicken Genome Sequencing Consortium. We compared this passerine microsatellite map with a recently published great reed warbler (Acrocephalus arundinaceus) linkage map derived from the segregation of marker alleles in a pedigree of a natural population. Twenty-four microsatellite markers were shared between the two maps, distributed across ten chromosomes. Synteny was maintained between the predicted passerine microsatellite map and the great reed warbler linkage map, confirming the validity and accuracy of our approach. Possible applications of the predicted passerine microsatellite map include genome mapping; quantitative trait locus (QTL) discovery; understanding heterozygosity-fitness correlations; investigating avian karyotype evolution; understanding microsatellite mutation processes; and for identifying loci conserved in multiple species, unlinked loci for use in genotyping sets and sex-linked markers.

Animals↗

Prophage Finder: a prophage loci prediction tool for prokaryotic genome sequences.

Prophage loci often remain under-annotated or even unrecognized in prokaryotic genome sequencing projects. A PHP application, Prophage Finder, has been developed and implemented to predict prophage loci, based upon clusters of phage-related gene products encoded within DNA sequences. This application provides results detailing several facets of these clusters to facilitate rapid prediction and analysis of prophage sequences. Prophage Finder was tested using previously annotated prokaryotic genomic sequences with manually curated prophage loci as benchmarks. Additional analyses from Prophage Finder searches of several draft prokaryotic genome sequences are available through the Web site (http://bioinformatics.uwp.edu/~phage/DOEResults.php) to illustrate the potential of this application.

Bacteria↗

Directed proteomic analysis of the human nucleolus.

BACKGROUND: The nucleolus is a subnuclear organelle containing the ribosomal RNA gene clusters and ribosome biogenesis factors. Recent studies suggest it may also have roles in RNA transport, RNA modification, and cell cycle regulation. Despite over 150 years of research into nucleoli, many aspects of their structure and function remain uncharacterized. RESULTS: We report a proteomic analysis of human nucleoli. Using a combination of mass spectrometry (MS) and sequence database searches, including online analysis of the draft human genome sequence, 271 proteins were identified. Over 30% of the nucleolar proteins were encoded by novel or uncharacterized genes, while the known proteins included several unexpected factors with no previously known nucleolar functions. MS analysis of nucleoli isolated from HeLa cells in which transcription had been inhibited showed that a subset of proteins was enriched. These data highlight the dynamic nature of the nucleolar proteome and show that proteins can either associate with nucleoli transiently or accumulate only under specific metabolic conditions. CONCLUSIONS: This extensive proteomic analysis shows that nucleoli have a surprisingly large protein complexity. The many novel factors and separate classes of proteins identified support the view that the nucleolus may perform additional functions beyond its known role in ribosome subunit biogenesis. The data also show that the protein composition of nucleoli is not static and can alter significantly in response to the metabolic state of the cell.

Amino Acid Sequence↗

Identification of endopeptidase genes from the genomic sequence of Lactobacillus helveticus CNRZ32 and the role of these genes in hydrolysis of model bitter peptides.

Genes encoding three putative endopeptidases were identified from a draft-quality genome sequence of Lactobacillus helveticus CNRZ32 and designated pepO3, pepF, and pepE2. The ability of cell extracts from Escherichia coli DH5alpha derivatives expressing CNRZ32 endopeptidases PepE, PepE2, PepF, PepO, PepO2, and PepO3 to hydrolyze the model bitter peptides, beta-casein (beta-CN) (f193-209) and alpha(S1)-casein (alpha(S1)-CN) (f1-9), under cheese-ripening conditions (pH 5.1, 4% NaCl, and 10 degrees C) was examined. CNRZ32 PepO3 was determined to be a functional paralog of PepO2 and hydrolyzed both peptides, while PepE and PepF had unique specificities towards alpha(S1)-CN (f1-9) and beta-CN (f193-209), respectively. CNRZ32 PepE2 and PepO did not hydrolyze either peptide under these conditions. To demonstrate the utility of these peptidases in cheese, PepE, PepO2, and PepO3 were expressed in Lactococcus lactis, a common cheese starter, using a high-copy vector pTRKH2 and under the control of the pepO3 promoter. Cell extracts of L. lactis derivatives expressing these peptidases were used to hydrolyze beta-CN (f193-209) and alpha(S1)-CN (f1-9) under cheese-ripening conditions in single-peptide reactions, in a defined peptide mix, and in Cheddar cheese serum. Peptides alpha(S1)-CN (f1-9), alpha(S1)-CN (f1-13), and alpha(S1)-CN (f1-16) were identified from Cheddar cheese serum and included in the defined peptide mix. Our results demonstrate that in all systems examined, PepO2 and PepO3 had the highest activity with beta-CN (f193-209) and alpha(S1)-CN (f1-9). Cheese-derived peptides were observed to affect the activity of some of the enzymes examined, underscoring the importance of incorporating such peptides in model systems. These data indicate that L. helveticus CNRZ32 endopeptidases PepO2 and PepO3 are likely to play a key role in this strain's ability to reduce bitterness in cheese.

Amino Acid Sequence↗

Computational inference of homologous gene structures in the human genome.

With the human genome sequence approaching completion, a major challenge is to identify the locations and encoded protein sequences of all human genes. To address this problem we have developed a new gene identification algorithm, GenomeScan, which combines exon-intron and splice signal models with similarity to known protein sequences in an integrated model. Extensive testing shows that GenomeScan can accurately identify the exon-intron structures of genes in finished or draft human genome sequence with a low rate of false-positives. Application of GenomeScan to 2.7 billion bases of human genomic DNA identified at least 20,000-25,000 human genes out of an estimated 30,000-40,000 present in the genome. The results show an accurate and efficient automated approach for identifying genes in higher eukaryotic genomes and provide a first-level annotation of the draft human genome.

Algorithms↗

The DNA sequence, annotation and analysis of human chromosome 3.

After the completion of a draft human genome sequence, the International Human Genome Sequencing Consortium has proceeded to finish and annotate each of the 24 chromosomes comprising the human genome. Here we describe the sequencing and analysis of human chromosome 3, one of the largest human chromosomes. Chromosome 3 comprises just four contigs, one of which currently represents the longest unbroken stretch of finished DNA sequence known so far. The chromosome is remarkable in having the lowest rate of segmental duplication in the genome. It also includes a chemokine receptor gene cluster as well as numerous loci involved in multiple human cancers such as the gene encoding FHIT, which contains the most common constitutive fragile site in the genome, FRA3B. Using genomic sequence from chimpanzee and rhesus macaque, we were able to characterize the breakpoints defining a large pericentric inversion that occurred some time after the split of Homininae from Ponginae, and propose an evolutionary history of the inversion.

Animals↗

EZ-Retrieve: a web-server for batch retrieval of coordinate-specified human DNA sequences and underscoring putative transcription factor-binding sites.

The availability of a draft human genome sequence and ability to monitor the transcription of thousands of genes with DNA microarrays has necessitated the need for new computational tools that can analyze cis-regulatory elements controlling genes that display similar expression patterns. We have developed a tool designated EZ-Retrieve that can: (i) retrieve any particular region of human genome sequence from the NCBI database and (ii) analyze retrieved sequences for putative transcription factor-binding sites (TFBSs) as they appear on the TRANSFAC database. The tool is web-based, user-friendly and offers both batch sequence retrieval and batch TFBS prediction. A major application of EZ-Retrieve is the analysis of co-expressed genes that are highlighted as expression clusters in DNA microarray experiments.

Animals↗

Endogenous retroviruses in the human genome sequence.

The human genome contains many endogenous retroviral sequences, and these have been suggested to play important roles in a number of physiological and pathological processes. Can the draft human genome sequences help us to define the role of these elements more closely?

Autoimmune Diseases↗

Comprehensive search for chicken W chromosome-linked genes expressed in early female embryos from the female-minus-male subtracted cDNA macroarray.

In order to seek chicken W chromosome-linked genes expressed significantly earlier than the time of gonadal differentiation, female-minus-male-subtracted cDNA macroarrays were prepared from day 2 (Hamburger-Hamilton stages 12-13), day 3 (stages 19-20) and day 4 (stages 24-25) embryos. From a total of 15-744 macroarrayed cDNA clones, 610 clones exhibiting significantly female-specific expression were selected. When each one of the 610 cDNA clones was used as a probe in Southern blot hybridization with male or female chicken genomic DNA, 62 clones, grouped into eight (A-H) types according to their patterns of hybridization, were considered to be derived from W chromosome-linked genes. When representative cDNA clones in each type were sequenced, clones derived from two known W-linked genes; SPIN-W and ATP5A1W , and from two hitherto unknown W-linked genes, represented by 2d-2D9 and 2d-2F9 clones, were identified and their localizations on the W chromosome were confirmed by fluorescence in-situ hybridization. The 2d-2D9 sequence has no significant homology with other genes in databases but 2d-2F9 has a region which shows partial homology to the consensus sequence of the AAA ATPase superfamily. Both 2d-2D9 and 2d-2F9 sequences are found in contigs of undetermined chromosome-linkage in the Draft Chicken Genome Sequence.

Animals↗

Comparative physical mapping of targeted regions of the rat genome.

The comparative mapping and sequencing of vertebrate genomes is now a key priority for the Human Genome Project. In addition to finishing the human genome sequence and generating a 'working draft' of the mouse genome sequence, significant attention is rapidly turning to the analysis of other model organisms, such as the laboratory rat (Rattus norvegicus). As a complement to genome-wide mapping and sequencing efforts, it is often important to generate detailed maps and sequence data for specific regions of interest. Using an adaptation of our previously described approach for constructing mouse comparative and physical maps, we have established a general strategy for targeted mapping of the rat genome. Specifically, we constructed a framework comparative map of human Chromosome (Chr) 7 and the orthologous regions of the rat genome, as well as two large (>1-Mb) P1-derived artificial chromosome (PAC)-based physical maps. Generation of these physical maps involved the use of mouse-derived probes that cross-hybridized with rat PAC clones. The first PAC map encompasses the cystic fibrosis transmembrane conductance regulator gene (Cftr), while the second map allows a three-species comparison of a genomic region containing intra- and inter-chromosomal evolutionary rearrangements. The studies reported here further demonstrate that cross-species hybridization between related animals, such as rat and mouse, can be readily used for the targeted construction of clone-based physical maps, thereby accelerating the analysis of biologically interesting regions of vertebrate genomes.

Animals↗

Can sequencing shed light on cell cycling?

Every organism must have cells that can replicate indefinitely. Can the draft human genome sequence tell us how the cell cycle works and how it evolved? We studied two protein families--the cyclins and their partners the cyclin-dependent kinases (Cdks)--and a conserved regulatory circuit, the spindle checkpoint. Disappointingly, we discovered a few novel cyclins and no new Cdks or components of the spindle checkpoint, and could shed little light on the organization of the cell cycle.

Amino Acid Sequence↗

A 7872 cDNA microarray and its use in bovine functional genomics.

The strategy used to create and annotate a 7872 cDNA microarray from cattle placenta and spleen cDNA sequences is described. This microarray contains approximately 6300 unique genes, as determined by BLASTN and TBLASTX similarity search against the human and mouse UniGene and draft human genome sequence databases (build 34). Sequences on the array were annotated with gene ontology (GO) terms, thereby facilitating data analysis and interpretation. A total of 3244 genes were annotated with GO terms. The array is rich in sequences encoding transcription factors, signal transducers and cell cycle regulators. Current research being conducted with this array is described, and an overview of planned improvements in our microarray platform for cattle functional genomics is presented.

Animals↗

Identification of SmtB/ArsR cis elements and proteins in archaea using the Prokaryotic InterGenic Exploration Database (PIGED).

Microbial genome sequencing projects have revealed an apparently wide distribution of SmtB/ArsR metal-responsive transcriptional regulators among prokaryotes. Using a position-dependent weight matrix approach, prokaryotic genome sequences were screened for SmtB/ArsR DNA binding sites using data derived from intergenic sequences upstream of orthologous genes encoding these regulators. Sixty SmtB/ArsR operators linked to metal detoxification genes, including nine among various archaeal species, are predicted among 230 annotated and draft prokaryotic genome sequences. Independent multiple sequence alignments of putative operator sites and corresponding winged helix-turn-helix motifs define sequence signatures for the DNA binding activity of this SmtB/ArsR subfamily. Prediction of an archaeal SmtB/ArsR based upon these signature sequences is confirmed using purified Methanosarcina acetivorans C2A protein and electrophoretic mobility shift assays. Tools used in this study have been incorporated into a web application, the Prokaryotic InterGenic Exploration Database (PIGED; http://bioinformatics.uwp.edu/~PIGED/home.htm), facilitating comparable studies. Use of this tool and establishment of orthology based on DNA binding signatures holds promise for deciphering potential cellular roles of various archaeal winged helix-turn-helix transcriptional regulators.

Archaea↗

A general approach to single-nucleotide polymorphism discovery.

Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence as a template on which to layer often unmapped, fragmentary sequence data and to use base quality values to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.

Algorithms↗