PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Identification and characterization of human DAPPER1 and DAPPER2 genes in silico.

WNT signals play key roles in carcinogenesis and embryogenesis through the specification of cell fate and polarity. Dishevelled (DVL) proteins are WNT signaling molecules implicated in beta-catenin pathway and PCP pathway. Xenopus Dapper and Frodo are Dvl-binding proteins, showing 89.8% total-amino-acid identity. Here, we identified and characterized human homologs of Xenopus Dapper and Frodo using bioinformatics. Human DAPPER1 gene was located within human genome draft sequence NT_025892.9 (nucleotide position 39378960-39387891 in the forward orientation), and human DAPPER2 gene within NT_007302.10 (nucleotide position 660279-672480 in the reverse orientation). DAPPER1 (799-amino-acids) and DAPPER2 (774-amino-acids) showed 28.8% total-amino-acid identity. Seven DAPPER homologous (DAPH) domains, including DAPH2 (leucine zipper), DAPH3 (serine rich) and DAPH7 (PDZ binding), were conserved between DAPPER1 and DAPPER2. Phylogenetic analysis of vertebrate Dapper proteins revealed that Xenopus Dapper and Frodo are orthologs of human DAPPER1. DAPPER1 mRNA was expressed in amnion, fetal brain, eye, heart, adult brain medulla, gastric cancer (signet ring cell features), RER+ colon tumor, acute lymphoblastic leukemia, germ cell tumor, chondrosarcoma, and parathyroid tumor. DAPPER2 mRNA was expressed in placenta, genitourinary tract tumor, and endometrial adenocarcinoma. DAPPER1 and DAPPER2 genes were mapped to human chromosome 14q22.3 and 6q27, respectively. Human chromosome 14q22.3 is deleted in astrocytoma, while human chromosome 6q27 is deleted in breast, ovarian, and gastric cancer. Based on evolutionary and functional conservation of WNT signaling molecules as well as human chromosomal localization, DAPPER1 and DAPPER2 genes are predicted to be potent cancer-associated genes.

Adaptor Proteins, Signal Transducing↗

[G-protein coupled peptide receptors and their ligands in human genome].

The G-protein coupled peptide receptors as well as their ligands, endogenous peptides, are involved in regulation of many important physiological processes in the organism and therefore represent attractive targets for pharmaceutical investigation and drug design. With the completion of the human draft genome sequencing, it has become possible to take a comprehensive picture of all genes encoding both peptide receptors and peptides themselves. In the present study a first attempt has been made to carry out a comprehensive analysis of G-protein coupled peptide receptors and their respective endogenous peptide ligands in the human genome. We searched the genome sequence by means of sequential application of standard bioinformatical methods (such as homology search, hierarchical cluster analysis, building of hmm-profiles etc.) with the goal of identifying all the components of peptide ligand/receptor system in the human genome. As a result of this search it was concluded that the probable number of functional peptide receptors in the human genome is 218, and the probable peptide precursors' number is 126 amino acid sequences. These two groups include, respectively, 12 novel G-protein coupled peptide receptors and 10 novel peptide precursors, discovered in the present study. The probable biological functions of newly discovered candidates were determined based on the sequence similarity to the earlier known proteins. Classification of all peptide GPCRs and their ligands based on the ligand specificity was performed for all probable G-protein coupled peptide receptors. The issue of ligand-receptor specificity in the human genome is also discussed.

Amino Acid Sequence↗

A draft sequence of the rice genome (Oryza sativa L. ssp. indica).

We have produced a draft sequence of the rice genome for the most widely cultivated subspecies in China, Oryza sativa L. ssp. indica, by whole-genome shotgun sequencing. The genome was 466 megabases in size, with an estimated 46,022 to 55,615 genes. Functional coverage in the assembled sequences was 92.0%. About 42.2% of the genome was in exact 20-nucleotide oligomer repeats, and most of the transposons were in the intergenic regions between genes. Although 80.6% of predicted Arabidopsis thaliana genes had a homolog in rice, only 49.4% of predicted rice genes had a homolog in A. thaliana. The large proportion of rice genes with no recognizable homologs is due to a gradient in the GC content of rice coding sequences.

Arabidopsis↗

Integration of cytogenetic landmarks into the draft sequence of the human genome.

We have placed 7,600 cytogenetically defined landmarks on the draft sequence of the human genome to help with the characterization of genes altered by gross chromosomal aberrations that cause human disease. The landmarks are large-insert clones mapped to chromosome bands by fluorescence in situ hybridization. Each clone contains a sequence tag that is positioned on the genomic sequence. This genome-wide set of sequence-anchored clones allows structural and functional analyses of the genome. This resource represents the first comprehensive integration of cytogenetic, radiation hybrid, linkage and sequence maps of the human genome; provides an independent validation of the sequence map and framework for contig order and orientation; surveys the genome for large-scale duplications, which are likely to require special attention during sequence assembly; and allows a stringent assessment of sequence differences between the dark and light bands of chromosomes. It also provides insight into large-scale chromatin structure and the evolution of chromosomes and gene families and will accelerate our understanding of the molecular bases of human disease and cancer.

Chromosome Aberrations↗

Genes for intermediate filament proteins and the draft sequence of the human genome: novel keratin genes and a surprisingly high number of pseudogenes related to keratin genes 8 and 18.

We screened the draft sequence of the human genome for genes that encode intermediate filament (IF) proteins in general, and keratins in particular. The draft covers nearly all previously established IF genes including the recent cDNA and gene additions, such as pancreatic keratin 23, synemin and the novel muscle protein syncoilin. In the draft, seven novel type II keratins were identified, presumably expressed in the hair follicle/epidermal appendages. In summary, 65 IF genes were detected, placing IF among the 100 largest gene families in humans. All functional keratin genes map to the two known keratin clusters on chromosomes 12 (type II plus keratin 18) and 17 (type I), whereas other IF genes are not clustered. Of the 208 keratin-related DNA sequences, only 49 reflect true keratin genes, whereas the majority describe inactive gene fragments and processed pseudogenes. Surprisingly, nearly 90% of these inactive genes relate specifically to the genes of keratins 8 and 18. Other keratin genes, as well as those that encode non-keratin IF proteins, lack either gene fragments/pseudogenes or have only a few derivatives. As parasitic derivatives of mature mRNAs, the processed pseudogenes of keratins 8 and 18 have invaded most chromosomes, often at several positions. We describe the limits of our analysis and discuss the striking unevenness of pseudogene derivation in the IF multigene family. Finally, we propose to extend the nomenclature of Moll and colleagues to any novel keratin.

Amino Acid Sequence↗

Sequence variation and disease in the wake of the draft human genome.

The sequencing phase of the human genome project will soon be over. In its wake, repertoires of sequence polymorphisms among the human population are being sampled and a battery of functional genomics projects, from gene and protein expression studies to whole proteome interaction experiments, are generating vast quantities of data. Now that the data, or the means to generate data, are available it is the application of this information in enhancing our understanding of biology that represents the next formidable challenge. Two prominent issues should be considered. First, existing data must be analysed using the best methods available. The prediction of enzymatic activity for bestrophin, whose gene is mutated in Best macular dystrophy, is described in this review. This is an example of the experimentally testable hypotheses that can result from such detailed and exhaustive analyses. Secondly, the torrents of data from high-throughput studies will need to be made more accessible to all using web-based resources that integrate and digest complementary data types. The internet sites that showcase the human genome sequence are blazing a new trail. Ultimately, the success of genome sequencing and functional genomics will be measured not by the quantity and accuracy of raw data generated, but how rapidly they can be harnessed to span the divide between genotype and phenotype.

Alleles↗

Integration of the cytogenetic and physical maps of chicken chromosome 17.

The chicken genome, like those of most avian species, contains numerous microchromosomes that cannot be distinguished by size alone. Unique properties attributed to the microchromosomes include high GC content and gene density, and an enhanced recombination rate. Previously, microchromosome GGA 17 was shown to align with the consensus genetic linkage group E41W17, and bacterial artificial chromosome (BAC) clones containing E41W17 markers were isolated and assigned on the physical BAC map as well as the recently assembled draft chicken genome sequence. For this study, these same BACS were utilized as probes for fluorescence in-situ hybridization (FISH) to develop the GGA 17 cytogenetic map. Here we detail the chromosome order of ten BAC DNAs, thereby deriving a cytogenetic map of GGA 17 that is simultaneously integrated with both the linkage map and genome sequence. The location of the FISH probes together with the morphological appearance of the chromosome suggested that GGA 17 is an acrocentric chromosome whose cytogenetic map orientation is reversed from that currently indicated by the linkage map and draft genome sequence. The reversed orientation and the centromere location of GGA 17 were confirmed experimentally by dual-colour FISH hybridization using terminal BACs and the centromere-specific CNM oligonucleotide as probes. An advantage of this cyto-genomic approach is the improved alignment of the sequence and linkage maps with cytogenetic features such as the centromere, telomeres, p and q arms, and staining patterns indicating GC versus AT content.

Animals↗

Genome-wide detection of alternative splicing in expressed sequences of human genes.

We have identified 6201 alternative splice relationships in human genes, through a genome-wide analysis of expressed sequence tags (ESTs). Starting with approximately 2.1 million human mRNA and EST sequences, we mapped expressed sequences onto the draft human genome sequence and only accepted splices that obeyed the standard splice site consensus. A large fraction (47%) of these were observed multiple times, indicating that they comprise a substantial fraction of the mRNA species. The vast majority of the detected alternative forms appear to be novel, and produce highly specific, biologically meaningful control of function in both known and novel human genes, e.g. specific removal of the lysosomal targeting signal from HLA-DM beta chain, replacement of the C-terminal transmembrane domain and cytoplasmic tail in an FC receptor beta chain homolog with a different transmembrane domain and cytoplasmic tail, likely modulating its signal transduction activity. Our data indicate that a large proportion of human genes, probably 42% or more, are alternatively spliced, but that this appears to be observed mainly in certain types of molecules (e.g. cell surface receptors) and systemic functions, particularly the immune system and nervous system. These results provide a comprehensive dataset for understanding the role of alternative splicing in the human genome, accessible at http://www.bioinformatics.ucla.edu/HASDB.

Alternative Splicing↗

Finishing the euchromatic sequence of the human genome.

The sequence of the human genome encodes the genetic instructions for human physiology, as well as rich information about human evolution. In 2001, the International Human Genome Sequencing Consortium reported a draft sequence of the euchromatic portion of the human genome. Since then, the international collaboration has worked to convert this draft into a genome sequence with high accuracy and nearly complete coverage. Here, we report the result of this finishing process. The current genome sequence (Build 35) contains 2.85 billion nucleotides interrupted by only 341 gaps. It covers approximately 99% of the euchromatic genome and is accurate to an error rate of approximately 1 event per 100,000 bases. Many of the remaining euchromatic gaps are associated with segmental duplications and will require focused work with new methods. The near-complete sequence, the first for a vertebrate, greatly improves the precision of biological analyses of the human genome including studies of gene number, birth and death. Notably, the human genome seems to encode only 20,000-25,000 protein-coding genes. The genome sequence reported here should serve as a firm foundation for biomedical research in the decades ahead.

Amino Acid Sequence↗

An intermediate grade of finished genomic sequence suitable for comparative analyses.

Although the cost of generating draft-quality genomic sequence continues to decline, refining that sequence by the process of "sequence finishing" remains expensive. Near-perfect finished sequence is an appropriate goal for the human genome and a small set of reference genomes; however, such a high-quality product cannot be cost-justified for large numbers of additional genomes, at least for the foreseeable future. Here we describe the generation and quality of an intermediate grade of finished genomic sequence (termed comparative-grade finished sequence), which is tailored for use in multispecies sequence comparisons. Our analyses indicate that this sequence is very high quality (with the residual gaps and errors mostly falling within repetitive elements) and reflects 99% of the total sequence. Importantly, comparative-grade sequence finishing requires approximately 40-fold less reagents and approximately 10-fold less personnel effort compared to the generation of near-perfect finished sequence, such as that produced for the human genome. Although applied here to finishing sequence derived from individual bacterial artificial chromosome (BAC) clones, one could envision establishing routines for refining sequences emanating from whole-genome shotgun sequencing projects to a similar quality level. Our experience to date demonstrates that comparative-grade sequence finishing represents a practical and affordable option for sequence refinement en route to comparative analyses.

Animals↗

Discovery of human inversion polymorphisms by comparative analysis of human and chimpanzee DNA sequence assemblies.

With a draft genome-sequence assembly for the chimpanzee available, it is now possible to perform genome-wide analyses to identify, at a submicroscopic level, structural rearrangements that have occurred between chimpanzees and humans. The goal of this study was to investigate chromosomal regions that are inverted between the chimpanzee and human genomes. Using the net alignments for the builds of the human and chimpanzee genome assemblies, we identified a total of 1,576 putative regions of inverted orientation, covering more than 154 mega-bases of DNA. The DNA segments are distributed throughout the genome and range from 23 base pairs to 62 mega-bases in length. For the 66 inversions more than 25 kilobases (kb) in length, 75% were flanked on one or both sides by (often unrelated) segmental duplications. Using PCR and fluorescence in situ hybridization we experimentally validated 23 of 27 (85%) semi-randomly chosen regions; the largest novel inversion confirmed was 4.3 mega-bases at human Chromosome 7p14. Gorilla was used as an out-group to assign ancestral status to the variants. All experimentally validated inversion regions were then assayed against a panel of human samples and three of the 23 (13%) regions were found to be polymorphic in the human genome. These polymorphic inversions include 730 kb (at 7p22), 13 kb (at 7q11), and 1 kb (at 16q24) fragments with a 5%, 30%, and 48% minor allele frequency, respectively. Our results suggest that inversions are an important source of variation in primate genome evolution. The finding of at least three novel inversion polymorphisms in humans indicates this type of structural variation may be a more common feature of our genome than previously realized.

Animals↗

NotI flanking sequences: a tool for gene discovery and verification of the human genome.

A set of 22 551 unique human NotI flanking sequences (16.2 Mb) was generated. More than 40% of the set had regions with significant similarity to known proteins and expressed sequences. The data demonstrate that regions flanking NotI sites are less likely to form nucleosomes efficiently and resemble promoter regions. The draft human genome sequence contained 55.7% of the NotI flanking sequences, Celera's database contained matches to 57.2% of the clones and all public databases (including non-human and previously sequenced NotI flanks) matched 89.2% of the NotI flanking sequences (identity > or =90% over at least 50 bp, data from December 2001). The data suggest that the shotgun sequencing approach used to generate the draft human genome sequence resulted in a bias against cloning and sequencing of NotI flanks. A rough estimation (based primarily on chromosomes 21 and 22) is that the human genome contains 15 000-20 000 NotI sites, of which 6000-9000 are unmethylated in any particular cell. The results of the study suggest that the existing tools for computational determination of CpG islands fail to identify a significant fraction of functional CpG islands, and unmethylated DNA stretches with a high frequency of CpG dinucleotides can be found even in regions with low CG content.

Cell Line, Transformed↗

Utilization of the human genome sequence localizes human papillomavirus type 16 DNA integrated into the TNFAIP2 gene in a fatal cervical cancer from a 39-year-old woman.

PURPOSE: The purpose of our study was to characterize a human papillomavirus (HPV) 16 DNA integration in the genome of a rapidly progressive, lethal cervical cancer in a 39-year-old woman. EXPERIMENTAL DESIGN: An HPV 16 integration site from cervical cancer tissue was cloned and analyzed using Southern blot hybridization, nucleotide sequencing, fluorescence in situ hybridization analysis for chromosomal localization and comparison with the draft human genome sequence. RESULTS: HPV 16 DNA (3826 bp) was integrated into the genome of the tumor sample and contained an intact upstream regulatory region and E6 and E7 open reading frames. Both 5' and 3' viral-cell junction regions contained direct repeat and palindrome sequences. The chromosomal location of the viral integration and cellular deletion was mapped to chromosome 14q32.3 using both a somatic cell hybrid panel and fluorescence in situ hybridization. Search of the draft human genome sequence confirmed the chromosomal location and revealed a disruption of the TNFAIP2 cytokine/retinoic acid-inducible gene. CONCLUSIONS: On the basis of the lack of sequence homology between the viral and cellular site of integration and the structure of the viral-cell junctions, it seems that HPV 16 DNA integrates into the host genome by a mechanism of nonhomologous recombination. We suggest that, taken together, maintenance of E6 and E7 expression, loss of the E2 gene and disruption of the TNFAIP2 gene through viral integration contributed to the rapid progression of cervical cancer in this patient. Availability of the human genome sequence will facilitate identification of cellular genes involved in cervical cancer by high-throughput analysis of viral integration sites.

Adult↗

Engineering a reduced Escherichia coli genome.

Our goal is to construct an improved Escherichia coli to serve both as a better model organism and as a more useful technological tool for genome science. We developed techniques for precise genomic surgery and applied them to deleting the largest K-islands of E. coli, identified by comparative genomics as recent horizontal acquisitions to the genome. They are loaded with cryptic prophages, transposons, damaged genes, and genes of unknown function. Our method leaves no scars or markers behind and can be applied sequentially. Twelve K-islands were successfully deleted, resulting in an 8.1% reduced genome size, a 9.3% reduction of gene count, and elimination of 24 of the 44 transposable elements of E. coli. These are particularly detrimental because they can mutagenize the genome or transpose into clones being propagated for sequencing, as happened in 18 places of the draft human genome sequence. We found no change in the growth rate on minimal medium, confirming the nonessential nature of these islands. This demonstration of feasibility opens the way for constructing a maximally reduced strain, which will provide a clean background for functional genomics studies, a more efficient background for use in biotechnology applications, and a unique tool for studies of genome stability and evolution.

Chromosome Deletion↗

Comparison of human genetic and sequence-based physical maps.

Recombination is the exchange of information between two homologous chromosomes during meiosis. The rate of recombination per nucleotide, which profoundly affects the evolution of chromosomal segments, is calculated by comparing genetic and physical maps. Human physical maps have been constructed using cytogenetics, overlapping DNA clones and radiation hybrids; but the ultimate and by far the most accurate physical map is the actual nucleotide sequence. The completion of the draft human genomic sequence provides us with the best opportunity yet to compare the genetic and physical maps. Here we describe our estimates of female, male and sex-average recombination rates for about 60% of the genome. Recombination rates varied greatly along each chromosome, from 0 to at least 9 centiMorgans per megabase (cM Mb(-1)). Among several sequence and marker parameters tested, only relative marker position along the metacentric chromosomes in males correlated strongly with recombination rate. We identified several chromosomal regions up to 6 Mb in length with particularly low (deserts) or high (jungles) recombination rates. Linkage disequilibrium was much more common and extended for greater distances in the deserts than in the jungles.

Female↗

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗