PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Shotgun sample sequence comparisons between mouse and human genomes.

A mixed 'clone-by-clone' and 'whole-genome shotgun' strategy will be used to determine the genomic sequence of the mouse. This method will allow a phase of rapid annotation of the contemporaneous human sequence draft, through whole-genome 'sample sequence comparisons'.

Animals↗

Protein chips based on recombinant antibody fragments: a highly sensitive approach as detected by mass spectrometry.

With the human genome in a first sequence draft and several other genomes being finished this year, the existing information gap between genomics and proteomics is becoming increasingly evident. The analysis of the proteome is, however, much more complicated because the synthesis and structural requirements of functional proteins are different from the easily handled oligonucleotides, for which a first analytical breakthrough already has come in the use of DNA chips. In comparison with the DNA microarrays, the protein arrays, or protein chips, offer the distinct possibility of developing a rapid global analysis of the entire proteome. Thus, the concept of comparing proteomic maps of healthy and diseased cells may allow us to understand cell signaling and metabolic pathways and will form a novel base for pharmaceutical companies to develop future therapeutics much more rapidly. This report demonstrates the possibilities of designing protein chips based on specially constructed, small recombinant antibody fragments using nano-structure surfaces with biocompatible characteristics, resulting in sensitive detection in the 600-amol range. The assay readout allows the determination of single or multiple antigen-antibody interactions. Mass identity of the antigens, currently with a resolution of 8000, enables the detection of structural modifications of single proteins.

Antibodies↗

A bigger mouse? The rat genome unveiled.

Rattus norvegicus is an important experimental organism and interesting to evolutionary biologists. The recently published draft rat genome sequence provides us with insights into both the rat's evolution and its physiology. We learn more about genome evolution and, in particular, the adaptive significance of gene family expansions and the evolution of rodent genomes, which appears to have decelerated since the divergence of mouse and rat. An important observation is that some regions of genomes, many in noncoding regions, show very high sequence conservation, while others show unexpectedly fast evolution. Both of these may be pointers to functional significance.

Animals↗

Capillary electrophoresis-based single strand DNA conformation analysis in high-throughput mutation screening.

The generation of the draft human genome sequence has created new possibilities for diagnosis, prevention, and treatment of human disease. One consequence of these new possibilities is an increasing need for methods and technology that can be used for high-throughput screening for mutations in large DNA sample materials. In recent years, a number of mutation screening methods have emerged that are based on the analysis of sequence-dependent changes in the conformation of single- and double-stranded DNA using capillary electrophoresis. Common features of these methods are high sensitivity and reproducibility as well as the possibility for automation and massive parallelization. Thus, at present they are among the most attractive technologies for high-throughput mutation screening. This review describes the recent advances in capillary electrophoresis-based single strand conformation polymorphism (CE-SSCP) for detection of unknown mutations, and assesses its practical usability for high-throughput mutation screening based on the available literature. In addition, future prospects are outlined in light of the recent advances in microchip-based capillary electrophoresis.

DNA Mutational Analysis↗

Functional subgenomics of Clostridium thermocellum cellulosomal genes: identification of the major catalytic components in the extracellular complex and detection of three new enzymes.

Clostridium thermocellum produces the most efficient enzyme-complex for the degradation of polysaccharides in biomass, the large extracellular cellulosome. The draft complete genomic sequence of Clostridium thermocellum was screened for open reading frames (ORF) containing cellulosomal dockerin sequences. Seventy-one putative cellulosomal genes were detected. One third of these ORFs may be involved in cellulose hydrolysis. Most of the others showed homology to hemicellulases, pectinases, chitinases, glycosidases or esterases potentially involved in the unwrapping of cellulose fibers. To identify the predominant catalytic components, cellulosomes were purified and the components were separated by an adapted two-dimensional gel electrophoresis technique. The apparent major spots were identified by MALDI-TOF/TOF. Ten of the components were previously known: the structural protein CipA, the endo-glucanases Cel8A, Cel5G, Cel9N, the cellobiohydrolases Cbh9A, Cel9K, Cel48S, the xylanases Xyn10C, Xyn10Z, and the chitinase Chi18A. In addition, three hitherto unknown major components were detected, Cel9R, Xyn10D and Xgh74A. These major components in the cellulosomal particles most probably constitute the essential enzymes for crystalline cellulose hydrolysis.

Catalysis↗

Rapid turnover and species-specificity of vomeronasal pheromone receptor genes in mice and rats.

Pheromones are used by individuals of the same species to elicit behavioral or physiological changes, and they are perceived primarily by the vomeronasal organ (VNO) in terrestrial vertebrates. VNO pheromone receptors are encoded by the V1r and V2r gene superfamilies in mammals. A comparison of the V1r and V2r repertoires between closely related species can provide significant insights into the evolutionary genetic mechanisms responsible for species-specific pheromone communications. A total of 137 putatively functional V1r genes of 12 families were previously identified from the mouse genome. We report the identification of 95 putatively functional V1r genes from the draft rat genome sequence. These genes map primarily to four blocks in two chromosomes. The rat V1r genes can be phylogenetically grouped into 10 families, which are shared with mouse, and 2 new families, which are rat-specific. Even in many shared families, gene numbers differ between the two species, apparently due to frequent gene duplication and pseudogenization after the separation of the two species. Molecular dating suggests that most of the rat V1r families emerged before or during the radiation of mammalian orders, but many duplications within families occurred as recently as in the past 10 million years (MY). Our results show that the evolution of the V1r repertoire is characterized by exceptionally fast gene turnover via gains and losses of individual genes, suggesting rapid and substantial changes in pheromone communication between species.

Animals↗

Cancer genomics.

The draft human genome sequence and the dissemination of high throughput technology provides opportunities for systematic analysis of cancer cells. Genome-wide mutation screens, high resolution analysis of chromosomal abberations and expression profiling all give comprehensive views of genetic alterations in cancer cells. From these analyses will come a complete list of the genetic changes that drive malignant transformation and of the therapeutic targets that may be exploited for clinical benefit.

Cell Transformation, Neoplastic↗

Retroviral insertion sites and cancer: fountain of all knowledge?

Retroviral gene tagging is enjoying a renaissance as a gene discovery method since the completion of the draft mouse genome sequence. The potential of this approach to elucidate the genetic basis of cancer is reviewed in the light of a series of recent papers that report the results of high-throughput screens.

Animals↗

Abundance of plastid DNA insertions in nuclear genomes of rice and Arabidopsis.

Pairwise comparison of whole plastid and draft nuclear genomic sequences of Arabidopsis thaliana and Oryza sativa L. ssp. indica shows that rice nuclear genomic sequences contain homologs of plastid DNA covering about 94 kb (83%) of plastid genome and including one or more full-length intact (without mutations resulting in premature stop codons) homologues of 26 known protein-coding (KPC) plastid genes. By contrast, only about 20 kb (16%) of chloroplast DNA, including a single intact plastid-derived KPC gene, is presented in the nucleus of A. thaliana. Sixteen rice plastid genes have at least one nuclear copy without any mutation or with only synonymous substitutions. Nuclear copies for other ten plastid genes contain both synonymous and non-synonymous substitutions. Multiple ESTs for 25 out of 26 KPC genes were also found, as well as putative promoters for some of them. The study of substitutions pattern shows that some of nuclear homologues of plastid genes may be functional and/or are under the pressure of the positive natural selection. The similar comparative analysis performed on rice chromosome 1 revealed 27 contigs containing plastid-derived sequences, totalling about 84 kb and covering two thirds of chloroplast DNA, with the intact nuclear copies of 26 different KPC genes. One of these contigs, AP003280, includes almost 57 kb (45%) of chloroplast genome with the intact copies of 22 KPC genes. At the same time, we observed that relative locations of homologues in plastid DNA and the nuclear genome are significantly different.

Arabidopsis↗

Mining the draft human genome.

Now that the draft human genome sequence is available, everyone wants to be able to use it. However, we have perhaps become complacent about our ability to turn new genomes into lists of genes. The higher volume of data associated with a larger genome is accompanied by a much greater increase in complexity. We need to appreciate both the scale of the challenge of vertebrate genome analysis and the limitations of current gene prediction methods and understanding.

Animals↗

Extensive genomic duplication during early chordate evolution.

Opinions on the hypothesis that ancient genome duplications contributed to the vertebrate genome range from strong skepticism to strong credence. Previous studies concentrated on small numbers of gene families or chromosomal regions that might not have been representative of the whole genome, or used subjective methods to identify paralogous genes and regions. Here we report a systematic and objective analysis of the draft human genome sequence to identify paralogous chromosomal regions (paralogons) formed during chordate evolution and to estimate the ages of duplicate genes. We found that the human genome contains many more paralogons than would be expected by chance. Molecular clock analysis of all protein families in humans that have orthologs in the fly and nematode indicated that a burst of gene duplication activity took place in the period 350 650 Myr ago and that many of the duplicate genes formed at this time are located within paralogons. Our results support the contention that many of the gene families in vertebrates were formed or expanded by large-scale DNA duplications in an early chordate. Considering the incompleteness of the sequence data and the antiquity of the event, the results are compatible with at least one round of polyploidy.

Animals↗

Efficient discovery of single-nucleotide polymorphisms in coding regions of human genes.

Single nucleotide polymorphisms in protein coding regions (cSNPs) are of great interest for their effects on phenotype and potential for mapping disease genes. We have identified 5,400 novel exonic SNPs from alignments of public EST data to the draft human genome sequence, and approximately 12,000 more novel exonic SNPs from EST cluster alignments. We found 82% of the genomic-aligned SNPs and 63% of the EST-only SNPs to be detectably polymorphic in 20 Finnish DNA samples. 37% of the SNPs mapped to known protein coding regions, yielding 6,500 distinct, novel cSNPs from the two datasets. These data reveal selection against mutations that alter protein structure, and distinct classes of genes under strongly positive vs. negative pressure from natural selection for amino acid replacement (detected by K(A)/K(S)ratio). We have searched these cSNPs for compatibility with the amino acid profile at each site and structural impact on protein core stability.

Chromosome Mapping↗

The genetics of type 2 diabetes.

Type 2 diabetes mellitus is not a single disease but a genetically heterogeneous group of metabolic disorders sharing glucose intolerance. The precise underlying biochemical defects are unknown and almost certainly include impairments of both insulin secretion and action. The rapidly increasing prevalence of T2D world wide makes it a major cause of morbidity and mortality. Understanding the genetic aetiology of T2D will facilitate its diagnosis, treatment and prevention. The results of linkage and association studies to date demonstrate that, as with other common diseases, multiple genes are involved in the susceptibility to T2D, each making a modest contribution to the overall risk. The completion of the draft human genome sequence and a brace of novel tools for genomic analysis promise to accelerate progress towards a more complete molecular description of T2D.

Animals↗

Comprehensive proteomic analysis of breast cancer cell membranes reveals unique proteins with potential roles in clinical cancer.

Proteins associated with cancer cell plasma membranes are rich in known drug and antibody targets as well as other proteins known to play key roles in the abnormal signal transduction processes required for carcinogenesis. We describe here a proteomics process that comprehensively annotates the protein content of breast tumor cell membranes and defines the clinical relevance of such proteins. Tumor-derived cell lines were used to ensure an enrichment for cancer cell-specific plasma membrane proteins because it is difficult to purify cancer cells and then obtain good membrane preparations from clinical material. Multiple cell lines with different molecular pathologies were used to represent the clinical heterogeneity of breast cancer. Peptide tandem mass spectra were searched against a comprehensive data base containing known and conceptual proteins derived from many public data bases including the draft human genome sequences. This plasma membrane-enriched proteome analysis created a data base of more than 500 breast cancer cell line proteins, 27% of which were of unknown function. The value of our approach is demonstrated by further detailed analyses of three previously uncharacterized proteins whose clinical relevance has been defined by their unique cancer expression profiles and the identification of protein-binding partners that elucidate potential functionality in cancer.

Amino Acid Sequence↗

The Gene Resource Locator: gene locus maps for transcriptome analysis.

Since the advent of the draft human genome sequence there has been growing interest in transcriptome analysis based on genomic data. The Gene Resource Locator (GRL) assembles gene maps that include information on gene-expression patterns, cis-elements in regulatory regions and alternatively spliced transcripts. The database was constructed using customized software, and currently contains 2.2 million alignments (exon-intron structures). The alignments have been annotated and integrated into a system that encompasses approximately 90 000 EST loci sharing common exons, 8091 alternatively spliced transcript groups, 10 801 expression-profile groups, 8066 candidate regulatory regions in full-length cDNAs, and 1 million SNP loci. We have used Flash technology to build a dynamic web viewer that facilitates browsing through the millions of alignments. All of the information is available through the World Wide Web at the Gene Resource Locator web site (http://grl.gi.k.u-tokyo.ac.jp).

Alternative Splicing↗

The human transcriptome map reveals extremes in gene density, intron length, GC content, and repeat pattern for domains of highly and weakly expressed genes.

The chromosomal gene expression profiles established by the Human Transcriptome Map (HTM) revealed a clustering of highly expressed genes in about 30 domains, called ridges. To physically characterize ridges, we constructed a new HTM based on the draft human genome sequence (HTMseq). Expression of 25,003 genes can be analyzed online in a multitude of tissues (http://bioinfo.amc.uva.nl/HTMseq). Ridges are found to be very gene-dense domains with a high GC content, a high SINE repeat density, and a low LINE repeat density. Genes in ridges have significantly shorter introns than genes outside of ridges. The HTMseq also identifies a significant clustering of weakly expressed genes in domains with fully opposite characteristics (antiridges). Both types of domains are open to tissue-specific expression regulation, but the maximal expression levels in ridges are considerably higher than in antiridges. Ridges are therefore an integral part of a higher order structure in the genome related to transcriptional regulation.

Base Composition↗

Retroelement distributions in the human genome: variations associated with age and proximity to genes.

Remnants of more than 3 million transposable elements, primarily retroelements, comprise nearly half of the human genome and have generated much speculation concerning their evolutionary significance. We have exploited the draft human genome sequence to examine the distributions of retroelements on a genome-wide scale. Here we show that genomic densities of 10 major classes of human retroelements are distributed differently with respect to surrounding GC content and also show that the oldest elements are preferentially found in regions of lower GC compared with their younger relatives. In addition, we determined whether retroelement densities with respect to genes could be accurately predicted based on surrounding GC content or if genes exert independent effects on the density distributions. This analysis revealed that all classes of long terminal repeat (LTR) retroelements and L1 elements, particularly those in the same orientation as the nearest gene, are significantly underrepresented within genes and older LTR elements are also underrepresented in regions within 5 kb of genes. Thus, LTR elements have been excluded from gene regions, likely because of their potential to affect gene transcription. In contrast, the density of Alu sequences in the proximity of genes is significantly greater than that predicted based on the surrounding GC content. Furthermore, we show that the previously described density shift of Alu repeats with age to domains of higher GC was markedly delayed on the Y chromosome, suggesting that recombination between chromosome pairs greatly facilitates genomic redistributions of retroelements. These findings suggest that retroelements can be removed from the genome, possibly through recombination resulting in re-creation of insert-free alleles. Such a process may provide an explanation for the shifting distributions of retroelements with time.

Alu Elements↗