PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Draft versus finished sequence data for DNA and protein diagnostic signature development.

Sequencing pathogen genomes is costly, demanding careful allocation of limited sequencing resources. We built a computational Sequencing Analysis Pipeline (SAP) to guide decisions regarding the amount of genomic sequencing necessary to develop high-quality diagnostic DNA and protein signatures. SAP uses simulations to estimate the number of target genomes and close phylogenetic relatives (near neighbors or NNs) to sequence. We use SAP to assess whether draft data are sufficient or finished sequencing is required using Marburg and variola virus sequences. Simulations indicate that intermediate to high-quality draft with error rates of 10(-3)-10(-5) (approximately 8x coverage) of target organisms is suitable for DNA signature prediction. Low-quality draft with error rates of approximately 1% (3x to 6x coverage) of target isolates is inadequate for DNA signature prediction, although low-quality draft of NNs is sufficient, as long as the target genomes are of high quality. For protein signature prediction, sequencing errors in target genomes substantially reduce the detection of amino acid sequence conservation, even if the draft is of high quality. In summary, high-quality draft of target and low-quality draft of NNs appears to be a cost-effective investment for DNA signature prediction, but may lead to underestimation of predicted protein signatures.

Computational Biology↗

Multiple gene organization of pufferfish Fugu rubripes tropomyosin isoforms and tissue distribution of their transcripts.

The Japanese pufferfish, torafugu (Fugu rubripes), has a haploid genome of about 400 Mb in size, which has been sequenced to approximately 90% coverage. Here we identified six Fugu tropomyosin (TPM) gene sequences by using the BLASTN program and the sequence of the white croaker TPM1 gene in our collection against the draft assembly of the Fugu genomic sequence database. TPM2, TPM3 and TPM4 genes were identified together with a set of two potentially duplicated genes of TPM1 (TPM1-1 and TPM1-2) as described in our previous report and TPM4 (TPM4-1 and TPM4-2) newly found in this study. The expression patterns of these Fugu TPM genes were determined by reverse transcription polymerase chain reaction (RT-PCR). A phylogenetic tree was constructed using the deduced amino acid sequences, which were encoded by the exons common to all vertebrate TPM genes. This indicated that the Fugu TPM1 and TPM4 genes had resulted from a gene duplication in the fish evolutionary lineage.

Alternative Splicing↗

Genome sequence of the Brown Norway rat yields insights into mammalian evolution.

The laboratory rat (Rattus norvegicus) is an indispensable tool in experimental medicine and drug development, having made inestimable contributions to human health. We report here the genome sequence of the Brown Norway (BN) rat strain. The sequence represents a high-quality 'draft' covering over 90% of the genome. The BN rat sequence is the third complete mammalian genome to be deciphered, and three-way comparisons with the human and mouse genomes resolve details of mammalian evolution. This first comprehensive analysis includes genes and proteins and their relation to human disease, repeated sequences, comparative genome-wide studies of mammalian orthologous chromosomal regions and rearrangement breakpoints, reconstruction of ancestral karyotypes and the events leading to existing species, rates of variation, and lineage-specific and lineage-independent evolutionary events such as expansion of gene families, orthology relations and protein evolution.

Animals↗

A first generation physical map of the medaka genome in BACs essential for positional cloning and clone-by-clone based genomic sequencing.

In order to realize the full potential of the medaka as a model system for developmental biology and genetics, characterized genomic resources need to be established, culminating in the sequence of the medaka genome. To facilitate the map-based cloning of genes underlying induced mutations and to provide templates for clone-based genomic sequencing, we have created a first-generation physical map of the medaka genome in bacterial artificial chromosome (BAC) clones. In particular, we exploited the synteny to the closely related genome of the pufferfish, Takifugu rubripes, by marker content mapping. As a first step, we clustered 103,144 public medaka EST sequences to obtain a set of 21,121 non-redundant sequence entities. Avoiding oversampling of gene-dense regions, 11,254 of EST clusters were successfully matched against the draft sequence of the fugu genome, and 2363 genes were selected for the BAC map project. We designed 35mer oligonucleotide probes from the selected genes and hybridized them against 64,500 BAC clones of strains Cab and Hd-rR, representing 14-fold coverage of the medaka genome. Our data set is further supplemented with 437 results generated from PCR-amplified inserts of medaka cDNA clones and BAC end-fragment markers. Our current, edited, first generation medaka BAC map consists of 902 map segments that cover about 74% of the medaka genome. The map contains 2721 markers. Of these, 2534 are from expressed sequences, equivalent to a non-redundant set of 2328 loci. The 934 markers (724 different) are anchored to the medaka genetic map. Thus, genetic map assignments provide immediate access to underlying clones and contigs, simplifying molecular access to candidate gene regions and their characterization.

Animals↗

An exhaustive DNA micro-satellite map of the human genome using high performance computing.

The current pace of the generation of sequence data requires the development of software tools that can rapidly provide full annotation of the data. We have developed a new method for rapid sequence comparison using the exact match algorithm without repeat masking. As a demonstration, we have identified all perfect simple tandem repeats (STR) within the draft sequence of the human genome. The STR elements (chromosome, position, length and repeat subunit) have been placed into a relational database. Repeat flanking sequence is also publicly accessible at http://grid.abcc.ncifcrf.gov. To illustrate the utility of this complete set of STR elements, we documented the increased density of potentially polymorphic markers throughout the genome. The new STR markers may be useful in disease association studies because so many STR elements manifest multiallelic polymorphism. Also, because triplet repeat expansions are important for human disease etiology, we identified trinucleotide repeats that exist within exons of known genes. This resulted in a list that includes all 14 genes known to undergo polynucleotide expansion, and 48 additional candidates. Several of these are non-polyglutamine triplet repeats. Other examinations of the STR database demonstrated repeats spanning splice junctions and identified SNPs within repeat elements.

Alleles↗

Genome-wide analysis of ATP-binding cassette (ABC) proteins in a model legume plant, Lotus japonicus: comparison with Arabidopsis ABC protein family.

ATP-binding cassette (ABC) proteins constitute a large family in plants with more than 120 members each in Arabidopsis and rice, and have various functions including the transport of auxin and alkaloid, as well as the regulation of stomata movement. In this report, we carried out genome-wide analysis of ABC protein genes in a model legume plant, Lotus japonicus. For analysis of the Lotus genome sequence, we devised a new method 'domain-based clustering analysis', where domain structures like the nucleotide-binding domain (NBD) and transmembrane domain (TMD), instead of full-length amino acid sequences, are used to compare phylogenetically each other. This method enabled us to characterize fragments of ABC proteins, which frequently appear in a draft sequence of the Lotus genome. We identified 91 putative ABC proteins in L. japonicus, i.e. 43 'full-size', 40 'half-size' and 18 'soluble' putative ABC proteins. The characteristic feature of the composition is that Lotus has extraordinarily many paralogs similar to AtMRP14 and AtPDR12, which are at least six and five members, respectively. Expression analysis of the latter genes performed with real-time quantitative reverse transcription-PCR revealed their putative involvement in the nodulation process.

ATP-Binding Cassette Transporters↗

Characterization of tissue expression and full-length coding sequence of a novel human gene mapping at 3q12.1 and transcribed in oligodendrocytes.

Macro-array differential hybridization of a collection of 5058 human gene transcripts represented in an IMAGE infant brain cDNA library has led to the identification of transcripts displaying preferential or specific expression in brain (Genome Res. 9 (1999) 195; http://idefix.upr420.vjf.cnrs.fr/IMAGE). Most of these genes correspond to as yet undescribed functions. Detailed characterization of the expression, sequence, and genome assignment of one of these genes named C3orf4, is reported here. The full-length sequence of the transcript was obtained by 5' extension RT-PCR. The gene transcript (2.8 kb) encodes a 253 amino acid long protein, with four transmembrane domains. The position of the C3orf4 gene was determined at 3q12.1 thanks to the draft sequence of the human genome. It is composed of five exons spanning more than 7 kb. No TATAA box but a CpG island was found upstream of the beginning of the gene. Northern blot analysis and in situ hybridization revealed a predominant expression in myelinated structures such as corpus callosum and spinal cord. RT-PCR showed expression of the C3orf4 gene in rat optic nerve and cultured oligodendrocytes, the myelinating cells of the central nervous system, but not in astrocytes. This work supports further investigations aimed at determining the role of the C3orf4 gene in myelinating cells.

Adult↗

The DNA sequence quality machine at IFOM: a simple Web-based tool for quantitative assessment of sequencing reactions.

DNA sequence quality is a factor of paramount importance in the world of modern genetic and genomics. Both the sequencing of Human Genome in the "post-draft" era [NHGRI Standard for quality of Human Genomic Sequences, Rev. 7 July (2002) where http://www.nhgri.nih.gov/Grant_info/Funding/ Statements/RFA/quality_standard.html is the HTTP address] and recent "high-throughput" approaches to genetic investigation such as SAGE [Velculescu, V.E., Zhang, L., Vogelstein, B. et al. (1995) "Serial analysis of gene expression", Science 270, 484-487] need a reliable, standardized measure of the quality of a sequencing reaction. The increasing importance of SNP studies also requires a stronger quality control on sequencing reactions by the final user. We propose here a simple, web-based tool for integrated sequence quality evaluation, high quality region quantitative value calculation and chromatogram display. This software is aimed at the small to medium DNA sequence laboratory or to the single researcher, interested in getting a quantitative measure of the sequence quality, browsing the chromatogram and checking the quality values base by base. The program is freely available from the IFOM bioinformatics web Server at http://bio.ifom-firc.it/Phred20/index.html.

Algorithms↗

Guide to the draft human genome.

There are a number of ways to investigate the structure, function and evolution of the human genome. These include examining the morphology of normal and abnormal chromosomes, constructing maps of genomic landmarks, following the genetic transmission of phenotypes and DNA sequence variations, and characterizing thousands of individual genes. To this list we can now add the elucidation of the genomic DNA sequence, albeit at 'working draft' accuracy. The current challenge is to weave together these disparate types of data to produce the information infrastructure needed to support the next generation of biomedical research. Here we provide an overview of the different sources of information about the human genome and how modern information technology, in particular the internet, allows us to link them together.

Amino Acid Sequence↗

Development of a comprehensive comparative radiation hybrid map of bovine chromosome 7 (BTA7) versus human chromosomes 1 (HSA1), 5 (HSA5) and 19 (HSA19).

In this study, we present a comprehensive 3,000-rad radiation hybrid (RH) map of bovine chromosome 7 (BTA7) with 108 markers including 54 genes or ESTs. For 52 of them, a human ortholog sequence was found either on HSA1 (one gene), HSA5 (31 genes) or HSA19 (19 genes and one non-annotated sequence) confirming previously described syntenies. Moreover, in order to refine boundaries of blocks of conserved synteny, nine new genes were mapped to the bovine genome on the basis of their localization on the human genome: six on BTA7 and originating from HSA1 (TRIM17), HSA5 (MAN2A1, LMNB1, SIAT8D and FLJ1159) and HSA19 (VAV1), and the three others (AP3B1, APC and CCNG1) on BTA10. The available draft of the human genome sequence allowed us to present a detailed picture of the distribution of conserved synteny segments between man and cattle. Finally, the INRA bovine BAC library was screened for most of the BTA7 markers considered in this study to provide anchors for the bovine physical map.

Animals↗

[Genetics of autoimmune diseases].

The publication of the human genome sequence this year was an important milestone in Biology. Autoimmune diseases arise through a complex interaction of genetics and environmental factors. They are influenced by more than one gene and do not exhibit a simple mode of inheritance. Some of the involved genes are thought to play a role in several autoimmune diseases whereas others seem to be specific for one of them. We have studied mainly genes located on the Major Histocompatibility Complex situated on the short arm of the 6th chromosome. This region is the most important genetic determinant for autoimmune diseases. It carries many genes playing a role in immune response and several have been observed to be associated with autoimmune diseases susceptibility. Although what has been published is just a draft of the human genome sequence, and a lot of work is needed to make sense of all these data and move into the study of gene function, the important technology breakthroughs of the last few years in genomics, proteomics and bioinformatics will speedup the study and results may come much earlier than could be imagined some years ago.

Autoimmune Diseases↗

Genome-wide prediction of human VNTRs.

Polymorphic minisatellites, also known as variable number of tandem repeats (VNTRs), are tandem repeat regions that show variation in the number of repeat units among chromosomes in a population. Currently, there are no general methods for predicting which minisatellites have a high probability of being polymorphic, given their sequence characteristics. An earlier approach has focused on potentially highly polymorphic and hypervariable minisatellites, which make up only a small fraction of all minisatellites in the human genome. We have developed a model, based on available minisatellite and VNTR sequence data, that predicts the probability that a minisatellite (unit size > or = 6 bp) identified by the computer program Tandem Repeats Finder is polymorphic (VNTR). According to the model, minisatellites with high copy number and high degree of sequence similarity are most likely to be VNTRs. This approach was used to scan the draft sequence of the human genome for VNTRs. A total of 157,549 minisatellite repeats were found, of which 29,224 are predicted to be VNTRs. Contrary to previous results, VNTRs appear to be widespread and abundant throughout the human genome, with an estimated density of 9.1 VNTRs/Mb.

Genome, Human↗

Functional genomics: lessons from yeast.

Functional genomics represents a systematic approach to elucidating the function of the novel genes revealed by complete genome sequences. Such an approach should adopt a hierarchical strategy since this will both limit the number of experiments to be performed and permit a closer and closer approximation to the function of any individual gene to be achieved. Moreover, hierarchical analyses have, in their early stages, tremendous integrative power and functional genomics aims at a comprehensive and integrative view of the workings of living cells. The first draft of the human genome sequence has just been produced, and the complete genome sequences of a number of eukaryotic human pathogens (including the parasitic protozoa Plasmodium, Leishmania, and Trypanosoma) will soon be available. However, the most rapid progress in the elucidation of gene function will initially be made using model organisms. Yeast is an excellent eukaryotic model and at least 40% of single-gene determinants of human heritable diseases find homologues in yeast. We have adopted a systematic approach to the functional analysis of the Saccharomyces cerevisiae genome. A number of the approaches for the functional analysis of novel yeast genes are discussed. The different approaches are grouped into four domains: genome, transcriptome, proteome, and metabolome. The utility of genetic, biochemical, and physico-chemical methods for the analysis of these domains is discussed, and the importance of framing precise biological questions, when using these comprehensive analytical methods, is emphasized. Finally, the prospects for elucidating the function of protozoan genes by using the methods pioneered with yeast, and even exploiting Saccharomyces itself, as a surrogate, are explored.

Fungal Proteins↗

Protein encoding genes in an ancient plant: analysis of codon usage, retained genes and splice sites in a moss, Physcomitrella patens.

BACKGROUND: The moss Physcomitrella patens is an emerging plant model system due to its high rate of homologous recombination, haploidy, simple body plan, physiological properties as well as phylogenetic position. Available EST data was clustered and assembled, and provided the basis for a genome-wide analysis of protein encoding genes. RESULTS: We have clustered and assembled Physcomitrella patens EST and CDS data in order to represent the transcriptome of this non-seed plant. Clustering of the publicly available data and subsequent prediction resulted in a total of 19,081 non-redundant ORF. Of these putative transcripts, approximately 30% have a homolog in both rice and Arabidopsis transcriptome. More than 130 transcripts are not present in seed plants but can be found in other kingdoms. These potential "retained genes" might have been lost during seed plant evolution. Functional annotation of these genes reveals unequal distribution among taxonomic groups and intriguing putative functions such as cytotoxicity and nucleic acid repair. Whereas introns in the moss are larger on average than in the seed plant Arabidopsis thaliana, position and amount of introns are approximately the same. Contrary to Arabidopsis, where CDS contain on average 44% G/C, in Physcomitrella the average G/C content is 50%. Interestingly, moss orthologs of Arabidopsis genes show a significant drift of codon fraction usage, towards the seed plant. While averaged codon bias is the same in Physcomitrella and Arabidopsis, the distribution pattern is different, with 15% of moss genes being unbiased. Species-specific, sensitive and selective splice site prediction for Physcomitrella has been developed using a dataset of 368 donor and acceptor sites, utilizing a support vector machine. The prediction accuracy is better than those achieved with tools trained on Arabidopsis data. CONCLUSION: Analysis of the moss transcriptome displays differences in gene structure, codon and splice site usage in comparison with the seed plant Arabidopsis. Putative retained genes exhibit possible functions that might explain the peculiar physiological properties of mosses. Both the transcriptome representation (including a BLAST and retrieval service) and splice site prediction have been made available on http://www.cosmoss.org, setting the basis for assembly and annotation of the Physcomitrella genome, of which draft shotgun sequences will become available in 2005.

Alternative Splicing↗

Genome comparisons highlight similarity and diversity within the eukaryotic kingdoms.

In 2000, the number of completely sequenced eukaryotic genomes increased to four. The addition of Drosophila and Arabidopsis into this cohort permits additional insights into the processes that have shaped evolution. Analysis and comparisons of both completed genomes and partially sequenced genomes have already shed light on mechanisms such as gene duplication and gene loss that have long been hypothesized to be major forces in speciation. Indeed, duplicate gene pairs in Saccharomyces, Arabidopsis, Caenorhabditis and Drosophila are high: 30%, 60%, 48% and 40%, respectively. Evidence of horizontal gene-transfer, thought to be a major evolutionary force in bacteria, has been found in Arabidopsis. The release of the 'first draft' of the human genome sequence in 2000 heralds a new stage of biological study. Understanding the as-yet-unannotated human genome will be largely based on conclusions, techniques and tools developed during the analysis and comparison of the genome of these four model organisms.

Animals↗

[Screening of novel genes differentially expressed in human renal cell carcinoma by suppression subtractive hybridization].

BACKGROUND AND OBJECTIVE: Identifying the differentially expressed genes in renal cell carcinoma (RCC) contributes to the elucidation of its genetic basis. However, the above knowledge has not yet been fully understood. The aims of this experiment were to screen novel genes differentially expressed in RCC tissues by suppression subtractive hybridization (SSH) and clone RCC-specific related genes. METHODS: To construct SSH library of RCC by using the mRNA from RCC tissues and matched normal kidney tissues as tester and driver, respectively. Partial positive clones in the library were selected randomly and sequenced, then analyze the sequences with the BLAST software. To confirm the location of the fragments of interest in human chromosome through comparing their sequences with the human genome draft. mRNA levels of the novel genes in RCC and matched normal kidney tissues were determined by Northern blot and semi-quantitative RT-PCR analysis. RESULTS: The SSH library contained 414 positive clones. Random analysis of 280 clones with enzyme restriction showed that 265 clones contained cDNA fragments distributed mainly between 300-900bp. Among 80 arbitrary clones with were derived from above 265 clones and sequenced, No. 28, 158, 170, and 249 clones are previously unknown genes and located in human chromosome 21q22, 4p15.3, 9q34, and 22q11.2 by electronic mapping, respectively. The consequence of semi-quantitative RT-PCR demonstrated that mRNA levels of the two novel genes were overexpressed in RCC compared to matched normal tissues by more than 2-6 folds. Northern blot analysis confirmed the above results. CONCLUSIONS: SSH is a reliable strategy for screening novel genes differentially expressed in RCC. The novel gene fragments can be used to clone their full length and further to study their functions.

Blotting, Northern↗

The human genome: an immuno-centric view of evolutionary strategies.

A hallmark of modern biology is the realization of the fundamental unity of biological processes in all life forms. Consequently, the complete genome sequencing of various bacteria, yeast (Saccharomyces cerevisiae), fly (Drosophila melanogaster) and worm (Caenorhabditis elegans) over the past five years has already had an impact on all of biology. "Model organisms" have contributed a great deal to immunology; for example, the Toll receptors of the fly provided the impetus for the investigation of Toll-like receptors, which proved to be fundamental elements in the mammalian innate immune system. The recent release of a draft sequence of the human genome provides the first panoramic view of the 30000-35000 human genes in the human genetic blueprint and provides a plethora of new details, the significance of which will take some time to appreciate. The over-riding concepts that emerge from these studies relate primarily to general evolutionary processes that are equally as relevant to immunology as they are to other disciplines of biology.

Amino Acid Sequence↗