PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Long-range heterogeneity at the 3' ends of human mRNAs.

The publication of a draft of the human genome and of large collections of transcribed sequences has made it possible to study the complex relationship between the transcriptome and the genome. In the work presented here, we have focused on mapping mRNA 3' ends onto the genome by use of the raw data generated by the expressed sequence tag (EST) sequencing projects. We find that at least half of the human genes encode multiple transcripts whose polyadenylation is driven by multiple signals. The corresponding transcript 3' ends are spread over distances in the kilobase range. This finding has profound implications for our understanding of gene expression regulation and of the diversity of human transcripts, for the design of cDNA microarray probes, and for the interpretation of gene expression profiling experiments.

3' Flanking Region↗

Identification of a long-chain polyunsaturated fatty acid acyl-coenzyme A synthetase from the diatom Thalassiosira pseudonana.

The draft genome of the diatom Thalassiosira pseudonana was searched for DNA sequences showing homology with long-chain acyl-coenzyme A synthetases (LACSs), since the corresponding enzyme may play a key role in the accumulation of health-beneficial polyunsaturated fatty acids (PUFAs) in triacylglycerol. Among the candidate genes identified, an open reading frame named TplacsA was found to be full length and constitutively expressed during cell cultivation. The predicted amino acid sequence of the corresponding protein, TpLACSA, exhibited typical features of acyl-coenzyme A (acyl-CoA) synthetases involved in the activation of long-chain fatty acids. Feeding experiments carried out in yeast (Saccharomyces cerevisiae) transformed with the algal gene showed that TpLACSA was able to activate a number of PUFAs, including eicosapentaenoic acid and docosahexaenoic acid (DHA). Determination of acyl-CoA synthetase activities by direct measurement of acyl-CoAs produced in the presence of different PUFA substrates showed that TpLACSA was most active toward DHA. Heterologous expression also revealed that TplacsA transformants were able to incorporate more DHA in triacylglycerols than the control yeast.

Acyl Coenzyme A↗

The repetitive landscape of the chicken genome.

Cot-based cloning and sequencing (CBCS) is a powerful tool for isolating and characterizing the various repetitive components of any genome, combining the established principles of DNA reassociation kinetics with high-throughput sequencing. CBCS was used to generate sequence libraries representing the high, middle, and low-copy fractions of the chicken genome. Sequencing high-copy DNA of chicken to about 2.7 x coverage of its estimated sequence complexity led to the initial identification of several new repeat families, which were then used for a survey of the newly released first draft of the complete chicken genome. The analysis provided insight into the diversity and biology of known repeat structures such as CR1 and CNM, for which only limited sequence data had previously been available. Cot sequence data also resulted in the identification of four novel repeats (Birddawg, Hitchcock, Kronos, and Soprano), two new subfamilies of CR1 repeats, and many elements absent from the chicken genome assembly. Multiple autonomous elements were found for a novel Mariner-like transposon, Galluhop, in addition to nonautonomous deletion derivatives. Phylogenetic analysis of the high-copy repeats CR1, Galluhop, and Birddawg provided insight into two distinct genome dispersion strategies. This study also exemplifies the power of the CBCS method to create representative databases for the repetitive fractions of genomes for which only limited sequence data is available.

Animals↗

SMASHing regulatory sites in DNA by human-mouse sequence comparisons.

Regulatory sequence elements provide important clues to understanding and predicting gene expression. Although the binding sites for hundreds of transcription factors are known, there has been no systematic attempt to incorporate this information in the annotation of the human genome. Cross species sequence comparisons are critical to a meaningful annotation of regulatory elements since they generally reside in conserved non-coding regions. To take advantage of the recently completed drafts of the mouse and human genomes for annotating transcription factor binding sites, we developed SMASH, a computational pipeline that identifies thousands of orthologous human/ mouse proteins, maps them to genomic sequences, extracts and compares upstream regions and annotates putative regulatory elements in conserved, non-coding, upstream regions. Our current dataset consists of approximately 2,500 human/mouse gene pairs. Transcription start sites were estimated by mapping quasi-full length cDNA sequences. SMASH uses a novel probabilistic method to identify putative conserved binding sites that takes into account the competition between transcription factors for binding DNA. SMASH presents the results via a genome browser web interface which displays the predicted regulatory information together with the current annotations for the human genome. Our results are validated by comparison to previously published experimental data. SMASH results compare favorably to other existing computational approaches.

Algorithms↗

A gene expression map for the euchromatic genome of Drosophila melanogaster.

We used a maskless photolithography method to produce DNA oligonucleotide microarrays with unique probe sequences tiled throughout the genome of Drosophila melanogaster and across predicted splice junctions. RNA expression of protein coding and nonprotein coding sequences was determined for each major stage of the life cycle, including adult males and females. We detected transcriptional activity for 93% of annotated genes and RNA expression for 41% of the probes in intronic and intergenic sequences. Comparison to genome-wide RNA interference data and to gene annotations revealed distinguishable levels of expression for different classes of genes and higher levels of expression for genes with essential cellular functions. Differential splicing was observed in about 40% of predicted genes, and 5440 previously unknown splice forms were detected. Genes within conserved regions of synteny with D. pseudoobscura had highly correlated expression; these regions ranged in length from 10 to 900 kilobase pairs. The expressed intergenic and intronic sequences are more likely to be evolutionarily conserved than nonexpressed ones, and about 15% of them appear to be developmentally regulated. Our results provide a draft expression map for the entire nonrepetitive genome, which reveals a much more extensive and diverse set of expressed sequences than was previously predicted.

Algorithms↗

Multiplexed discovery of sequence polymorphisms using base-specific cleavage and MALDI-TOF MS.

The completion of the Human Genome Project provides researchers with a reference sequence that covers about 99% of the gene-containing regions and is more than 99.9% accurate. Sequence drafts and completed sequences for several other species are also available to researchers worldwide. The ongoing effort to provide more and more genomic reference information now enables the detection of deviations from this 'genetic blueprint'. Comparative sequencing projects will play a major role in elucidating the meaning of the genetic code and in establishing a correlation between genotype and phenotype. As part of this effort, a number of projects will focus on distinct functional aspects, like resequencing of exons or HLA determining regions. Typically these target regions are short in length and their analysis does not require long read length. To find an efficient solution for these applications, we developed a novel method that allows simultaneous analysis of multiple independent target regions (Multiplexed Comparative Sequence Analysis) by employing base-specific cleavage biochemistry and MALDI TOF-MS analysis.

Humans↗

High throughput genotyping technologies.

A comprehensive genetic map containing several hundred microsatellite markers resulted from a large microsatellite mapping project. This was the first real study that introduced high throughput methods to the genetic community. This map and the concurrent technological advances, which will briefly be reviewed, led to further numerous mapping investigations of simple and complex diseases. The annotated draft sequence of approximately three billion base pairs (bp) of the human genome has been completed much sooner than many imagined, due to considerable technological advancements and the international enterprise that resulted. This was a major development for the genetics community, but is only the precursor to the next phase of studying and understanding the variation within the human genome. The awareness of the differences may help us understand the effects on the genetics of the variation between individuals and disease. It is these variations at the nucleotide level that determine the physiological differences, or phenotypes of each individual, including all biological functions at the cellular and body level. Single nucleotide polymorphisms (SNPs) will provide the next high density map, and be the genetic tool to study these genetic variations. There are many sources of SNPs and exhaustive numbers of methods of SNP detection to be considered. The focus in this paper will be on the merits of selected, varied SNP typing methodologies that are emerging to genotype many individuals with the required huge number of SNPs to make the study of complex diseases and pharmacogenomics a practical and economically viable option.

Genotype↗

The genome of black cottonwood, Populus trichocarpa (Torr. & Gray).

We report the draft genome of the black cottonwood tree, Populus trichocarpa. Integration of shotgun sequence assembly with genetic mapping enabled chromosome-scale reconstruction of the genome. More than 45,000 putative protein-coding genes were identified. Analysis of the assembled genome revealed a whole-genome duplication event; about 8000 pairs of duplicated genes from that event survived in the Populus genome. A second, older duplication event is indistinguishably coincident with the divergence of the Populus and Arabidopsis lineages. Nucleotide substitution, tandem gene duplication, and gross chromosomal rearrangement appear to proceed substantially more slowly in Populus than in Arabidopsis. Populus has more protein-coding genes than Arabidopsis, ranging on average from 1.4 to 1.6 putative Populus homologs for each Arabidopsis gene. However, the relative frequency of protein domains in the two genomes is similar. Overrepresented exceptions in Populus include genes associated with lignocellulosic wall biosynthesis, meristem development, disease resistance, and metabolite transport.

Arabidopsis↗

Target validation and functional analyses using antisense oligonucleotides.

The human genome project (HGP) has been described as the single most important project in biology and the biomedical sciences to date. In February 2001, the efforts of the HGP resulted in the publication of a 'working draft' of the entire human genome and it is expected that final sequencing and annotation of the genome will be completed by 2003. Researchers are now focusing efforts on the identification of the function of the reported 30,000 human genes. During the past few years, antisense oligomers have been widely used as potent tools for functional genomics and drug target validation. This article describes the emerging and established antisense technologies that will be used to continue the efforts to unlock the function of the human genome and to discover novel drug targets for the treatment of human diseases.

Journal Article↗

Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model.

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked seven representative methods spanning gene-cluster, compacted coloured de Bruijn graph, one hybrid approach and one multiple sequence alignment method. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas compacted coloured de Bruijn graphs expanded, with distinct degree-prevalence fingerprints across tools. In contrast, the multiple sequence alignment method could not be evaluated across fragmented inputs because it did not run reliably on draft-assembly datasets. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one compacted coloured de Bruijn graph implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation by gene-cluster-based tools does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that, in this repeat-rich O157:H7 benchmark dataset, assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy.

Escherichia coli O157↗

In-depth view of structure, activity, and evolution of rice chromosome 10.

Rice is the world's most important food crop and a model for cereal research. At 430 megabases in size, its genome is the most compact of the cereals. We report the sequence of chromosome 10, the smallest of the 12 rice chromosomes (22.4 megabases), which contains 3471 genes. Chromosome 10 contains considerable heterochromatin with an enrichment of repetitive elements on 10S and an enrichment of expressed genes on 10L. Multiple insertions from organellar genomes were detected. Collinearity was apparent between rice chromosome 10 and sorghum and maize. Comparison between the draft and finished sequence demonstrates the importance of finished sequence.

Chromosomes, Plant↗

Segmental duplications: organization and impact within the current human genome project assembly.

Segmental duplications play fundamental roles in both genomic disease and gene evolution. To understand their organization within the human genome, we have developed the computational tools and methods necessary to detect identity between long stretches of genomic sequence despite the presence of high copy repeats and large insertion-deletions. Here we present our analysis of the most recent genome assembly (January 2001) in which we focus on the global organization of these segments and the role they play in the whole-genome assembly process. Initially, we considered only large recent duplication events that fell well-below levels of draft sequencing error (alignments 90%-98% similar and > or =1 kb in length). Duplications (90%-98%; > or =1 kb) comprise 3.6% of all human sequence. These duplications show clustering and up to 10-fold enrichment within pericentromeric and subtelomeric regions. In terms of assembly, duplicated sequences were found to be over-represented in unordered and unassigned contigs indicating that duplicated sequences are difficult to assign to their proper position. To assess coverage of these regions within the genome, we selected BACs containing interchromosomal duplications and characterized their duplication pattern by FISH. Only 47% (106/224) of chromosomes positive by FISH had a corresponding chromosomal position by comparison. We present data that indicate that this is attributable to misassembly, misassignment, and/or decreased sequencing coverage within duplicated regions. Surprisingly, if we consider putative duplications >98% identity, we identify 10.6% (286 Mb) of the current assembly as paralogous. The majority of these alignments, we believe, represent unmerged overlaps within unique regions. Taken together the above data indicate that segmental duplications represent a significant impediment to accurate human genome assembly, requiring the development of specialized techniques to finish these exceptional regions of the genome. The identification and characterization of these highly duplicated regions represents an important step in the complete sequencing of a human reference genome.

Base Sequence↗

From mapping to sequencing, post-sequencing and beyond.

The Rice Genome Research Program (RGP) in Japan has been collaborating with the international community in elucidating a complete high-quality sequence of the rice genome. As the pioneer in large-scale analysis of the rice genome, the RGP has successfully established the fundamental tools for genome research such as a genetic map, a yeast artificial chromosome (YAC)-based physical map, a transcript map and a phage P1 artificial chromosome (PAC)/bacterial artificial chromosome (BAC) sequence-ready physical map, which serve as common resources for genome sequencing. Among the 12 rice chromosomes, the RGP is in charge of sequencing six chromosomes covering 52% of the 390 Mb total length of the genome. The contribution of the RGP to the realization of decoding the rice genome sequence with high accuracy and deciphering the genetic information in the genome will have a great impact in understanding the biology of the rice plant that provides a major food source for almost half of the world's population. A high-quality draft sequence (phase 2) was completed in December 2002. Since then, much of the finished quality sequence (phase 3) has become available in public databases. With the completion of sequencing in December 2004, it is expected that the genome sequence would facilitate innovative research in functional and applied genomics. A map-based genome sequence is indispensable for further improvement of current rice varieties and for development of novel varieties carrying agronomically important traits such as high yield potential and tolerance to both biotic and abiotic stresses. In addition to genome sequencing, various related projects have been initiated to generate valuable resources, which could serve as indispensable tools in clarifying the structure and function of the rice genome. These resources have been made available to the scientific community through the Rice Genome Resource Center (RGRC) of the National Institute of Agrobiological Sciences (NIAS) to enable rapid progress in research that will lead to thorough understanding of the rice plant. As the next trend in rice genome research will focus on determining the function of about 40,000-50,000 genes predicted in the genome as well as applying various genomics tools in rice breeding, an unlimited access to rice DNA and seed stocks will provide a broad community of scientists with the necessary materials for formulating new concepts, developing innovative research and making new scientific discoveries in rice genomics.

Centromere↗

Physical and transcript map of the hereditary prostate cancer region at xq27.

We have recently mapped a locus for hereditary prostate cancer (termed HPCX) to the long arm of the X chromosome (Xq25-q27) through a genome-wide linkage study. Here we report the construction of an approximately 9-Mb sequence-ready bacterial clone contig map of Xq26.3-q27.3. The contig was constructed by screening BAC/PAC libraries with markers spaced at approximately 85-kb intervals. We identified overlapping clones by end-sequencing framework clones to generate 407 new sequence-tagged sites, followed by PCR verification of overlaps. Contig assembly was based on clone restriction fingerprinting and the landmark information. We identified a minimal overlap contig for genomic sequencing, which has yielded 7.7 Mb of finished sequence and 1.5 Mb of draft sequence. The transcriptional mapping effort localized 57 known and predicted genes by database searching, STS content mapping, and sequencing, followed by sequence annotation. These transcriptional units represent candidate genes for HPCX and multiple other hereditary diseases at Xq26.3-q27.3.

Chromosome Mapping↗

Chromosomal mapping of 170 BAC clones in the ascidian Ciona intestinalis.

The draft genome ( approximately 160 Mb) of the urochordate ascidian Ciona intestinalis has been sequenced by the whole-genome shotgun method and should provide important insights into the origin and evolution of chordates as well as vertebrates. However, because this genomic data has not yet been mapped onto chromosomes, important biological questions including regulation of gene expression at the genome-wide level cannot yet be addressed. Here, we report the molecular cytogenetic characterization of all 14 pairs of C. intestinalis chromosomes, as well as initial large-scale mapping of genomic sequences onto chromosomes by fluorescent in situ hybridization (FISH). Two-color FISH using 170 bacterial artificial chromosome (BAC) clones and construction of joined scaffolds using paired BAC end sequences allowed for mapping of up to 65% of the deduced 117-Mb nonrepetitive sequence onto chromosomes. This map lays the foundation for future studies of the protochordate C. intestinalis genome at the chromosomal level.

Animals↗

Genomics. Public-private project to deliver mouse genome in 6 months.

Research on the mouse genome lurched into the fast lane last week, as private donors joined the U.S. government to step on the gas. A public-private consortium announced on 6 October that it's kicking $58 million into a new fund that will pay to sequence the DNA of the "black six" (C57BL/6J) strain of laboratory mouse. The consortium aims to produce a draft version of the genome by the end of February.

Animals↗