PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

A genetic and structural analysis of the N-glycosylation capabilities.

The recent draft sequencing of the rice (Oryza sativa) genome has enabled a genetic analysis of the glycosylation capabilities of an agroeconomically important group of plants, the monocotyledons. In this study, we have not only identified genes putatively encoding enzymes involved in N-glycosylation, but have examined by MALDI-TOF MS the structures of the N-glycans of rice and other monocotyledons (maize, wheat and dates; Zea mays, Triticum aestivum and Phoenix dactylifera); these data show that within the plant kingdom the types of N-glycans found are very similar between monocotyledons, dicotyledons and gymnosperms. Subsequently, we constructed expression vectors for the key enzymes forming plant-typical structures in rice, N-acetylglucosaminyltransferase I (GlcNAc-TI; EC 2.4.1.101), core alpha1,3-fucosyltransferase (FucTA; EC 2.4.1.214) and beta1,2-xylosyltransferase (EC 2.4.2.38) and successfully expressed them in Pichia pastoris. Rice GlcNAc-TI, FucTA and xylosyltransferase are therefore the first monocotyledon glycosyltransferases involved in N-glycan biosynthesis to be characterised in a recombinant form.

Amino Acid Sequence↗

Inventory and comparative analysis of rice and Arabidopsis ATP-binding cassette (ABC) systems.

ATP-binding cassette (ABC) proteins constitute a large superfamily found in all kingdoms of living organisms. The recent completion of two draft sequences of the rice (Oryza sativa) genome allowed us to analyze and classify its ABC proteins and to compare to those in Arabidopsis thaliana. We identified a similar number of ABC proteins in rice and Arabidopsis (121 versus 120), despite the rice genome being more than three times the size of Arabidopsis. Both Arabidopsis and rice have representative members in all seven major subfamilies of ABC ATPases (A to G) commonly found in eukaryotes. This comparative analysis allowed the detection of 29 potential orthologous sequences in Arabidopsis and rice. However, plant share with prokaryotes a specific set of ABC systems that is not detected in animals. These ABC systems might be inherited from the cyanobacterial ancestor of chloroplasts. The present work provides the first complete inventory of rice ABC proteins and an updated inventory of those proteins in Arabidopsis.

ATP-Binding Cassette Transporters↗

Novel approaches for identifying genes regulating lymphocyte development and function.

The draft sequence of the human and mouse genomes provides an unparalleled opportunity for understanding the genetic control of immune-cell development. Strategies can begin with a gene sequence and pursue a putative immune-system function by employing mRNA-expression profiling or creating gene knockouts in embryonic stem cells. The latter can be produced by utilising the Cre/Lox system, a tetracycline operon, a gene-trap method or chemical mutagenesis. Alternatively, mutant phenotypes (derived using the mutagen ethylnitrosourea) can be traced back to gene sequences.

Animals↗

Methods for comparing sources of strand compositional asymmetry in microbial chromosomes.

Significant compositional biases in bacterial chromosomes have been explained by replication- and transcription-coupled repair mechanisms, the latter causing GC skew to indicate the direction of replication when gene polarity is correspondingly entrained. Correlations between indicators of replication direction, skew, and transcription polarity are computed for the complete nucleotide sequences of 20 microbial chromosomes and interpreted through statistical tests. A second quantitative method, previously applied to the first complete draft of the Escherichia coli K12 genome, characterizes the sequences by average skew and net skew due to replication. These methods generally agree in finding the coexistence of replication- and translation-coupled effects and in identifying atypical sequences in which one influence is clearly dominant. The replication-dominated class is exemplified by two chlamydial sequences and the transcription-dominated class by three archaea. The preference for leading-strand transcription in two mycoplasmas is stronger than the skew implies. These concordant methods provide an objective framework for comparing sources of strand compositional asymmetry and interpreting skew diagrams.

Base Composition↗

Assembly of the working draft of the human genome with GigAssembler.

The data for the public working draft of the human genome contains roughly 400,000 initial sequence contigs in approximately 30,000 large insert clones. Many of these initial sequence contigs overlap. A program, GigAssembler, was built to merge them and to order and orient the resulting larger sequence contigs based on mRNA, paired plasmid ends, EST, BAC end pairs, and other information. This program produced the first publicly available assembly of the human genome, a working draft containing roughly 2.7 billion base pairs and covering an estimated 88% of the genome that has been used for several recent studies of the genome. Here we describe the algorithm used by GigAssembler.

Algorithms↗

Integration of microsatellite-based genetic maps for the turkey (Meleagris gallopavo).

Integration of turkey genetic maps and their associated markers is essential to increase marker density in support of map-based genetic studies. The objectives of this study were to integrate 2 microsatellite-based turkey genetic maps--the Roslin map and the University of Minnesota (UMN) map--by genotyping markers from the Roslin study on the mapping families of the UMN study. A total of 279 markers was tested, and 240 were subsequently screened for polymorphisms in the UMN/Nicholas Turkey Breeding Farms (NTBF) mapping families. Of the 240 markers, 89 were genetically informative and were used for genotyping the F2 offspring. Significant genetic linkages (log of odds > 3.0) were found for 84 markers from the Roslin study. BLASTn comparison of marker sequences with the draft assembly of the chicken genome found 263 significant matches. The combination of genetic and in silico mapping allowed for the alignment of all linkage groups of the Roslin map with those of the UMN map. With the addition of the markers from the Roslin map, 438 markers are now genetically linked in the UMN/NTBF families, and more than 1700 turkey sequences have now been assigned to likely positions in the chicken-genome sequence.

Animals↗

Identification of a fatty acid Delta11-desaturase from the microalga Thalassiosira pseudonana.

A set of genomic DNA sequences putatively encoding front-end desaturases were identified by in silico analysis of the draft genome of the marine microalga Thalassiosira pseudonana. Among these candidate genes, an open reading frame named TpdesN was found to be full-length, intronless, and constitutively expressed during cell cultivation. The predicted amino acid sequence of the corresponding protein, TpDESN, exhibited typical features of desaturases involved in the production of polyunsaturated fatty acids (PUFAs) in algae, i.e. a cytochrome b5-like domain at the N-terminus and three conserved histidine-rich motifs in the desaturase domain. Expression of TpDESN in Saccharomyces cerevisiae revealed that this enzyme was not involved in PUFA synthesis, but specifically desaturated palmitic acid 16:0 to 16:1Delta11. To our knowledge, until this report, Delta11-desaturase activity had only been detected in insect cells.

Amino Acid Motifs↗

Identification of a novel Bardet-Biedl syndrome protein, BBS7, that shares structural features with BBS1 and BBS2.

Bardet-Biedl syndrome (BBS) is a genetically heterogeneous disorder, the primary features of which include obesity, retinal dystrophy, polydactyly, hypogenitalism, learning difficulties, and renal malformations. Conventional linkage and positional cloning have led to the mapping of six BBS loci in the human genome, four of which (BBS1, BBS2, BBS4, and BBS6) have been cloned. Despite these advances, the protein sequences of the known BBS genes have provided little or no insight into their function. To delineate functionally important regions in BBS2, we performed phylogenetic and genomic studies in which we used the human and zebrafish BBS2 peptide sequences to search dbEST and the translation of the draft human genome. We identified two novel genes that we initially named "BBS2L1" and "BBS2L2" and that exhibit modest similarity with two discrete, overlapping regions of BBS2. In the present study, we demonstrate that BBS2L1 mutations cause BBS, thereby defining a novel locus for this syndrome, BBS7, whereas BBS2L2 has been shown independently to be BBS1. The motif-based identification of a novel BBS locus has enabled us to define a potential functional domain that is present in three of the five known BBS proteins and, therefore, is likely to be important in the pathogenesis of this complex syndrome.

Adaptor Proteins, Signal Transducing↗

The phusion assembler.

The Phusion assembler has assembled the mouse genome from the whole-genome shotgun (WGS) dataset collected by the Mouse Genome Sequencing Consortium, at ~7.5x sequence coverage, producing a high-quality draft assembly 2.6 gigabases in size, of which 90% of these bases are in 479 scaffolds. For the mouse genome, which is a large and repeat-rich genome, the input dataset was designed to include a high proportion of paired end sequences of various size selected inserts, from 2-200 kbp lengths, into various host vector templates. Phusion uses sequence data, called reads, and information about reads that share common templates, called read pairs, to drive the assembly of this large genome to highly accurate results. The preassembly stage, which clusters the reads into sensible groups, is a key element of the entire assembler, because it permits a simple approach to parallelization of the assembly stage, as each cluster can be treated independent of the others. In addition to the application of Phusion to the mouse genome, we will also present results from the WGS assembly of Caenorhabditis briggsae sequenced to about 11x coverage. The C. briggsae assembly was accessioned through EMBL, http://www.ebi.ac.uk/services/index.html, using the series CAAC01000001-CAAC01000578, however, the Phusion mouse assembly described here was not accessioned. The mouse data was generated by the Mouse Genome Sequencing Consortium. The C. briggsae sequence was generated at The Wellcome Trust Sanger Institute and the Genome Sequencing Center, Washington University School of Medicine.

Animals↗

Genome-wide comparative analysis of the transposable elements in the related species Arabidopsis thaliana and Brassica oleracea.

Transposable elements (TEs) are the major component of plant genomes where they contribute significantly to the >1,000-fold genome size variation. To understand the dynamics of TE-mediated genome expansion, we have undertaken a comparative analysis of the TEs in two related organisms: the weed Arabidopsis thaliana (125 megabases) and Brassica oleracea ( approximately 600 megabases), a species with many crop plants. Comparison of the whole genome sequence of A. thaliana with a partial draft of B. oleracea has permitted an estimation of the patterns of TE amplification, diversification, and loss that has occurred in related species since their divergence from a common ancestor. Although we find that nearly all TE lineages are shared, the number of elements in each lineage is almost always greater in B. oleracea. Class 1 (retro) elements are the most abundant TE class in both species with LTR and non-LTR elements comprising the largest fraction of each genome. However, several families of class 2 (DNA) elements have amplified to very high copy number in B. oleracea where they have contributed significantly to genome expansion. Taken together, the results of this analysis indicate that amplification of both class 1 and class 2 TEs is responsible, in part, for B. oleracea genome expansion since divergence from a common ancestor with A. thaliana. In addition, the observation that B. oleracea and A. thaliana share virtually all TE lineages makes it unlikely that wholesale removal of TEs is responsible for the compact genome of A. thaliana.

Arabidopsis↗

Fugu and human sequence comparison identifies novel human genes and conserved non-coding sequences.

The compact genome of the pufferfish, Fugu rubripes, has been proposed as a 'reference' genome to aid in annotating and analysing the human genome. We have annotated and compared 85 kb of Fugu sequence containing 17 genes with its homologous loci in the human draft genome and identified three 'novel' human genes that were missed or incompletely predicted by the previous gene prediction methods. Two of the novel genes contain zinc finger domains and are designated ZNF366 and ZNF367. They map to human chromosomes 5q13.2 and 9q22.32, respectively. The third novel gene, designated C9orf21, maps to chromosome 9q22.32. This gene is unique to vertebrates, and the protein encoded by it does not contain any known domains. We could not find human homologs for two Fugu genes, a novel chemokine gene and a kinase gene. These genes are either specific to teleosts or lost in the human lineage. The Fugu-human comparison identified several conserved non-coding sequences in the promoter and intronic regions. These sequences, conserved during 450 million years of vertebrate evolution, are likely to be involved in gene regulation. The 85 kb Fugu locus is dispersed over four human loci, occupying about 1.5 Mb. Contiguity is conserved in the human genome between six out of 16 Fugu gene pairs. These contiguous chromosomal segments should share a common evolutionary history dating back to the common ancestor of mammals and teleosts. We propose contiguity as strong evidence to identify orthologous genes in distant organisms. This study confirms the utility of the Fugu as a supplementary tool to uncover and confirm novel genes and putative gene regulatory regions in the human genome.

Amino Acid Sequence↗

Diagnostic approach to children with birth defects.

Clinical genetics deals with the diagnosis, management and prevention of genetically determined disorders. Our current understanding of the role genes play in the pathogenesis of everything from fetal malformation to neurodegenerative and malignant disorders of late adulthood make it somewhat difficult to draw a clear boundary for this rapidly expanding specialty. With the recent completion of a preliminary draft of the entire sequence of the human genome it is not unreasonable to dream of novel therapeutic approaches such as "gene therapy", to cure disorders heretofore treatable with supportive measures only. Nevertheless, the clinical assessment of the patient will continue to be the cornerstone of good practice of medicine. In this article we review a clinical approach to the diagnostic challenge presented by children with birth defects. The principles we illustrate apply to other aspects of "genetic medicine" as well.

Abnormalities, Multiple↗

Strategies and tools for whole-genome alignments.

The availability of the assembled mouse genome makes possible, for the first time, an alignment and comparison of two large vertebrate genomes. We investigated different strategies of alignment for the subsequent analysis of conservation of genomes that are effective for assemblies of different quality. These strategies were applied to the comparison of the working draft of the human genome with the Mouse Genome Sequencing Consortium assembly, as well as other intermediate mouse assemblies. Our methods are fast and the resulting alignments exhibit a high degree of sensitivity, covering more than 90% of known coding exons in the human genome. We obtained such coverage while preserving specificity. With a view towards the end user, we developed a suite of tools and Web sites for automatically aligning and subsequently browsing and working with whole-genome comparisons. We describe the use of these tools to identify conserved non-coding regions between the human and mouse genomes, some of which have not been identified by other methods.

Algorithms↗

Comparative DNA sequence analysis of mapped wheat ESTs reveals the complexity of genome relationships between rice and wheat.

The use of DNA sequence-based comparative genomics for evolutionary studies and for transferring information from model species to related large-genome species has revolutionized molecular genetics and breeding strategies for improving those crops. Comparative sequence analysis methods can be used to cross-reference genes between species maps, enhance the resolution of comparative maps, study patterns of gene evolution, identify conserved regions of the genomes, and facilitate interspecies gene cloning. In this study, 5,780 Triticeae ESTs that have been physically mapped using wheat ( Triticum aestivum L.) deletion lines and segregating populations were compared using NCBI BLASTN to the first draft of the public rice ( Oryza sativa L.) genome sequence data from 3,280 ordered BAC/PAC clones. A rice genome view of the homoeologous wheat genome locations based on sequence analysis shows general similarity to the previously published comparative maps based on Southern analysis of RFLP. For most rice chromosomes there is a preponderance of wheat genes from one or two wheat chromosomes. The physical locations of non-conserved regions were not consistent across rice chromosomes. Some wheat ESTs with multiple wheat genome locations are associated with the non-conserved regions of similarity between rice and wheat. The inverse view, showing the relationship between the wheat deletion map and rice genomic sequence, revealed the breakdown of gene content and order at the resolution conferred by the physical chromosome deletions in the wheat genome. An average of 35% of the putative single copy genes that were mapped to the most conserved bins matched rice chromosomes other than the one that was most similar. This suggests that there has been an abundance of rearrangements, insertions, deletions, and duplications eroding the wheat-rice genome relationship that may complicate the use of rice as a model for cross-species transfer of information in non-conserved regions.

Chromosome Mapping↗

HTS in the new millennium: the role of pharmacology and flexibility.

Over the past decade, high throughput screening (HTS) has become the focal point for discovery programs within the pharmaceutical industry. The role of this discipline has been and remains the rapid and efficient identification of lead chemical matter within chemical libraries for therapeutics development. Recent advances in molecular and computational biology, i.e., genomic sequencing and bioinformatics, have resulted in the announcement of publication of the first draft of the human genome. While much work remains before a complete and accurate genomic map will be available, there can be no doubt that the number of potential therapeutic intervention points will increase dramatically, thereby increasing the workload of early discovery groups. One current drug discovery paradigm integrates genomics, protein biosciences and HTS in establishing what the authors refer to as the "gene-to-screen" process. Adoption of the "gene-to-screen" paradigm results in a dramatic increase in the efficiency of the process of converting a novel gene coding for a putative enzymatic or receptor function into a robust and pharmacologically relevant high throughput screen. This article details aspects of the identification of lead chemical matter from HTS. Topics discussed include portfolio composition (molecular targets amenable to small molecule drug discovery), screening file content, assay formats and plating densities, and the impact of instrumentation on the ability of HTS to identify lead chemical matter.

Animals↗

Sequence information can be obtained from single DNA molecules.

The completion of the human genome draft has taken several years and is only the beginning of a period in which large amounts of DNA and RNA sequence information will be required from many individuals and species. Conventional sequencing technology has limitations in cost, speed, and sensitivity, with the result that the demand for sequence information far outstrips current capacity. There have been several proposals to address these issues by developing the ability to sequence single DNA molecules, but none have been experimentally demonstrated. Here we report the use of DNA polymerase to obtain sequence information from single DNA molecules by using fluorescence microscopy. We monitored repeated incorporation of fluorescently labeled nucleotides into individual DNA strands with single base resolution, allowing the determination of sequence fingerprints up to 5 bp in length. These experiments show that one can study the activity of DNA polymerase at the single molecule level with single base resolution and a high degree of parallelization, thus providing the foundation for a practical single molecule sequencing technology.

Base Sequence↗

Whole-genome sequence of Streptococcus agalactiae strain GIFTS31 isolated from streptococcosis-infected Nile tilapia in Bangladesh.

Streptococcus agalactiae strain GIFTS31 was isolated from a Nile tilapia infected with streptococcosis in Gazipur, Bangladesh. The draft genome of GIFTS31 comprises 2,039,674 bp with a GC content of 35% and encodes 1,957 predicted protein-coding sequences. The genome sequence provides valuable insights into the pathogenic potential of fish-associated S. agalactiae.

Streptococcus agalactiae↗

Interrogating the human genome using uninterpreted mass spectrometry data.

The public availability of a draft assembly of the human genome has enabled us to demonstrate, for the first time, the feasibility of searching a complete, unmasked eukaryotic genome using uninterpreted mass spectrometry data. A complex LC-MS/MS data set, containing peptides from at least 22 human proteins, was searched against a comprehensive, nonidentical protein database, an expressed sequence tag (EST) database, and the International Human Genome Project draft assembly of the human genome. The results from the three searches are compared in detail, and the merits of the different databases for this application are discussed. In the case of the EST database, the UniGene index provided a method of simplifying and summarising the search results. In the case of the genomic DNA, the presence of introns prevented matching of roughly one quarter of the spectra, but the technique can provide primary experimental verification of predicted coding sequences, and has the potential to identify novel coding sequences.

Algorithms↗