PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

From mapping to sequencing, post-sequencing and beyond.

The Rice Genome Research Program (RGP) in Japan has been collaborating with the international community in elucidating a complete high-quality sequence of the rice genome. As the pioneer in large-scale analysis of the rice genome, the RGP has successfully established the fundamental tools for genome research such as a genetic map, a yeast artificial chromosome (YAC)-based physical map, a transcript map and a phage P1 artificial chromosome (PAC)/bacterial artificial chromosome (BAC) sequence-ready physical map, which serve as common resources for genome sequencing. Among the 12 rice chromosomes, the RGP is in charge of sequencing six chromosomes covering 52% of the 390 Mb total length of the genome. The contribution of the RGP to the realization of decoding the rice genome sequence with high accuracy and deciphering the genetic information in the genome will have a great impact in understanding the biology of the rice plant that provides a major food source for almost half of the world's population. A high-quality draft sequence (phase 2) was completed in December 2002. Since then, much of the finished quality sequence (phase 3) has become available in public databases. With the completion of sequencing in December 2004, it is expected that the genome sequence would facilitate innovative research in functional and applied genomics. A map-based genome sequence is indispensable for further improvement of current rice varieties and for development of novel varieties carrying agronomically important traits such as high yield potential and tolerance to both biotic and abiotic stresses. In addition to genome sequencing, various related projects have been initiated to generate valuable resources, which could serve as indispensable tools in clarifying the structure and function of the rice genome. These resources have been made available to the scientific community through the Rice Genome Resource Center (RGRC) of the National Institute of Agrobiological Sciences (NIAS) to enable rapid progress in research that will lead to thorough understanding of the rice plant. As the next trend in rice genome research will focus on determining the function of about 40,000-50,000 genes predicted in the genome as well as applying various genomics tools in rice breeding, an unlimited access to rice DNA and seed stocks will provide a broad community of scientists with the necessary materials for formulating new concepts, developing innovative research and making new scientific discoveries in rice genomics.

Centromere↗

Physical and transcript map of the hereditary prostate cancer region at xq27.

We have recently mapped a locus for hereditary prostate cancer (termed HPCX) to the long arm of the X chromosome (Xq25-q27) through a genome-wide linkage study. Here we report the construction of an approximately 9-Mb sequence-ready bacterial clone contig map of Xq26.3-q27.3. The contig was constructed by screening BAC/PAC libraries with markers spaced at approximately 85-kb intervals. We identified overlapping clones by end-sequencing framework clones to generate 407 new sequence-tagged sites, followed by PCR verification of overlaps. Contig assembly was based on clone restriction fingerprinting and the landmark information. We identified a minimal overlap contig for genomic sequencing, which has yielded 7.7 Mb of finished sequence and 1.5 Mb of draft sequence. The transcriptional mapping effort localized 57 known and predicted genes by database searching, STS content mapping, and sequencing, followed by sequence annotation. These transcriptional units represent candidate genes for HPCX and multiple other hereditary diseases at Xq26.3-q27.3.

Chromosome Mapping↗

Chromosomal mapping of 170 BAC clones in the ascidian Ciona intestinalis.

The draft genome ( approximately 160 Mb) of the urochordate ascidian Ciona intestinalis has been sequenced by the whole-genome shotgun method and should provide important insights into the origin and evolution of chordates as well as vertebrates. However, because this genomic data has not yet been mapped onto chromosomes, important biological questions including regulation of gene expression at the genome-wide level cannot yet be addressed. Here, we report the molecular cytogenetic characterization of all 14 pairs of C. intestinalis chromosomes, as well as initial large-scale mapping of genomic sequences onto chromosomes by fluorescent in situ hybridization (FISH). Two-color FISH using 170 bacterial artificial chromosome (BAC) clones and construction of joined scaffolds using paired BAC end sequences allowed for mapping of up to 65% of the deduced 117-Mb nonrepetitive sequence onto chromosomes. This map lays the foundation for future studies of the protochordate C. intestinalis genome at the chromosomal level.

Animals↗

Genomics. Public-private project to deliver mouse genome in 6 months.

Research on the mouse genome lurched into the fast lane last week, as private donors joined the U.S. government to step on the gas. A public-private consortium announced on 6 October that it's kicking $58 million into a new fund that will pay to sequence the DNA of the "black six" (C57BL/6J) strain of laboratory mouse. The consortium aims to produce a draft version of the genome by the end of February.

Animals↗

The evolution of vertebrate Toll-like receptors.

The complete sequences of Takifugu Toll-like receptor (TLR) loci and gene predictions from many draft genomes enable comprehensive molecular phylogenetic analysis. Strong selective pressure for recognition of and response to pathogen-associated molecular patterns has maintained a largely unchanging TLR recognition in all vertebrates. There are six major families of vertebrate TLRs. This repertoire is distinct from that of invertebrates. TLRs within a family recognize a general class of pathogen-associated molecular patterns. Most vertebrates have exactly one gene ortholog for each TLR family. The family including TLR1 has more species-specific adaptations than other families. A major family including TLR11 is represented in humans only by a pseudogene. Coincidental evolution plays a minor role in TLR evolution. The sequencing phase of this study produced finished genomic sequences for the 12 Takifugu rubripes TLRs. In addition, we have produced >70 gene models, including sequences from the opossum, chicken, frog, dog, sea urchin, and sea squirt.

Animals↗

An integrated culturomic and genomic database and analysis platform for methanogenic archaea.

Methanogenic archaea research is challenged by limited strain resources, fragmented genomic data, inconsistent genome quality, substantial uncultured lineages, and difficulties in laboratory culturing, hindering advances in biogas production, climate mitigation, and microbial ecology. These archaea play crucial roles in global carbon cycling and anaerobic environments, yet scattered data and unculturable strains limit systematic studies and applications. To address this, we created MethArDB (Methanogenic Archaeal Genome Database), a specialized database for methanogenic archaea, compiling 3919 genomes, 87 host-associated plasmids, and 42 phages, with standardized quality classifications (complete, scaffold, draft), protein sequences, and metadata on geography, habitats, metabolism, and inheritable elements. Integrated MethArCT (Methanogenic Archaeal Culturomics Toolkit) employs a dual-threshold orthologous/paralogous protein analysis to evaluate metabolic pathway completeness, predicting cultivation parameters and suggesting candidate cultivation strategies, including potential medium formulations and conditions, to support strain isolation. Overall, MethArDB and MethArCT form an integrated platform combining genomics and culturomics to facilitate methanogenic archaea research. Database URL:  http://methardb.cn.

Genome, Archaeal↗

Whole-genome sequences of Pseudomonas aeruginosa strains isolated from patients and inanimate hospital reservoirs in Chattogram, Bangladesh.

Antimicrobial-resistant Pseudomonas aeruginosa (PA) is a critical concern for the global population. This study identified four isolates and reported the draft genomes annotated as ST645, ST2238, and the high-risk clone ST773. The isolates featured diverse antimicrobial resistance (AMR) gene profiles that emphasize the need for continuous molecular surveillance in Bangladesh.

Chattogram↗

The circadian clock of the unicellular eukaryotic model organism Chlamydomonas reinhardtii.

The green unicellular alga Chlamydomonas reinhardtii, also called 'green yeast', emerged in the past years as a model organism for specific scientific questions such as chloroplast biogenesis and function, the composition of the flagella including its basal apparatus, or the mechanism of the circadian clock. Sequencing of its chloroplast and mitochondrial genomes have already been completed and a first draft of its nuclear genome has also been released recently. In C. reinhardtii several circadian rhythms are physiologically well characterized, and one of them has even been shown to operate in outer space. Circadian expression patterns of nuclear and plastid genes have been studied. The mode of regulation of these genes occurs at the transcriptional level, although there is also evidence for posttranscriptional control. A clock-controlled, phylogenetically conserved RNA-binding protein was characterized in this alga, which interacts with several mRNAs that all contain a common cis-acting motif. Its function within the circadian system is currently under investigation. This review summarizes the current state of the knowledge about the circadian system in C. reinhardtii and points out its potential for future studies.

Animals↗

Genomic and transcriptional analysis of protein heterogeneity of the honeybee venom allergen Api m 6.

Several components of honeybee venom are known to cause allergenic responses in humans and other vertebrates. One such component, the minor allergen Api m 6, has been known to show amino acid variation but the genetic mechanism for this variation is unknown. Here we show that Api m 6 is derived from a single locus, and that substantial protein-level variation has a simple genome-level cause, without the need to invoke multiple loci or alternatively spliced exons. Api m 6 sits near a misassembled section of the honeybee genome sequence, and we propose that a substantial number of indels at and near Api m 6 might be the root cause of this misassembly. We suggest that genes such as Api m 6 with coding-region or untranslated region indels might have had a strong effect on the assembly of this draft of the honeybee genome.

Allergens↗

BAC end sequences and a physical map reveal transposable element content and clustering patterns in the genome of Magnaporthe grisea.

Transposable elements (TEs) are viewed as major contributors to the evolution of fungal genomes. Genomic resources such as BAC libraries are an underutilized resource for studying genome-wide TE distribution. Using the BAC end sequences and physical map that are available for the rice blast fungus, Magnaporthe grisea, we describe a likelihood ratio test designed to identify clustering of TEs in the genome. A significant variation in the distribution of three TEs, MAGGY, MGL, and Pot2 was observed among the fingerprint contigs of the physical map. We utilized a draft sequence of M. grisea chromosome 7 to validate our results and found a similar pattern of clustering. By examining individual BAC end sequences, we found evidence for 11 unique integrations of MAGGY or MGL into Pot2 but no evidence for the reciprocal integration of Pot2 into another TE. This suggests that: (a) the presence of Pot2 in the genome predates that of the other TEs, (b) Pot2 was less transpositionally active than other TEs, or (c) that MAGGY and MGL have integration site preference for Pot2. High transition/transversion mutation ratios as well as bias in transition site context was observed in MAGGY and MGL elements, but not in Pot2 elements. These features are consistent with the effects of a Repeat-Induced Point (RIP) mutation-like process occurring in MAGGY and MGL elements. This study illustrates the general utility of a physical map and BAC end sequences for the study of genome-wide repetitive DNA content and organization.

Chromosomes, Artificial, Bacterial↗

A second generation radiation hybrid map to aid the assembly of the bovine genome sequence.

BACKGROUND: Several approaches can be used to determine the order of loci on chromosomes and hence develop maps of the genome. However, all mapping approaches are prone to errors either arising from technical deficiencies or lack of statistical support to distinguish between alternative orders of loci. The accuracy of the genome maps could be improved, in principle, if information from different sources was combined to produce integrated maps. The publicly available bovine genomic sequence assembly with 6x coverage (Btau_2.0) is based on whole genome shotgun sequence data and limited mapping data however, it is recognised that this assembly is a draft that contains errors. Correcting the sequence assembly requires extensive additional mapping information to improve the reliability of the ordering of sequence scaffolds on chromosomes. The radiation hybrid (RH) map described here has been contributed to the international sequencing project to aid this process. RESULTS: An RH map for the 30 bovine chromosomes is presented. The map was built using the Roslin 3000-rad RH panel (BovGen RH map) and contains 3966 markers including 2473 new loci in addition to 262 amplified fragment-length polymorphisms (AFLP) and 1231 markers previously published with the first generation RH map. Sequences of the mapped loci were aligned with published bovine genome maps to identify inconsistencies. In addition to differences in the order of loci, several cases were observed where the chromosomal assignment of loci differed between maps. All the chromosome maps were aligned with the current 6x bovine assembly (Btau_2.0) and 2898 loci were unambiguously located in the bovine sequence. The order of loci on the RH map for BTA 5, 7, 16, 22, 25 and 29 differed substantially from the assembled bovine sequence. From the 2898 loci unambiguously identified in the bovine sequence assembly, 131 mapped to different chromosomes in the BovGen RH map. CONCLUSION: Alignment of the BovGen RH map with other published RH and genetic maps showed higher consistency in marker order and chromosome assignment than with the current 6x sequence assembly. This suggests that the bovine sequence assembly could be significantly improved by incorporating additional independent mapping information.

Animals↗

Unravelling the role of the ToxR-like transcriptional regulator WmpR in the marine antifouling bacterium Pseudoalteromonas tunicata.

The dark-green-pigmented marine bacterium Pseudoalteromonas tunicata produces several target-specific compounds that act against a range of common fouling organisms, including bacteria, fungi, protozoa, invertebrate larvae and algal spores. The ToxR-like regulator WmpR has previously been shown to regulate expression of bioactive compounds, type IV pili and biofilm formation phenotypes which all appear at the onset of stationary phase. In this study a comparison of survival under starvation or stress between the wild-type P. tunicata strain and a wmpR mutant (D2W2) does not suggest a role for WmpR in regulating starvation- and stress-resistant phenotypes such as those that may be required in stationary phase. Both proteomic [2-dimensional PAGE (2D-PAGE)] and transcriptomic (RNA arbitrarily primed PCR) studies were used to discover members of the WmpR regulon. 2D-PAGE identified 11 proteins that were differentially expressed by WmpR. Peptide sequence data were obtained for six of these proteins and identified using the draft P. tunicata genome as being involved in protein synthesis, amino acid transamination and ubiquinone biosynthesis, as well as hypothetical proteins. The transcriptomic analysis identified three genes significantly up-regulated by WmpR, including a TonB-dependent outer-membrane protein, a non-ribosomal peptide synthetase and a hypothetical protein. Under iron-limitation the wild-type showed greater survival than D2W2, indicating the importance of WmpR under these conditions. Results from these studies show that WmpR controls the expression of genes encoding proteins involved in iron acquisition and uptake, amino acid metabolism and ubiquinone biosynthesis in addition to a number of proteins with as yet unknown functions.

Amino Acids↗

Genomewide identification of Pseudomonas syringae pv. tomato DC3000 promoters controlled by the HrpL alternative sigma factor.

The ability of Pseudomonas syringae pv. tomato DC3000 to parasitize tomato and Arabidopsis thaliana depends on genes activated by the HrpL alternative sigma factor. To support various functional genomic analyses of DC3000, and specifically, to identify genes involved in pathogenesis, we developed a draft sequence of DC3000 and used an iterative process involving computational and gene expression techniques to identify virulence-implicated genes downstream of HrpL-responsive promoters. Hypersensitive response and pathogenicity (Hrp) promoters are known to control genes encoding the Hrp (type III protein secretion) machinery and a few type III effector proteins in DC3000. This process involved (i) identification of 9 new virulence-implicated genes in the Hrp regulon by miniTn5gus mutagenesis, (ii) development of a hidden Markov model (HMM) trained with known and transposon-identified Hrp promoter sequences, (iii) HMM identification of promoters upstream of 12 additional virulence-implicated genes, and (iv) microarray and RNA blot analyses of the HrpL-dependent expression of a representative subset of these DC3000 genes. We found that the Hrp regulon encodes candidates for 4 additional type III secretion machinery accessory factors, homologs of the effector proteins HopPsyA, AvrPpiB1 (2 copies), AvrPpiC2, AvrPphD (2 copies), AvrPphE, AvrPphF, and AvrXv3, and genes associated with the production or metabolism of virulence factors unrelated to the Hrp type III secretion system, including syringomycin synthetase (SyrE), N(epsilon)-(indole-3-acetyl)-l-lysine synthetase (IaaL), and a subsidiary regulon controlling coronatine production. Additional candidate effector genes, hopPtoA2, hopPtoB2, and an avrRps4 homolog, were preceded by Hrp promoter-like sequences, but these had HMM expectation values of relatively low significance and were not detectably activated by HrpL.

Bacterial Proteins↗

Mulan: multiple-sequence local alignment and visualization for studying function and evolution.

Multiple-sequence alignment analysis is a powerful approach for understanding phylogenetic relationships, annotating genes, and detecting functional regulatory elements. With a growing number of partly or fully sequenced vertebrate genomes, effective tools for performing multiple comparisons are required to accurately and efficiently assist biological discoveries. Here we introduce Mulan (http://mulan.dcode.org/), a novel method and a network server for comparing multiple draft and finished-quality sequences to identify functional elements conserved over evolutionary time. Mulan brings together several novel algorithms: the TBA multi-aligner program for rapid identification of local sequence conservation, and the multiTF program for detecting evolutionarily conserved transcription factor binding sites in multiple alignments. In addition, Mulan supports two-way communication with the GALA database; alignments of multiple species dynamically generated in GALA can be viewed in Mulan, and conserved transcription factor binding sites identified with Mulan/multiTF can be integrated and overlaid with extensive genome annotation data using GALA. Local multiple alignments computed by Mulan ensure reliable representation of short- and large-scale genomic rearrangements in distant organisms. Mulan allows for interactive modification of critical conservation parameters to differentially predict conserved regions in comparisons of both closely and distantly related species. We illustrate the uses and applications of the Mulan tool through multispecies comparisons of the GATA3 gene locus and the identification of elements that are conserved in a different way in avians than in other genomes, allowing speculation on the evolution of birds. Source code for the aligners and the aligner-evaluation software can be freely downloaded from http://www.bx.psu.edu/miller_lab/.

Animals↗

454 sequencing put to the test using the complex genome of barley.

BACKGROUND: During the past decade, Sanger sequencing has been used to completely sequence hundreds of microbial and a few higher eukaryote genomes. In recent years, a number of alternative technologies became available, among them adaptations of the pyrosequencing procedure (i.e. "454 sequencing"), promising an approximately 100-fold increase in throughput over Sanger technology--an advancement which is needed to make large and complex genomes more amenable to full genome sequencing at affordable costs. Although several studies have demonstrated its potential usefulness for sequencing small and compact microbial genomes, it was unclear how the new technology would perform in large and highly repetitive genomes such as those of wheat or barley. RESULTS: To study its performance in complex genomes, we used 454 technology to sequence four barley Bacterial Artificial Chromosome (BAC) clones and compared the results to those from ABI-Sanger sequencing. All gene containing regions were covered efficiently and at high quality with 454 sequencing whereas repetitive sequences were more problematic with 454 sequencing than with ABI-Sanger sequencing. 454 sequencing provided a much more even coverage of the BAC clones than ABI-Sanger sequencing, resulting in almost complete assembly of all genic sequences even at only 9 to 10-fold coverage. To obtain highly advanced working draft sequences for the BACs, we developed a strategy to assemble large parts of the BAC sequences by combining comparative genomics, detailed repeat analysis and use of low-quality reads from 454 sequencing. Additionally, we describe an approach of including small numbers of ABI-Sanger sequences to produce hybrid assemblies to partly compensate the short read length of 454 sequences. CONCLUSION: Our data indicate that 454 pyrosequencing allows rapid and cost-effective sequencing of the gene-containing portions of large and complex genomes and that its combination with ABI-Sanger sequencing and targeted sequence analysis can result in large regions of high-quality finished genomic sequences.

Base Pairing↗

The Ensembl genome database project.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of the human genome sequence, with confirmed gene predictions that have been integrated with external data sources, and is available as either an interactive web site or as flat files. It is also an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements from sequence analysis to data storage and visualisation. The Ensembl site is one of the leading sources of human genome sequence annotation and provided much of the analysis for publication by the international human genome project of the draft genome. The Ensembl system is being installed around the world in both companies and academic sites on machines ranging from supercomputers to laptops.

Computational Biology↗