PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

[Nutrition genomics].

The importance of nutrition for human health and its influence on the onset and course of many diseases are nowadays considered as proven. Only the recent development of molecular biology and biochemical methods allows the elucidation of the molecular mechanisms of diet constituent actions and their subsequent effect on homeostatic mechanisms in health and disease states. The availability of the draft human genome sequence as well as the genome sequences of model organisms, combined with the functional and integrative genomics approaches of systems biology, bring about the possibility to identify alleles and haplotypes responsible for specific reaction to the dietary challenge in susceptible individuals. Such complex interactions are studied within the newly conceived field, the nutrition genomics (nutrigenomics). Using the tools of highly parallel analyses of transcriptome, proteome and metabolome, the nutrition genomics pursues its ultimate goal, i.e. the individualized diet, respecting not only quantitative and qualitative nutritional needs and the actual health status, but also the genetic predispositions of an individual. This approach should lead to prevention of the onset of such diseases as obesity, hypertension or type 2 diabetes, or enhance the efficiency of their therapy.

Animals↗

Identification of nine human-specific frameshift mutations by comparative analysis of the human and the chimpanzee genome sequences.

MOTIVATION: The recent release of the draft sequence of the chimpanzee genome is an invaluable resource for finding genome-wide genetic differences that might explain phenotypic differences between humans and chimpanzees. AVAILABILITY: In this paper, we describe a simple procedure to identify potential human-specific frameshift mutations that occurred after the divergence of human and chimpanzee. The procedure involves collecting human coding exons bearing insertions or deletions compared with the chimpanzee genome and identification of homologs from other species, in support of the mutations being human-specific. Using this procedure, we identified nine genes, BASE, DNAJB3, FLJ33674, HEJ1, NTSR2, RPL13AP, SCGB1D4, WBSCR27 and ZCCHC13, that show human-specific alterations including truncations of the C-terminus. In some cases, the frameshift mutation results in gene inactivation or decay. In other cases, the altered protein seems to be functional. This study demonstrates that even the unfinished chimpanzee genome sequence can be useful in identifying modification of genes that are specific to the human lineage and, therefore, could potentially be relevant to the study of the acquisition of human-specific traits.

Amino Acid Sequence↗

The genome sequence of silkworm, Bombyx mori.

We performed threefold shotgun sequencing of the silkworm (Bombyx mori) genome to obtain a draft sequence and establish a basic resource for comprehensive genome analysis. By using the newly developed RAMEN assembler, the sequence data derived from whole-genome shotgun (WGS) sequencing were assembled into 49,345 scaffolds that span a total length of 514 Mb including gaps and 387 Mb without gaps. Because the genome size of the silkworm is estimated to be 530 Mb, almost 97% of the genome has been organized in scaffolds, of which 75% has been sequenced. By carrying out a BLAST search for 50 characteristic Bombyx genes and 11,202 non-redundant expressed sequence tags (ESTs) in a Bombyx EST database against the WGS sequence data, we evaluated the validity of the sequence for elucidating the majority of silkworm genes. Analysis of the WGS data revealed that the silkworm genome contains many repetitive sequences with an average length of <500 bp. These repetitive sequences appear to have been derived from truncated transposons, which are interspersed at 2.5- to 3-kb intervals throughout the genome. This pattern suggests that silkworm may have an active mechanism that promotes removal of transposons from the genome. We also found evidence for insertions of mitochondrial DNA fragments at 9 sites. A search for Bombyx orthologs to Drosophila genes controlling sex determination in the WGS data revealed 11 Bombyx genes and suggested that the sex-determining systems differ profoundly between the two species.

Animals↗

Genomics of the human carnitine acyltransferase genes.

Five genes in the human genome are known to encode different active forms of related carnitine acyltransferases: CPT1A for liver-type carnitine palmitoyltransferase I, CPT1B for muscle-type carnitine palmitoyltransferase I, CPT2 for carnitine palmitoyltransferase II, CROT for carnitine octanoyltransferase, and CRAT for carnitine acetyltransferase. Only from two of these genes (CPT1B and CPT2) have full genomic structures been described. Data from the human genome sequencing efforts now reveal drafts of the genomic structure of CPT1A and CRAT, the latter not being known from any other mammal. Furthermore, cDNA sequences of human CROT were obtained recently, and database analysis revealed a completed bacterial artificial chromosome sequence that contains the entire CROT gene and several exons of the flanking genes P53TG and PGY3. The genomic location of CROT is at chromosome 7q21.1. There is a putative CPT1-like pseudogene in the carnitine/choline acyltransferase family at chromosome 19. Here we give a brief overview of the functional relations between the different carnitine acyltransferases and some of the common features of their genes. We will highlight the phylogenetics of the human carnitine acyltransferase genes in relation to the fungal genes YAT1 and CAT2, which encode cytosolic and mitochondrial/peroxisomal carnitine acetyltransferases, respectively.

Carnitine Acyltransferases↗

Decoding the human genome sequence.

The year 2000 is marked by the production of the sequence of the human genome. A 'working draft' of high quality sequence covering 90% of the genome has been determined and a quarter is in finished form, including the first two completed chromosomes. All sequence data from the project is made freely available to the community via the Internet, for further analysis and exploitation. The challenge which lies ahead is to decipher the information. Knowledge of the human genome sequence will enable us to understand how the genetic information determines the development, structure and function of the human body. We will be able to explore how variations within our DNA sequence cause disease, how they affect our interaction with our environment and ultimately to develop new and effective ways to improve human health.

Conserved Sequence↗

Chicken genome sequence: a centennial gift to poultry genetics.

A draft sequence of the chicken genome will be available by early 2004. This event conveniently marks the start of the second century of poultry genetics, coming 100 years after the use of the chicken to demonstrate Mendelian inheritance in animals by William Bateson. How will the second, post-genomic century of poultry genetics differ from the first? A whole genome shotgun (WGS) approach is being used to obtain the chicken sequence, with the goal of generating approximately six-fold coverage of the genome. Bacterial artificial chromosome (BAC) and fosmid clone end sequences, along with a BAC contig map integrated with genetic linkage and radiation hybrid maps, will form the platform for assembly of the WGS data. Rapid progress in global analysis of chicken gene expression patterns is also being made. Comparative genomics will link these new discoveries to the knowledge base for all other animal species. It's hoped that the genome sequence will also provide common ground on which to unite studies of the chicken as a model species with those aimed at agriculturally-relevant applications. The current status of chicken genomics will be assessed with projections for its near and long term future.

Animals↗

Whole-genome sequencing and analysis of the endophytic fungus Alternaria alternata Y-2 from Leymus chinensis.

To explore the genetic basis and functional potential of beneficial symbiosis between the endophytic fungus Alternaria alternata Y-2 and its host Leymus chinensis, we performed Illumina-based draft whole-genome sequencing and systematic bioinformatic analysis. Although this assembly does not reach telomere-to-telomere completeness, it provides high-quality gene-level information for gene prediction, functional annotation, carbohydrate-active enzyme (CAZyme) identification, and secondary metabolite biosynthetic gene cluster analysis. The final genome size of A. alternata Y-2 was 34,383,676&#xa0;bp with a GC content of 51.0%, containing 12,724 predicted protein-coding genes, 90 tRNAs, and 12 rRNAs. BUSCO assessment showed 98.9% completeness, supporting the high quality of this draft genome. A total of 12,627 genes were successfully annotated in the NCBI NR database, and 17,183 genes were functionally categorized using GO terms. In total, 448 CAZyme genes and 21 secondary metabolite biosynthetic gene clusters were identified, which are potentially involved in lignocellulose degradation, cellular redox homeostasis and biosynthesis of bioactive metabolites. Based on ITS sequence alignment, NR annotation, and phylogenetic analysis of single-copy orthologous genes, the strain was confidently identified as A. alternata. This study firstly reports the draft genome of an endophytic A. alternata strain derived from L. chinensis and provides valuable genetic resources for exploring the endophytic lifestyle, stress tolerance, and bioactive metabolite potential of this fungus.

Alternaria↗

GAMOLA: a new local solution for sequence annotation and analyzing draft and finished prokaryotic genomes.

Laboratories working with draft phase genomes have specific software needs, such as the unattended processing of hundreds of single scaffolds and subsequent sequence annotation. In addition, it is critical to follow the "movement" and the manual annotation of single open reading frames (ORFs) within the successive sequence updates. Even with finished genomes, regular database updates can lead to significant changes in the annotation of single ORFs. In functional genomics it is important to mine data and identify new genetic targets rapidly and easily. Often there is no need for sophisticated relational databases (RDB) that greatly reduce the system-independent access of the results. Another aspect is the internet dependency of most software packages. If users are working with confidential data, this dependency poses a security issue. GAMOLA was designed to handle the numerous scaffolds and changing contents of draft phase genomes in an automated process and stores the results for each predicted ORF in flatfile databases. In addition, annotation transfers, ORF designation tracking, Blast comparisons, and primer design for whole genome microarrays have been implemented. The software is available under the license of North Carolina State University. A website and a downloadable example are accessible under (http://fsweb2.schaub. ncsu.edu/TRKwebsite/index.htm).

Algorithms↗

Spidey: a tool for mRNA-to-genomic alignments.

We have developed a computer program that aligns spliced sequences to genomic sequences, using local alignment algorithms and heuristics to put together a global spliced alignment. Spidey can produce reliable alignments quickly, even when confronted with noise from alternative splicing, polymorphisms, sequencing errors, or evolutionary divergence. We show how Spidey was used to align reference sequences to known genomic sequences and then to the draft human genome, to align mRNAs to gene clusters, and to align mouse mRNAs to human genomic sequence. We compared Spidey to two other spliced alignment programs; Spidey generally performed quite well in a very reasonable amount of time.

Algorithms↗

[Plant genome sequencing: a prelude to the study of its expression].

This report illustrates development of plant sequencing programmes. So far Arabidopsis genome has been completely sequenced and a draft of the rice genome is available. The Arabidopsis programmes stimulated sequencing of EST (expressed sequence tags) from numerous cultivated species thus creating an enormous resource. The major challenge is now to correctly annotate all the genes in Arabidopsis and find out a biological and biochemical function for each one. The availability of EST and genome sequence now allows one to analyse the expression of genes at the level of the whole genome.

Arabidopsis↗

A comparative map of bovine chromosome 19 based on a combination of mapping on a bacterial artificial chromosome scaffold map, a whole genome radiation hybrid panel and the human draft sequence.

We have constructed a medium density physical map of bovine chromosome 19 using a combination of mapping loci on both a bovine bacterial artificial chromosome (BAC) scaffold map and a whole genome radiation hybrid (WGRH) panel. The resulting map contains 70 loci spanning the length of bovine chromosome 19. Three contiguous groups of BACs were identified on the basis of multiple loci mapping to individual BAC clones. Bovine chromosome 19 was found in this study to be comprised almost entirely from regions of human chromosome 17, with a small region putatively assigned to human chromosome 10. Fourteen breakpoints between the bovine and human chromosomes were detected, with a possibility of five more based on ordering of the WGRH map.

Animals↗

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500&#xa0;m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed >&#x2009;99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071&#x1d40; (=&#x2009;ATCC 10145&#x1d40;), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33&#xa0;Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8&#xa0;kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~&#x2009;22&#xa0;kb, ~&#x2009;17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family↗

ECR Browser: a tool for visualizing and accessing data from comparisons of multiple vertebrate genomes.

With an increasing number of vertebrate genomes being sequenced in draft or finished form, unique opportunities for decoding the language of DNA sequence through comparative genome alignments have arisen. However, novel tools and strategies are required to accommodate this large volume of genomic information and to facilitate the transfer of predictions generated by comparative sequence alignment to researchers focused on experimental annotation of genome function. Here, we present the ECR Browser, a tool that provides easy and dynamic access to whole genome alignments of human, mouse, rat and fish sequences. This web-based tool (http://ecrbrowser.dcode.org) provides the starting point for discovery of novel genes, identification of distant gene regulatory elements and prediction of transcription factor binding sites. The genome alignment portal of the ECR Browser also permits fast and automated alignments of any user-submitted sequence to the genome of choice. The interconnection of the ECR Browser with other DNA sequence analysis tools creates a unique portal for studying and exploring vertebrate genomes.

Animals↗

Impact of human genome sequencing for in silico target discovery.

The year 2000 stands as a landmark in modern biology: the first draft of the human genome sequence has been completed. For the pharmaceutical industry, this achievement provides tremendous opportunities because the genomic sequence exposes all human drug targets for therapeutic intervention. The challenge for the pharmaceutical companies is to exploit this definitive resource for the identification of potential molecular targets, rapid characterization of their function and validation of their involvement in disease pathology. Bioinformatics approaches provide increasingly crucial tools to systematically support this exploratory target drug discovery activity.

Journal Article↗

[Genomic medicine].

The International Human Genome Sequencing Consortium announced they have got a draft sequence of the human genome that covers about 94% of the human genome in February, 2001. This statement implies that we will soon have the complete human genome sequence and that we will be able to use genomic information in every field including medicine, drug production as well as prophylaxis of many diseases. Two major advances in technology have made it possible to apply genomic information for medicine: microarray technology and high-throughput sequencing technology. Microarray provides us a new method to classify various diseases on the basis of gene expression and high-throughput sequencing enables us to draw high-resolution SNPs map.

Chromosome Mapping↗

Computational comparison of human genomic sequence assemblies for a region of chromosome 4.

Much of the available human genomic sequence data exist in a fragmentary draft state following the completion of the initial high-volume sequencing performed by the International Human Genome Sequencing Consortium (IHGSC) and Celera Genomics (CG). We compared six draft genome assemblies over a region of chromosome 4p (D4S394-D4S403), two consecutive releases by the IHGSC at University of California, Santa Cruz (UCSC), two consecutive releases from the National Centre for Biotechnology Information (NCBI), the public release from CG, and a hybrid assembly we have produced using IHGSC and CG sequence data. This region presents particular problems for genomic sequence assembly algorithms as it contains a large tandem repeat and is sparsely covered by draft sequences. The six assemblies differed both in terms of their relative coverage of sequence data from the region and in their estimated rates of misassembly. The CG assembly method attained the lowest level of misassembly, whereas NCBI and UCSC assemblies had the highest levels of coverage. All assemblies examined included <60% of the publicly available sequence from the region. At least 6% of the sequence data within the CG assembly for the D4S394-D4S403 region was not present in publicly available sequence data. We also show that even in a problematic region, existing software tools can be used with high-quality mapping data to produce genomic sequence contigs with a low rate of rearrangements.

Chromosomes, Human, Pair 4↗

Strain-specific genomic regions of Ruminococcus flavefaciens FD-1 as revealed by combinatorial random-phase genome sequencing and suppressive subtractive hybridization.

Two closely related strains of the Gram-positive, cellulolytic ruminal bacterium Ruminococcus flavefaciens were compared at the genomic level by suppressive subtractive hybridization. The two strains investigated in this study differ by 1.94% in their respective 16S rDNA genes. Three hundred and eighty-four PCR-amplified products were cloned and then screened for their strain identity by dot blot hybridization. Based on redundancy percentages of the clones sequenced, 9.5% of the genome of the R. flavefaciens FD-1 strain is not present in the JM1 strain. The majority of identities of individual cloned subtracted products (642 bp average length) bore no relation to deposited sequences in GenBank (42% of the subtracted library), whereas of those with putative assigned functions 7% are loosely associated with fibre-degradation, 6% with insertion elements, transposons and phage-like ORFs, 5% with cell membrane associated proteins and 3% with signal transduction. Subtracted sequences were then supplemented with the draft (2 x coverage) genome sequence of R. flavefaciens FD-1 to indicate potential regions of rearrangement within the genome, including a novel insertion sequence.

Base Sequence↗

Much ado about bacteria-to-vertebrate lateral gene transfer.

When the International Human Genome Sequencing Consortium (IHGSC) published its draft of the human genome in February 2001, several genes were identified as possible bacteria-to-vertebrate transfers (BVTs). These genes were identified by their highly significant sequence similarity to bacterial genes in BLAST searches, and by their lack of matches among non-vertebrate eukaryote genes. Many were later rejected as BVTs by several methods, including recovery of probable orthologs from the genomes of incompletely sequenced eukaryotes. Whereas the BVT issue has received considerable attention, there has been no compilation of all potential BVTs considered to date, nor any proposal of a single comprehensive method for rigorously establishing the veracity of a putative BVT. In reviewing the work to date, we list all of the proteins examined and propose systematic tests to investigate whether a vertebrate gene proposed as a BVT is indeed of bacterial origin. We use the proposed strategy to test--and reject--one of the BVTs from the original IHGSC list.

Animals↗