PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Human-zebrafish non-coding conserved elements act in vivo to regulate transcription.

Whole genome comparisons of distantly related species effectively predict biologically important sequences--core genes and cis-acting regulatory elements (REs)--but require experimentation to verify biological activity. To examine the efficacy of comparative genomics in identification of active REs from anonymous, non-coding (NC) sequences, we generated a novel alignment of the human and draft zebrafish genomes, and contrasted this set to existing human and fugu datasets. We tested the transcriptional regulatory potential of candidate sequences using two in vivo assays. Strict selection of non-genic elements which are deeply conserved in vertebrate evolution identifies 1744 core vertebrate REs in human and two fish genomes. We tested 16 elements in vivo for cis-acting gene regulatory properties using zebrafish transient transgenesis and found that 10 (63%) strongly modulate tissue-specific expression of a green fluorescent protein reporter vector. We also report a novel quantitative enhancer assay with potential for increased throughput based on normalized luciferase activity in vivo. This complementary system identified 11 (69%; including 9 of 10 GFP-confirmed elements) with cis-acting function. Together, these data support the utility of comparative genomics of distantly related vertebrate species to identify REs and provide a scaleable, in vivo quantitative assay to define functional activity of candidate REs.

Animals↗

Composition and evolution of the V2r vomeronasal receptor gene repertoire in mice and rats.

Pheromones are chemicals produced and detected by conspecifics to elicit social/sexual physiological and behavioral responses, and they are perceived primarily by the vomeronasal organ (VNO) in terrestrial vertebrates. Two large superfamilies of G protein-coupled receptors, V1rs and V2rs, have been identified as pheromone receptors in vomeronasal sensory neurons. Based on a computational analysis of the mouse and rat genome sequences, we report the first global draft of the V2r gene repertoire, composed of approximately 200 genes and pseudogenes. Rodent V2rs are subject to rapid gene births/deaths and accelerated amino acid substitutions, likely reflecting the species-specific nature of pheromones. Vertebrate V2rs appear to have originated twice prior to the emergence of the VNO in ancestral tetrapods, explaining seemingly inconsistent observations among different V2rs. The identification of the entire V2r repertoire opens the door to genomic-level studies of the structure, function, and evolution of this diverse group of sensory receptors.

Amino Acid Sequence↗

A novel gene family NBPF: intricate structure generated by gene duplications during primate evolution.

Partial and complete genome duplications occurred during evolution and resulted in the creation of new genes and gene families. We identified a novel and intricate human gene family located primarily in regions of segmental duplications on human chromosome 1. We named it NBPF, for neuroblastoma breakpoint family, because one of its members is disrupted by a chromosomal translocation in a neuroblastoma patient. The NBPF genes have a repetitive structure with high intragenic and intergenic sequence similarity in both coding and noncoding regions. These similarities might expose these genomic regions to illegitimate recombination, resulting in structural variation in the NBPF genes. The encoded proteins contain a highly conserved domain of unknown function, which we have named the NBPF repeat. In silico analysis combined with the isolation of multiple full-length cDNA clones showed that several members of this gene family are abundantly expressed in a large variety of tissues and cell lines. Strikingly, no discernable orthologues could be identified in the completed genomes of fruit fly, nematode, mouse, or rat, but sequences with low homology could be isolated from the draft canine and bovine genomes. Interestingly, this gene family shows primate-specific duplications that result in species-specific arrays of NBPF homologous sequences. Overall, this novel NBPF family reflects the continuous evolution of primate genomes that resulted in large physiological differences, and its potential role in this process is discussed.

Amino Acid Sequence↗

The myosin light chain kinase gene is not duplicated in mouse: partial structure and chromosomal localization of Mylk.

The gene encoding myosin light chain kinase (MYLK) is duplicated on human chromosome 3 (HSA3; 3p13;3q21) and on a chromosome with conserved synteny to HSA3 in most non-human primate species. In human, the functional copy resides on 3q21, whereas the 3p13 site contains a pseudogene. To trace the origin of the duplication, we characterized the mouse gene Mylk. A single sequence corresponding to the functional Mylk was detected. We sequenced a 180-kb bacterial artificial chromosome clone containing the 24 first exons of Mylk; the complete mouse gene is expected to span >200 kb. Comparisons with the draft of the human genome revealed that the sequence and structure of MYLK are conserved in mammals. Fluorescence in situ hybridization (FISH) analysis indicated that the mouse gene localizes to a single site on chromosome 16B4-B5, a region with conserved synteny with HSA3q. Our study provides information on both the structure and the evolution of MYLK in mammals and suggests that it was duplicated after the divergence of rodents and primates.

Amino Acid Sequence↗

Medical education in the 'postgenomic era'.

The sequence of the human genome is very nearly in hand; the first draft has been completed, and the finished sequence will be available years ahead of schedule. Already, advances in medical genetics have affected the day-to-day practice of medicine by providing more powerful approaches to diagnosis of genetic disorders and cancer. But the full impact of the integration of genetics into medical practice lies before us.

Education, Medical↗

Optimization of protoplast based DNA isolation and genome analysis in a gamma-irradiated Aspergillus niger mutant strain.

Aspergillus niger is an important industrial fungus widely used for citric acid production and a range of biotechnological applications. In this study, a protoplast-based DNA isolation protocol was optimized for a gamma-irradiated A. niger AN-L103_M1 mutant strain, followed by whole-genome sequencing and functional genome analysis. Protoplast yield was strongly influenced by enzyme concentration and the molarity of the osmotic stabilizer. The highest yield was achieved at an enzyme concentration of 50&#xa0;mg/mL (2.487&#x2009;&#xb1;&#x2009;0.04&#x2009;&#xd7;&#x2009;10&#x2078; cells/mL) and 0.8&#xa0;M KCl (2.550&#x2009;&#xb1;&#x2009;0.06&#x2009;&#xd7;&#x2009;10&#x2078; cells/mL), with both factors showing significant effects (p&#x2009;<&#x2009;0.0001) in GraphPad Prism 11.0.0. Whole-genome sequencing performed using an Illumina NovaSeq 6000 platform yielded a 37.06&#xa0;Mb draft genome assembled into 537 contigs, with an N50 of 363,084&#xa0;bp and a GC content of 48.2%. BUSCO 14 analysis showed high completeness (97.95% complete BUSCOs). Functional annotation and KEGG pathway mapping identified genes involved in glycolysis, the tricarboxylic acid cycle, and citrate biosynthesis, while biosynthetic gene cluster analysis revealed diverse potential for secondary metabolite production. These findings provide an optimized workflow for protoplast-based DNA isolation and genome-scale functional analysis in A. niger, proposing a basis for future comparative genomics, transformation studies, and experimentally validated metabolic engineering.

Aspergillus niger↗

A case for evolutionary genomics and the comprehensive examination of sequence biodiversity.

Comparative analysis is one of the most powerful methods available for understanding the diverse and complex systems found in biology, but it is often limited by a lack of comprehensive taxonomic sampling. Despite the recent development of powerful genome technologies capable of producing sequence data in large quantities (witness the recently completed first draft of the human genome), there has been relatively little change in how evolutionary studies are conducted. The application of genomic methods to evolutionary biology is a challenge, in part because gene segments from different organisms are manipulated separately, requiring individual purification, cloning, and sequencing. We suggest that a feasible approach to collecting genome-scale data sets for evolutionary biology (i.e., evolutionary genomics) may consist of combination of DNA samples prior to cloning and sequencing, followed by computational reconstruction of the original sequences. This approach will allow the full benefit of automated protocols developed by genome projects to be realized; taxon sampling levels can easily increase to thousands for targeted genomes and genomic regions. Sequence diversity at this level will dramatically improve the quality and accuracy of phylogenetic inference, as well as the accuracy and resolution of comparative evolutionary studies. In particular, it will be possible to make accurate estimates of normal evolution in the context of constant structural and functional constraints (i.e., site-specific substitution probabilities), along with accurate estimates of changes in evolutionary patterns, including pairwise coevolution between sites, adaptive bursts, and changes in selective constraints. These estimates can then be used to understand and predict the effects of protein structure and function on sequence evolution and to predict unknown details of protein structure, function, and functional divergence. In order to demonstrate the practicality of these ideas and the potential benefit for functional genomic analysis, we describe a pilot project we are conducting to simultaneously sequence large numbers of vertebrate mitochondrial genomes.

Animals↗

ToxoDB: accessing the Toxoplasma gondii genome.

ToxoDB (http://ToxoDB.org) provides a genome resource for the protozoan parasite Toxoplasma gondii. Several sequencing projects devoted to T. gondii have been completed or are in progress: an EST project (http://genome.wustl.edu/est/index.php?toxoplasma=1), a BAC clone end-sequencing project (http://www.sanger.ac.uk/Projects/T_gondii/) and an 8X random shotgun genomic sequencing project (http://www.tigr.org/tdb/e2k1/tga1/). ToxoDB was designed to provide a central point of access for all available T. gondii data, and a variety of data mining tools useful for the analysis of unfinished, un-annotated draft sequence during the early phases of the genome project. In later stages, as more and different types of data become available (microarray, proteomic, SNP, QTL, etc.) the database will provide an integrated data analysis platform facilitating user-defined queries across the different data types.

Animals↗

A genetic and structural analysis of the N-glycosylation capabilities.

The recent draft sequencing of the rice (Oryza sativa) genome has enabled a genetic analysis of the glycosylation capabilities of an agroeconomically important group of plants, the monocotyledons. In this study, we have not only identified genes putatively encoding enzymes involved in N-glycosylation, but have examined by MALDI-TOF MS the structures of the N-glycans of rice and other monocotyledons (maize, wheat and dates; Zea mays, Triticum aestivum and Phoenix dactylifera); these data show that within the plant kingdom the types of N-glycans found are very similar between monocotyledons, dicotyledons and gymnosperms. Subsequently, we constructed expression vectors for the key enzymes forming plant-typical structures in rice, N-acetylglucosaminyltransferase I (GlcNAc-TI; EC 2.4.1.101), core alpha1,3-fucosyltransferase (FucTA; EC 2.4.1.214) and beta1,2-xylosyltransferase (EC 2.4.2.38) and successfully expressed them in Pichia pastoris. Rice GlcNAc-TI, FucTA and xylosyltransferase are therefore the first monocotyledon glycosyltransferases involved in N-glycan biosynthesis to be characterised in a recombinant form.

Amino Acid Sequence↗

Inventory and comparative analysis of rice and Arabidopsis ATP-binding cassette (ABC) systems.

ATP-binding cassette (ABC) proteins constitute a large superfamily found in all kingdoms of living organisms. The recent completion of two draft sequences of the rice (Oryza sativa) genome allowed us to analyze and classify its ABC proteins and to compare to those in Arabidopsis thaliana. We identified a similar number of ABC proteins in rice and Arabidopsis (121 versus 120), despite the rice genome being more than three times the size of Arabidopsis. Both Arabidopsis and rice have representative members in all seven major subfamilies of ABC ATPases (A to G) commonly found in eukaryotes. This comparative analysis allowed the detection of 29 potential orthologous sequences in Arabidopsis and rice. However, plant share with prokaryotes a specific set of ABC systems that is not detected in animals. These ABC systems might be inherited from the cyanobacterial ancestor of chloroplasts. The present work provides the first complete inventory of rice ABC proteins and an updated inventory of those proteins in Arabidopsis.

ATP-Binding Cassette Transporters↗

Novel approaches for identifying genes regulating lymphocyte development and function.

The draft sequence of the human and mouse genomes provides an unparalleled opportunity for understanding the genetic control of immune-cell development. Strategies can begin with a gene sequence and pursue a putative immune-system function by employing mRNA-expression profiling or creating gene knockouts in embryonic stem cells. The latter can be produced by utilising the Cre/Lox system, a tetracycline operon, a gene-trap method or chemical mutagenesis. Alternatively, mutant phenotypes (derived using the mutagen ethylnitrosourea) can be traced back to gene sequences.

Animals↗

Methods for comparing sources of strand compositional asymmetry in microbial chromosomes.

Significant compositional biases in bacterial chromosomes have been explained by replication- and transcription-coupled repair mechanisms, the latter causing GC skew to indicate the direction of replication when gene polarity is correspondingly entrained. Correlations between indicators of replication direction, skew, and transcription polarity are computed for the complete nucleotide sequences of 20 microbial chromosomes and interpreted through statistical tests. A second quantitative method, previously applied to the first complete draft of the Escherichia coli K12 genome, characterizes the sequences by average skew and net skew due to replication. These methods generally agree in finding the coexistence of replication- and translation-coupled effects and in identifying atypical sequences in which one influence is clearly dominant. The replication-dominated class is exemplified by two chlamydial sequences and the transcription-dominated class by three archaea. The preference for leading-strand transcription in two mycoplasmas is stronger than the skew implies. These concordant methods provide an objective framework for comparing sources of strand compositional asymmetry and interpreting skew diagrams.

Base Composition↗

Assembly of the working draft of the human genome with GigAssembler.

The data for the public working draft of the human genome contains roughly 400,000 initial sequence contigs in approximately 30,000 large insert clones. Many of these initial sequence contigs overlap. A program, GigAssembler, was built to merge them and to order and orient the resulting larger sequence contigs based on mRNA, paired plasmid ends, EST, BAC end pairs, and other information. This program produced the first publicly available assembly of the human genome, a working draft containing roughly 2.7 billion base pairs and covering an estimated 88% of the genome that has been used for several recent studies of the genome. Here we describe the algorithm used by GigAssembler.

Algorithms↗

Integration of microsatellite-based genetic maps for the turkey (Meleagris gallopavo).

Integration of turkey genetic maps and their associated markers is essential to increase marker density in support of map-based genetic studies. The objectives of this study were to integrate 2 microsatellite-based turkey genetic maps--the Roslin map and the University of Minnesota (UMN) map--by genotyping markers from the Roslin study on the mapping families of the UMN study. A total of 279 markers was tested, and 240 were subsequently screened for polymorphisms in the UMN/Nicholas Turkey Breeding Farms (NTBF) mapping families. Of the 240 markers, 89 were genetically informative and were used for genotyping the F2 offspring. Significant genetic linkages (log of odds > 3.0) were found for 84 markers from the Roslin study. BLASTn comparison of marker sequences with the draft assembly of the chicken genome found 263 significant matches. The combination of genetic and in silico mapping allowed for the alignment of all linkage groups of the Roslin map with those of the UMN map. With the addition of the markers from the Roslin map, 438 markers are now genetically linked in the UMN/NTBF families, and more than 1700 turkey sequences have now been assigned to likely positions in the chicken-genome sequence.

Animals↗

Identification of a fatty acid Delta11-desaturase from the microalga Thalassiosira pseudonana.

A set of genomic DNA sequences putatively encoding front-end desaturases were identified by in silico analysis of the draft genome of the marine microalga Thalassiosira pseudonana. Among these candidate genes, an open reading frame named TpdesN was found to be full-length, intronless, and constitutively expressed during cell cultivation. The predicted amino acid sequence of the corresponding protein, TpDESN, exhibited typical features of desaturases involved in the production of polyunsaturated fatty acids (PUFAs) in algae, i.e. a cytochrome b5-like domain at the N-terminus and three conserved histidine-rich motifs in the desaturase domain. Expression of TpDESN in Saccharomyces cerevisiae revealed that this enzyme was not involved in PUFA synthesis, but specifically desaturated palmitic acid 16:0 to 16:1Delta11. To our knowledge, until this report, Delta11-desaturase activity had only been detected in insect cells.

Amino Acid Motifs↗

Identification of a novel Bardet-Biedl syndrome protein, BBS7, that shares structural features with BBS1 and BBS2.

Bardet-Biedl syndrome (BBS) is a genetically heterogeneous disorder, the primary features of which include obesity, retinal dystrophy, polydactyly, hypogenitalism, learning difficulties, and renal malformations. Conventional linkage and positional cloning have led to the mapping of six BBS loci in the human genome, four of which (BBS1, BBS2, BBS4, and BBS6) have been cloned. Despite these advances, the protein sequences of the known BBS genes have provided little or no insight into their function. To delineate functionally important regions in BBS2, we performed phylogenetic and genomic studies in which we used the human and zebrafish BBS2 peptide sequences to search dbEST and the translation of the draft human genome. We identified two novel genes that we initially named "BBS2L1" and "BBS2L2" and that exhibit modest similarity with two discrete, overlapping regions of BBS2. In the present study, we demonstrate that BBS2L1 mutations cause BBS, thereby defining a novel locus for this syndrome, BBS7, whereas BBS2L2 has been shown independently to be BBS1. The motif-based identification of a novel BBS locus has enabled us to define a potential functional domain that is present in three of the five known BBS proteins and, therefore, is likely to be important in the pathogenesis of this complex syndrome.

Adaptor Proteins, Signal Transducing↗

The phusion assembler.

The Phusion assembler has assembled the mouse genome from the whole-genome shotgun (WGS) dataset collected by the Mouse Genome Sequencing Consortium, at ~7.5x sequence coverage, producing a high-quality draft assembly 2.6 gigabases in size, of which 90% of these bases are in 479 scaffolds. For the mouse genome, which is a large and repeat-rich genome, the input dataset was designed to include a high proportion of paired end sequences of various size selected inserts, from 2-200 kbp lengths, into various host vector templates. Phusion uses sequence data, called reads, and information about reads that share common templates, called read pairs, to drive the assembly of this large genome to highly accurate results. The preassembly stage, which clusters the reads into sensible groups, is a key element of the entire assembler, because it permits a simple approach to parallelization of the assembly stage, as each cluster can be treated independent of the others. In addition to the application of Phusion to the mouse genome, we will also present results from the WGS assembly of Caenorhabditis briggsae sequenced to about 11x coverage. The C. briggsae assembly was accessioned through EMBL, http://www.ebi.ac.uk/services/index.html, using the series CAAC01000001-CAAC01000578, however, the Phusion mouse assembly described here was not accessioned. The mouse data was generated by the Mouse Genome Sequencing Consortium. The C. briggsae sequence was generated at The Wellcome Trust Sanger Institute and the Genome Sequencing Center, Washington University School of Medicine.

Animals↗