PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

WindowMasker: window-based masker for sequenced genomes.

MOTIVATION: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes. RESULTS: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis. AVAILABILITY: WM is included in the NCBI C++ toolkit. The source code for the entire toolkit is available at ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools++/CURRENT/. Once the toolkit source is unpacked, the instructions for building WindowMasker application in the UNIX environment can be found in file src/app/winmasker/README.build. SUPPLEMENTARY INFORMATION: Supplementary data are available at ftp://ftp.ncbi.nlm.nih.gov/pub/agarwala/windowmasker/windowmasker_suppl.pdf

Algorithms↗

Canonical TTAGG-repeat telomeres and telomerase in the honey bee, Apis mellifera.

The draft assembly of the honey bee Apis mellifera genome sequence reveals that the 17 centromeric-distal telomeres are of a simple, shared, and canonical structure, with 3-4 kb of a unique subtelomeric sequence, followed by several kilobases of TTAGG or variant telomeric repeats. This simple subtelomeric structure differs from the centromeric-proximal telomeres on the short arms of the 15 acrocentric chromosomes, which are apparently composed primarily of the 176-bp AluI tandem repeat. This dichotomy between the distal and proximal telomeres may involve differential participation of the telomeres of the 15 acrocentric chromosomes in the Rabl configuration after mitosis and the chromosome bouquet in meiotic prophase I. As expected from the presence of canonical TTAGG telomeric repeats, we identified a candidate telomerase gene in the bee, as well as the silkmoth Bombyx mori and the flour beetle Tribolium castaneum.

Animals↗

Reference based annotation with GeneMapper.

We introduce GeneMapper, a program for transferring annotations from a well annotated genome to other genomes. Drawing on high quality curated annotations, GeneMapper enables rapid and accurate annotation of newly sequenced genomes and is suitable for both finished and draft genomes. GeneMapper uses a profile based approach for mapping genes into multiple species, improving upon the standard pairwise approach. GeneMapper is freely available for academic use.

Algorithms↗

Transposable elements create distinct genomic niches for effector evolution among Magnaporthe oryzae lineages.

BACKGROUND: Plant-pathogen interactions are characterized by evolutionary arms races. At the molecular level, fungal effectors can target important plant functions, while plants evolve to improve effector recognition. Rapid evolution in genes encoding effectors can be facilitated by transposable elements (TEs). In Magnaporthe oryzae, the causal agent of blast disease in several cereals and grasses, TEs play important roles in chromosomal evolution as well as the gain or loss of effector genes in host specialized lineages. However, a global understanding of TE dynamics driving effector evolution at population scale and across lineages is lacking. RESULTS: Here, we focus on 16 AVR effector loci assessed across a global sampling of 11 reference genomes and 447 newly generated draft genome assemblies from publicly available short-read sequencing data across all major M. oryzae lineages and outgroups. We classified each effector based on evidence for duplication, deletion and translocation processes among lineages. Next, we determined AVR gain and loss dynamics across lineages allowing for a broad categorization of effector dynamics. Each AVR was integrated in a distinct genomic niche determined by the TE activity profile contributing to the diversification at the locus. We quantified TE contributions to effector niches and found that TE identity helped diversify AVR loci. We used the large genomic dataset to recapitulate the evolution of the rice blast AVR1-CO39 locus. CONCLUSIONS: Taken together, our work demonstrates how TE dynamics are an integral component of M. oryzae effector evolution, likely facilitating escape from host recognition. In-depth tracking of effector loci is a valuable tool to predict the durability of host resistance.

Ascomycota↗

Fluorescent in situ hybridization to ascidian chromosomes.

The draft genome of the ascidian Ciona intestinalis has been sequenced. Mapping of the genome sequence to the Ciona 14 haploid chromosomes is essential for future studies of the genome-wide control of gene expression in this basal chordate. Here we describe an efficient protocol for fluorescent in situ hybridization for mapping genes to the Ciona chromosomes. We demonstrate how the locations of two BAC clones can be mapped relative to each other. We also show that this method is efficient for coupling two so-far independent scaffolds into one longer scaffold when two BAC clones represent sequences located at either end of the two scaffolds.

Animals↗

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗

Mining the human genome using microarrays of open reading frames.

To test the hypothesis that the human genome project will uncover many genes not previously discovered by sequencing of expressed sequence tags (ESTs), we designed and produced a set of microarrays using probes based on open reading frames (ORFs) in 350 Mb of finished and draft human sequence. Our approach aims to identify all genes directly from genomic sequence by querying gene expression. We analysed genomic sequence with a suite of ORF prediction programs, selected approximately one ORF per gene, amplified the ORFs from genomic DNA and arrayed the amplicons onto treated glass slides. Of the first 10,000 arrayed ORFs, 31% are completely novel and 29% are similar, but not identical, to sequences in public databases. Approximately one-half of these are expressed in the tissues we queried by microarray. Subsequent verification by other techniques confirmed expression of several of the novel genes. Expressed sequence tags (ESTs) have yielded vast amounts of data, but our results indicate that many genes in the human genome will only be found by genomic sequencing.

Cell Line↗

[The canine genome: alternative model for the functional analysis of mammalian genes].

The pace of genome sequencing has been tremendously accelerated during the last few years leading to the determination of dozens of entire bacterial genome sequences in addition to several eukaryotic genome sequences and to the publication in 2000 of a draft of the human one. Nowadays scientists have to face a new challenge that corresponds to the elucidation of the function(s) of the thousands of genes uncovered by sequencing. Obviously this task will necessitate a large panel of methodologies. Since its domestication, dog has been the subject of intense breeding and selection practices that result in the creation of many breeds that differ one from the others by a huge variation in shape, size, coat colour, aptitude. Unfortunately these breeding practices along the selection of specific alleles governing those characters have co-selected various deleterious or morbid alleles and nowadays most of the canine breeds suffers from many different diseases of genetic origin. In addition many breeds have developed susceptibility toward many diseases very often similar to those affecting humans such as cancers, heart diseases, allergies.... In this paper we present arguments in favour of the utilisation of the canine model to sort out through linkage disequilibrium studies the phenotype/genotype relationship as an aid to understand the function(s) of the thousands of genes uncovered by sequencing.

Animals↗

The sarcomeric myosin heavy chain gene family in the dog: analysis of isoform diversity and comparison with other mammalian species.

Sarcomeric myosin heavy chains (MyHC) are the major contractile proteins of cardiac and skeletal muscles and belong to class II MyHC. In this study the sequences of nine sarcomeric MyHC isoforms were obtained by combining assembled contigs of the dog genome draft available in the NCBI database. With this information available the dog becomes the second species, after human, for which the sequences of all members of the sarcomeric MyHC gene family are identified. The newly determined sequences of canine MyHC isoforms were aligned with their orthologs in mammals, forming a set of 38 isoforms, to search for the molecular features that determine the structural and functional specificity of each type of isoform. In this way the structural motifs that allow identification of each isoform and are likely determinants of functional properties were identified in six specific regions (surface loop 1, loop 2, loop 3, converter, MLC binding region, and S2 proximal segment).

Amino Acid Sequence↗

Mouse models of Down syndrome: how useful can they be? Comparison of the gene content of human chromosome 21 with orthologous mouse genomic regions.

With an incidence of approximately 1 in 700 live births, Down syndrome (DS) remains the most common genetic cause of mental retardation. The phenotype is assumed to be due to overexpression of some number of the >300 genes encoded by human chromosome 21. Mouse models, in particular the chromosome 16 segmental trisomies, Ts65Dn and Ts1Cje, are indispensable for DS-related studies of gene-phenotype correlations. Here we compare the updated gene content of the finished sequence of human chromosome 21 (364 genes and putative genes) with the gene content of the homologous mouse genomic regions (291 genes and putative genes) obtained from annotation of the public sector C57Bl/6 draft sequence. Annotated genes fall into one of three classes. First, there are 170 highly conserved, human/mouse orthologues. Second, there are 83 minimally conserved, possible orthologues. Included among the conserved and minimally conserved genes are 31 antisense transcripts. Third, there are species-specific genes: 111 spliced human transcripts show no orthologues in the syntenic mouse regions although 13 have homologous sequences elsewhere in the mouse genomic sequence, and 38 spliced mouse transcripts show no identifiable human orthologues. While these species-specific genes are largely based solely on spliced EST data, a majority can be verified in RNA expression experiments. In addition, preliminary data suggest that many human-specific transcripts may represent a novel class of primate-specific genes. Lastly, updated functional annotation of orthologous genes indicates genes encoding components of several cellular pathways are dispersed throughout the orthologous mouse chromosomal regions and are not completely represented in the Down syndrome segmental mouse models. Together, these data point out the potential for existing mouse models to produce extraneous phenotypes and to fail to produce DS-relevant phenotypes.

Alternative Splicing↗

Oxford Nanopore Sequencing of Clinical DNA for Identification and Comparative Genomic Analysis of Erysipelothrix piscisicarius.

The genus Erysipelothrix comprises facultative anaerobic, nonspore-forming, gram-positive bacteria that can cause skin infections and severe diseases such as septicemia and endocarditis in humans. Although E. rhusiopathiae is the primary pathogen, other species may also be involved, necessitating accurate identification. However, 16S rDNA sequencing lacks sufficient resolution to differentiate among Erysipelothrix species. In this study, we used Oxford Nanopore Technology (ONT) to directly sequence low-quality DNA extracted from heart valve tissue of a 66-year-old female patient with a fatal case of septicemia and aortic endocarditis. In contrast to 16S rDNA Illumina sequencing and matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS), which incorrectly identified the pathogen as E. rhusiopathiae, direct sequencing via ONT precisely identified E. piscisicarius as the cause of infection. About 1.47 Mb genome was retrieved from nanopore direct sequencing. Within the E. piscisicarius genome, we detected genes associated with virulence. Phylogenetic analysis showed that our strain clustered with a human-derived E. piscisicarius strain from China and swine-derived strains from Brazil. In conclusion, this study demonstrated that ONT can be used to sequence low-quality DNA extracted directly from patient specimens, obtain a draft bacterial genome, and reliably distinguish between pathogenic species.

Aged↗

Organization and evolution of a gene-rich region of the mouse genome: a 12.7-Mb region deleted in the Del(13)Svea36H mouse.

Del(13)Svea36H (Del36H) is a deletion of approximately 20% of mouse chromosome 13 showing conserved synteny with human chromosome 6p22.1-6p22.3/6p25. The human region is lost in some deletion syndromes and is the site of several disease loci. Heterozygous Del36H mice show numerous phenotypes and may model aspects of human genetic disease. We describe 12.7 Mb of finished, annotated sequence from Del36H. Del36H has a higher gene density than the draft mouse genome, reflecting high local densities of three gene families (vomeronasal receptors, serpins, and prolactins) which are greatly expanded relative to human. Transposable elements are concentrated near these gene families. We therefore suggest that their neighborhoods are gene factories, regions of frequent recombination in which gene duplication is more frequent. The gene families show different proportions of pseudogenes, likely reflecting different strengths of purifying selection and/or gene conversion. They are also associated with relatively low simple sequence concentrations, which vary across the region with a periodicity of approximately 5 Mb. Del36H contains numerous evolutionarily conserved regions (ECRs). Many lie in noncoding regions, are detectable in species as distant as Ciona intestinalis, and therefore are candidate regulatory sequences. This analysis will facilitate functional genomic analysis of Del36H and provides insights into mouse genome evolution.

Animals↗

[Rapid identification of human testis spermatocyte apoptosis-related gene, TSARG2, by nested PCR and draft human genome searching].

Cloning apoptosis-related novel genes is a key to further understanding of apoptosis mechanism and the biology process of germ cells, and is of momentous significance on clarifying physiological and pathological process of spermatogenesis. To rapidly attain human novel gene full-length cDNA sequence, the gene-specific primers and the vector-specific primers were designed for nested PCR, and draft human genome searching was performed to rapidly identify the TSARG2 (GenBank accession number AY040204) 5' end from a human testis cDNA library, by using a cDNA fragment (GenBank accession number BE644542) as an electronic probe, which was significantly changed in cryptorchidism and represented a novel gene. Furthermore, a mouse homologue of this gene was identified (GenBank accession number AF395083) by lab on-line. TSARG2 with a 1 233 bp length was composed of 6 exons and spanned about 115 kb of genomic DNA, The putative protein encoded by this gene was 305 amino acid with a theoretical molecular weight of 34 751 dalton and did not share significant homology with any known protein in databases. TSARG2 was expressed in many tissues and mapped to chromosome 4q33-34.1 by database analyses. Therefore, we propose that nested-PCR and draft human genome searching are rapid, sensitive, accurate and efficient method for isolating gene 5' end, even full-length gene from cDNA library.

Amino Acid Sequence↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

A second gene for peroxisomal HMG-CoA reductase? A genomic reassessment.

HMG-CoA reductase (HMGCR) catalyzes the conversion of HMG-CoA to mevalonate, the rate-limiting step of eukaryotic isoprenoid biosynthesis, and is the main target of cholesterol-lowering drugs. The classical form of the enzyme is a transmembrane-protein anchored to the endoplasmic reticulum. However, during the last years several lines of evidence pointed to the existence of a second isoform of HMGCR localized in peroxisomes, where mevalonate is converted further to farnesyl diphosphate. This finding is relevant for our understanding of the complex regulation and compartmentalization of the cholesterogenic pathway. Here we review experimental evidence suggesting that the peroxisomal activity might be due to a second HMGCR gene in mammals. We then present a comprehensive analysis of completely sequenced eukaryotic genomes, as well as the human and mouse genome drafts. Our results provide evidence for a large number of independent duplications of HMGCR in all eukaryotic kingdoms, but not for a second gene in mammals. We conclude that the peroxisomal HMGCR activity in mammals is due to alternative targeting of the ER enzyme to peroxisomes by an as yet uncharacterized mechanism.

Animals↗

The human genome project: implications for the endocrinologist.

The sequencing of the human genome is a major achievement of our time. This article reviews the process and current status of the working draft sequence, ways to predict genes and assign function, and conclusions for human biology. Gene density is uneven and related to chromosome banding patterns, and the estimate of approximately 30,000 genes is lower than expected. Genetic maps for men and women differ from each other and from the physical map. Single nucleotide polymorphisms occur at an average spacing of 1 kb. Human populations are 99.99% identical, and most sequences are shared between people from different continents. To illustrate the tools for accessing the human genome sequence, searches were performed for genes encoding three categories of growth-related proteins, insulin-like growth factor-I (IGF-I) receptor, IGF-binding proteins and growth hormone receptor. The results revealed novel details about their genomic organization and new predicted transcripts. Impacts on medicine are promised in the fields of diagnostics (development of new tests), therapeutics (identification of new potential drug targets) and pharmacogenomics (streamlining of drug discovery and personalized medicine). Associated ethical, legal and social implications and controversies include genetic determinism, informed consent, privacy and confidentiality, ownership of genetic information in the biotechnology marketplace, and access to genetic healthcare.

Endocrinology↗

Transposable element (TE) display and rapid detection of TE insertion polymorphism in the Anopheles gambiae species complex.

Transposable element (TE) display was shown to be a highly specific and reproducible method of detecting the insertion sites of TEs in individuals of the African malaria mosquito, Anopheles gambiae, and its sibling species, A. arabiensis. Relatively high levels of insertion polymorphism were observed during the TE display of several families of miniature inverted-repeat TEs (MITEs) that have variable copy numbers. The genomic locations of selected insertion sites were identified by matching the sequences of their corresponding bands in a TE display gel to specific regions of the draft A. gambiae genome assembly. We discuss different scenarios in which TE display will provide powerful dominant and co-dominant genetic markers to study the behaviour of TEs in A. gambiae populations and to illustrate the complex population genetics of this intriguing disease vector. We suggest that TE display can also provide tools for a phylogenetic analysis of the A. gambiae complex.

Animals↗