PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Organization and evolution of a gene-rich region of the mouse genome: a 12.7-Mb region deleted in the Del(13)Svea36H mouse.

Del(13)Svea36H (Del36H) is a deletion of approximately 20% of mouse chromosome 13 showing conserved synteny with human chromosome 6p22.1-6p22.3/6p25. The human region is lost in some deletion syndromes and is the site of several disease loci. Heterozygous Del36H mice show numerous phenotypes and may model aspects of human genetic disease. We describe 12.7 Mb of finished, annotated sequence from Del36H. Del36H has a higher gene density than the draft mouse genome, reflecting high local densities of three gene families (vomeronasal receptors, serpins, and prolactins) which are greatly expanded relative to human. Transposable elements are concentrated near these gene families. We therefore suggest that their neighborhoods are gene factories, regions of frequent recombination in which gene duplication is more frequent. The gene families show different proportions of pseudogenes, likely reflecting different strengths of purifying selection and/or gene conversion. They are also associated with relatively low simple sequence concentrations, which vary across the region with a periodicity of approximately 5 Mb. Del36H contains numerous evolutionarily conserved regions (ECRs). Many lie in noncoding regions, are detectable in species as distant as Ciona intestinalis, and therefore are candidate regulatory sequences. This analysis will facilitate functional genomic analysis of Del36H and provides insights into mouse genome evolution.

Animals↗

[Rapid identification of human testis spermatocyte apoptosis-related gene, TSARG2, by nested PCR and draft human genome searching].

Cloning apoptosis-related novel genes is a key to further understanding of apoptosis mechanism and the biology process of germ cells, and is of momentous significance on clarifying physiological and pathological process of spermatogenesis. To rapidly attain human novel gene full-length cDNA sequence, the gene-specific primers and the vector-specific primers were designed for nested PCR, and draft human genome searching was performed to rapidly identify the TSARG2 (GenBank accession number AY040204) 5' end from a human testis cDNA library, by using a cDNA fragment (GenBank accession number BE644542) as an electronic probe, which was significantly changed in cryptorchidism and represented a novel gene. Furthermore, a mouse homologue of this gene was identified (GenBank accession number AF395083) by lab on-line. TSARG2 with a 1 233 bp length was composed of 6 exons and spanned about 115 kb of genomic DNA, The putative protein encoded by this gene was 305 amino acid with a theoretical molecular weight of 34 751 dalton and did not share significant homology with any known protein in databases. TSARG2 was expressed in many tissues and mapped to chromosome 4q33-34.1 by database analyses. Therefore, we propose that nested-PCR and draft human genome searching are rapid, sensitive, accurate and efficient method for isolating gene 5' end, even full-length gene from cDNA library.

Amino Acid Sequence↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

A second gene for peroxisomal HMG-CoA reductase? A genomic reassessment.

HMG-CoA reductase (HMGCR) catalyzes the conversion of HMG-CoA to mevalonate, the rate-limiting step of eukaryotic isoprenoid biosynthesis, and is the main target of cholesterol-lowering drugs. The classical form of the enzyme is a transmembrane-protein anchored to the endoplasmic reticulum. However, during the last years several lines of evidence pointed to the existence of a second isoform of HMGCR localized in peroxisomes, where mevalonate is converted further to farnesyl diphosphate. This finding is relevant for our understanding of the complex regulation and compartmentalization of the cholesterogenic pathway. Here we review experimental evidence suggesting that the peroxisomal activity might be due to a second HMGCR gene in mammals. We then present a comprehensive analysis of completely sequenced eukaryotic genomes, as well as the human and mouse genome drafts. Our results provide evidence for a large number of independent duplications of HMGCR in all eukaryotic kingdoms, but not for a second gene in mammals. We conclude that the peroxisomal HMGCR activity in mammals is due to alternative targeting of the ER enzyme to peroxisomes by an as yet uncharacterized mechanism.

Animals↗

The human genome project: implications for the endocrinologist.

The sequencing of the human genome is a major achievement of our time. This article reviews the process and current status of the working draft sequence, ways to predict genes and assign function, and conclusions for human biology. Gene density is uneven and related to chromosome banding patterns, and the estimate of approximately 30,000 genes is lower than expected. Genetic maps for men and women differ from each other and from the physical map. Single nucleotide polymorphisms occur at an average spacing of 1 kb. Human populations are 99.99% identical, and most sequences are shared between people from different continents. To illustrate the tools for accessing the human genome sequence, searches were performed for genes encoding three categories of growth-related proteins, insulin-like growth factor-I (IGF-I) receptor, IGF-binding proteins and growth hormone receptor. The results revealed novel details about their genomic organization and new predicted transcripts. Impacts on medicine are promised in the fields of diagnostics (development of new tests), therapeutics (identification of new potential drug targets) and pharmacogenomics (streamlining of drug discovery and personalized medicine). Associated ethical, legal and social implications and controversies include genetic determinism, informed consent, privacy and confidentiality, ownership of genetic information in the biotechnology marketplace, and access to genetic healthcare.

Endocrinology↗

Transposable element (TE) display and rapid detection of TE insertion polymorphism in the Anopheles gambiae species complex.

Transposable element (TE) display was shown to be a highly specific and reproducible method of detecting the insertion sites of TEs in individuals of the African malaria mosquito, Anopheles gambiae, and its sibling species, A. arabiensis. Relatively high levels of insertion polymorphism were observed during the TE display of several families of miniature inverted-repeat TEs (MITEs) that have variable copy numbers. The genomic locations of selected insertion sites were identified by matching the sequences of their corresponding bands in a TE display gel to specific regions of the draft A. gambiae genome assembly. We discuss different scenarios in which TE display will provide powerful dominant and co-dominant genetic markers to study the behaviour of TEs in A. gambiae populations and to illustrate the complex population genetics of this intriguing disease vector. We suggest that TE display can also provide tools for a phylogenetic analysis of the A. gambiae complex.

Animals↗

Plant Gene and Alternatively Spliced Variant Annotator. A plant genome annotation pipeline for rice gene and alternatively spliced variant identification with cross-species expressed sequence tag conservation from seven plant species.

The completion of the rice (Oryza sativa) genome draft has brought unprecedented opportunities for genomic studies of the world's most important food crop. Previous rice gene annotations have relied mainly on ab initio methods, which usually yield a high rate of false-positive predictions and give only limited information regarding alternative splicing in rice genes. Comparative approaches based on expressed sequence tags (ESTs) can compensate for the drawbacks of ab initio methods because they can simultaneously identify experimental data-supported genes and alternatively spliced transcripts. Furthermore, cross-species EST information can be used to not only offset the insufficiency of same-species ESTs but also derive evolutionary implications. In this study, we used ESTs from seven plant species, rice, wheat (Triticum aestivum), maize (Zea mays), barley (Hordeum vulgare), sorghum (Sorghum bicolor), soybean (Glycine max), and Arabidopsis (Arabidopsis thaliana), to annotate the rice genome. We developed a plant genome annotation pipeline, Plant Gene and Alternatively Spliced Variant Annotator (PGAA). Using this approach, we identified 852 genes (931 isoforms) not annotated in other widely used databases (i.e. the Institute for Genomic Research, National Center for Biotechnology Information, and Rice Annotation Project) and found 87% of them supported by both rice and nonrice EST evidence. PGAA also identified more than 44,000 alternatively spliced events, of which approximately 20% are not observed in the other three annotations. These novel annotations represent rich opportunities for rice genome research, because the functions of most of our annotated genes are currently unknown. Also, in the PGAA annotation, the isoforms with non-rice-EST-supported exons are significantly enriched in transporter activity but significantly underrepresented in transcription regulator activity. We have also identified potential lineage-specific and conserved isoforms, which are important markers in evolutionary studies. The data and the Web-based interface, RiceViewer, are available for public access at http://RiceViewer.genomics.sinica.edu.tw/.

Base Sequence↗

Molecular cloning of a putative Ciona intestinalis cionin receptor, a new member of the CCK/gastrin receptor family.

Cionin, a peptide showing similarities with cholecystokinin and gastrin has been shown to be expressed in the gut and neural ganglion of the protochordate Ciona intestinalis. The present report describes the cloning of a putative cionin receptor (CioR), a new member of the CCK/gastrin family from the gastrointestinal tract of C. intestinalis. mRNA from the stomach of C. intestinalis was isolated using a modified RNA extraction procedure and, subsequently, reverse-transcribed into single-stranded cDNA by means of rapid amplification of 5'- and 3'-cDNA ends (RACE-PCR), followed by full-length PCR amplification. The cloned full-length PCR amplicons contained a short upstream open-reading frame (uORF) coding for a putative 16 amino acid long peptide, followed by a long open reading frame encoding a 526 amino acid putative CioR protein. At the amino acid level, the putative CioR protein shared 35-40% homology with cloned mammalian, chicken, and Xenopus laevis CCK receptors. Phylogenetic analysis revealed that the chicken and X. laevis CCK receptors are orthologues of the mammalian CCK2 receptors whereas CioR protein forms a clade with vertebrate cholecystokinin receptors. Moreover, we found that the CioR cDNA and deduced amino acid sequences were found to correspond to the annotated CCK/gastrin-like receptor gene on Scaffold 117 (C. intestinalis draft genome project, Joint Genome Institute database; http://www.jgi.doe.gov).

Amino Acid Sequence↗

A Sanger/pyrosequencing hybrid approach for the generation of high-quality draft assemblies of marine microbial genomes.

Since its introduction a decade ago, whole-genome shotgun sequencing (WGS) has been the main approach for producing cost-effective and high-quality genome sequence data. Until now, the Sanger sequencing technology that has served as a platform for WGS has not been truly challenged by emerging technologies. The recent introduction of the pyrosequencing-based 454 sequencing platform (454 Life Sciences, Branford, CT) offers a very promising sequencing technology alternative for incorporation in WGS. In this study, we evaluated the utility and cost-effectiveness of a hybrid sequencing approach using 3730xl Sanger data and 454 data to generate higher-quality lower-cost assemblies of microbial genomes compared to current Sanger sequencing strategies alone.

Biotechnology↗

Long-range heterogeneity at the 3' ends of human mRNAs.

The publication of a draft of the human genome and of large collections of transcribed sequences has made it possible to study the complex relationship between the transcriptome and the genome. In the work presented here, we have focused on mapping mRNA 3' ends onto the genome by use of the raw data generated by the expressed sequence tag (EST) sequencing projects. We find that at least half of the human genes encode multiple transcripts whose polyadenylation is driven by multiple signals. The corresponding transcript 3' ends are spread over distances in the kilobase range. This finding has profound implications for our understanding of gene expression regulation and of the diversity of human transcripts, for the design of cDNA microarray probes, and for the interpretation of gene expression profiling experiments.

3' Flanking Region↗

Identification of a long-chain polyunsaturated fatty acid acyl-coenzyme A synthetase from the diatom Thalassiosira pseudonana.

The draft genome of the diatom Thalassiosira pseudonana was searched for DNA sequences showing homology with long-chain acyl-coenzyme A synthetases (LACSs), since the corresponding enzyme may play a key role in the accumulation of health-beneficial polyunsaturated fatty acids (PUFAs) in triacylglycerol. Among the candidate genes identified, an open reading frame named TplacsA was found to be full length and constitutively expressed during cell cultivation. The predicted amino acid sequence of the corresponding protein, TpLACSA, exhibited typical features of acyl-coenzyme A (acyl-CoA) synthetases involved in the activation of long-chain fatty acids. Feeding experiments carried out in yeast (Saccharomyces cerevisiae) transformed with the algal gene showed that TpLACSA was able to activate a number of PUFAs, including eicosapentaenoic acid and docosahexaenoic acid (DHA). Determination of acyl-CoA synthetase activities by direct measurement of acyl-CoAs produced in the presence of different PUFA substrates showed that TpLACSA was most active toward DHA. Heterologous expression also revealed that TplacsA transformants were able to incorporate more DHA in triacylglycerols than the control yeast.

Acyl Coenzyme A↗

The repetitive landscape of the chicken genome.

Cot-based cloning and sequencing (CBCS) is a powerful tool for isolating and characterizing the various repetitive components of any genome, combining the established principles of DNA reassociation kinetics with high-throughput sequencing. CBCS was used to generate sequence libraries representing the high, middle, and low-copy fractions of the chicken genome. Sequencing high-copy DNA of chicken to about 2.7 x coverage of its estimated sequence complexity led to the initial identification of several new repeat families, which were then used for a survey of the newly released first draft of the complete chicken genome. The analysis provided insight into the diversity and biology of known repeat structures such as CR1 and CNM, for which only limited sequence data had previously been available. Cot sequence data also resulted in the identification of four novel repeats (Birddawg, Hitchcock, Kronos, and Soprano), two new subfamilies of CR1 repeats, and many elements absent from the chicken genome assembly. Multiple autonomous elements were found for a novel Mariner-like transposon, Galluhop, in addition to nonautonomous deletion derivatives. Phylogenetic analysis of the high-copy repeats CR1, Galluhop, and Birddawg provided insight into two distinct genome dispersion strategies. This study also exemplifies the power of the CBCS method to create representative databases for the repetitive fractions of genomes for which only limited sequence data is available.

Animals↗

SMASHing regulatory sites in DNA by human-mouse sequence comparisons.

Regulatory sequence elements provide important clues to understanding and predicting gene expression. Although the binding sites for hundreds of transcription factors are known, there has been no systematic attempt to incorporate this information in the annotation of the human genome. Cross species sequence comparisons are critical to a meaningful annotation of regulatory elements since they generally reside in conserved non-coding regions. To take advantage of the recently completed drafts of the mouse and human genomes for annotating transcription factor binding sites, we developed SMASH, a computational pipeline that identifies thousands of orthologous human/ mouse proteins, maps them to genomic sequences, extracts and compares upstream regions and annotates putative regulatory elements in conserved, non-coding, upstream regions. Our current dataset consists of approximately 2,500 human/mouse gene pairs. Transcription start sites were estimated by mapping quasi-full length cDNA sequences. SMASH uses a novel probabilistic method to identify putative conserved binding sites that takes into account the competition between transcription factors for binding DNA. SMASH presents the results via a genome browser web interface which displays the predicted regulatory information together with the current annotations for the human genome. Our results are validated by comparison to previously published experimental data. SMASH results compare favorably to other existing computational approaches.

Algorithms↗

A gene expression map for the euchromatic genome of Drosophila melanogaster.

We used a maskless photolithography method to produce DNA oligonucleotide microarrays with unique probe sequences tiled throughout the genome of Drosophila melanogaster and across predicted splice junctions. RNA expression of protein coding and nonprotein coding sequences was determined for each major stage of the life cycle, including adult males and females. We detected transcriptional activity for 93% of annotated genes and RNA expression for 41% of the probes in intronic and intergenic sequences. Comparison to genome-wide RNA interference data and to gene annotations revealed distinguishable levels of expression for different classes of genes and higher levels of expression for genes with essential cellular functions. Differential splicing was observed in about 40% of predicted genes, and 5440 previously unknown splice forms were detected. Genes within conserved regions of synteny with D. pseudoobscura had highly correlated expression; these regions ranged in length from 10 to 900 kilobase pairs. The expressed intergenic and intronic sequences are more likely to be evolutionarily conserved than nonexpressed ones, and about 15% of them appear to be developmentally regulated. Our results provide a draft expression map for the entire nonrepetitive genome, which reveals a much more extensive and diverse set of expressed sequences than was previously predicted.

Algorithms↗

Multiplexed discovery of sequence polymorphisms using base-specific cleavage and MALDI-TOF MS.

The completion of the Human Genome Project provides researchers with a reference sequence that covers about 99% of the gene-containing regions and is more than 99.9% accurate. Sequence drafts and completed sequences for several other species are also available to researchers worldwide. The ongoing effort to provide more and more genomic reference information now enables the detection of deviations from this 'genetic blueprint'. Comparative sequencing projects will play a major role in elucidating the meaning of the genetic code and in establishing a correlation between genotype and phenotype. As part of this effort, a number of projects will focus on distinct functional aspects, like resequencing of exons or HLA determining regions. Typically these target regions are short in length and their analysis does not require long read length. To find an efficient solution for these applications, we developed a novel method that allows simultaneous analysis of multiple independent target regions (Multiplexed Comparative Sequence Analysis) by employing base-specific cleavage biochemistry and MALDI TOF-MS analysis.

Humans↗

High throughput genotyping technologies.

A comprehensive genetic map containing several hundred microsatellite markers resulted from a large microsatellite mapping project. This was the first real study that introduced high throughput methods to the genetic community. This map and the concurrent technological advances, which will briefly be reviewed, led to further numerous mapping investigations of simple and complex diseases. The annotated draft sequence of approximately three billion base pairs (bp) of the human genome has been completed much sooner than many imagined, due to considerable technological advancements and the international enterprise that resulted. This was a major development for the genetics community, but is only the precursor to the next phase of studying and understanding the variation within the human genome. The awareness of the differences may help us understand the effects on the genetics of the variation between individuals and disease. It is these variations at the nucleotide level that determine the physiological differences, or phenotypes of each individual, including all biological functions at the cellular and body level. Single nucleotide polymorphisms (SNPs) will provide the next high density map, and be the genetic tool to study these genetic variations. There are many sources of SNPs and exhaustive numbers of methods of SNP detection to be considered. The focus in this paper will be on the merits of selected, varied SNP typing methodologies that are emerging to genotype many individuals with the required huge number of SNPs to make the study of complex diseases and pharmacogenomics a practical and economically viable option.

Genotype↗

The genome of black cottonwood, Populus trichocarpa (Torr. & Gray).

We report the draft genome of the black cottonwood tree, Populus trichocarpa. Integration of shotgun sequence assembly with genetic mapping enabled chromosome-scale reconstruction of the genome. More than 45,000 putative protein-coding genes were identified. Analysis of the assembled genome revealed a whole-genome duplication event; about 8000 pairs of duplicated genes from that event survived in the Populus genome. A second, older duplication event is indistinguishably coincident with the divergence of the Populus and Arabidopsis lineages. Nucleotide substitution, tandem gene duplication, and gross chromosomal rearrangement appear to proceed substantially more slowly in Populus than in Arabidopsis. Populus has more protein-coding genes than Arabidopsis, ranging on average from 1.4 to 1.6 putative Populus homologs for each Arabidopsis gene. However, the relative frequency of protein domains in the two genomes is similar. Overrepresented exceptions in Populus include genes associated with lignocellulosic wall biosynthesis, meristem development, disease resistance, and metabolite transport.

Arabidopsis↗