PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

The human transcriptome map reveals extremes in gene density, intron length, GC content, and repeat pattern for domains of highly and weakly expressed genes.

The chromosomal gene expression profiles established by the Human Transcriptome Map (HTM) revealed a clustering of highly expressed genes in about 30 domains, called ridges. To physically characterize ridges, we constructed a new HTM based on the draft human genome sequence (HTMseq). Expression of 25,003 genes can be analyzed online in a multitude of tissues (http://bioinfo.amc.uva.nl/HTMseq). Ridges are found to be very gene-dense domains with a high GC content, a high SINE repeat density, and a low LINE repeat density. Genes in ridges have significantly shorter introns than genes outside of ridges. The HTMseq also identifies a significant clustering of weakly expressed genes in domains with fully opposite characteristics (antiridges). Both types of domains are open to tissue-specific expression regulation, but the maximal expression levels in ridges are considerably higher than in antiridges. Ridges are therefore an integral part of a higher order structure in the genome related to transcriptional regulation.

Base Composition↗

Retroelement distributions in the human genome: variations associated with age and proximity to genes.

Remnants of more than 3 million transposable elements, primarily retroelements, comprise nearly half of the human genome and have generated much speculation concerning their evolutionary significance. We have exploited the draft human genome sequence to examine the distributions of retroelements on a genome-wide scale. Here we show that genomic densities of 10 major classes of human retroelements are distributed differently with respect to surrounding GC content and also show that the oldest elements are preferentially found in regions of lower GC compared with their younger relatives. In addition, we determined whether retroelement densities with respect to genes could be accurately predicted based on surrounding GC content or if genes exert independent effects on the density distributions. This analysis revealed that all classes of long terminal repeat (LTR) retroelements and L1 elements, particularly those in the same orientation as the nearest gene, are significantly underrepresented within genes and older LTR elements are also underrepresented in regions within 5 kb of genes. Thus, LTR elements have been excluded from gene regions, likely because of their potential to affect gene transcription. In contrast, the density of Alu sequences in the proximity of genes is significantly greater than that predicted based on the surrounding GC content. Furthermore, we show that the previously described density shift of Alu repeats with age to domains of higher GC was markedly delayed on the Y chromosome, suggesting that recombination between chromosome pairs greatly facilitates genomic redistributions of retroelements. These findings suggest that retroelements can be removed from the genome, possibly through recombination resulting in re-creation of insert-free alleles. Such a process may provide an explanation for the shifting distributions of retroelements with time.

Alu Elements↗

Beyond complementation. Map-based cloning in Chlamydomonas reinhardtii.

Chlamydomonas reinhardtii is an excellent model system for plant biologists because of its ease of manipulation, facile genetics, and the ability to transform the nuclear, chloroplast, and mitochondrial genomes. Numerous forward genetics studies have been performed in Chlamydomonas, in many cases to elucidate the regulation of photosynthesis. One of the resultant challenges is moving from mutant phenotype to the gene mutation causing that phenotype. To date, complementation has been the primary method for gene cloning, but this is impractical in several situations, for example, when the complemented strain cannot be readily selected or in the case of recessive suppressors that restore photosynthesis. New tools, including a molecular map consisting of 506 markers and an 8X-draft nuclear genome sequence, are now available, making map-based cloning increasingly feasible. Here we discuss advances in map-based cloning developed using the strains mcd4 and mcd5, which carry recessive nuclear suppressors restoring photosynthesis to chloroplast mutants. Tools that have not been previously applied to Chlamydomonas, such as bulked segregant analysis and marker duplexing, are being implemented to increase the speed at which one can go from mutant phenotype to gene. In addition to assessing and applying current resources, we outline anticipated future developments in map-based cloning in the context of the newly extended Chlamydomonas genome initiative.

Animals↗

Fatty acid desaturases from the microalga Thalassiosira pseudonana.

Analysis of a draft nuclear genome sequence of the diatom Thalassiosira pseudonana revealed the presence of 11 open reading frames showing significant similarity to functionally characterized fatty acid front-end desaturases. The corresponding genes occupy discrete chromosomal locations as determined by comparison with the recently published genome sequence. Phylogenetic analysis showed that two of the T. pseudonana desaturase (Tpdes) sequences grouped with proteobacterial desaturases that lack a fused cytochrome b5 domain. Among the nine remaining gene sequences, temporal expression analysis revealed that seven were expressed in T. pseudonana cells. One of these, TpdesN, was previously characterized as encoding a Delta11-desaturase active on palmitic acid. From the six remaining putative desaturase genes, we report here that three, TpdesI, TpdesO and TpdesK, respectively encode Delta6-, Delta5- and Delta4-desaturases involved in production of the health beneficial polyunsaturated fatty acid DHA (docosahexaenoic acid). Furthermore, we show that one of the remaining genes, TpdesB, encodes a Delta8-sphingolipid desaturase with strong preference for dihydroxylated substrates.

Chromatography, Gas↗

Integration targeting by avian sarcoma-leukosis virus and human immunodeficiency virus in the chicken genome.

We have analyzed the placement of sites of integration of avian sarcoma-leukosis virus (ASLV) and human immunodeficiency virus (HIV) DNA in the draft chicken genome sequence, with the goals of assessing species-specific effects on integration and allowing comparison to the distribution of chicken endogenous retroviruses (ERVs). We infected chicken embryo fibroblasts (CEF) with ASLV or HIV and sequenced 863 junctions between host and viral DNA. The relationship with cellular gene activity was analyzed by transcriptional profiling of uninfected or ASLV-infected CEF cells. ASLV weakly favored integration in active transcription units (TUs), and HIV strongly favored active TUs, trends seen previously for integration in human cells. The ERVs, in contrast, accumulated mostly outside TUs, including ERVs related to ASLV. The minority of ERVs present within TUs were mainly in the antisense orientation; consequently, the viral splicing and polyadenylation signals would not disrupt cellular mRNA synthesis. In contrast, de novo ASLV integration sites within TUs showed no orientation bias. Comparing the distribution of de novo ASLV integration sites to ERVs indicated that purifying selection against gene disruption, and not initial integration targeting, probably determined the ERV distribution. Further analysis indicated that ERVs in humans, mice, and rats showed similar distributions, suggesting purifying selection dictated their distributions as well.

Alpharetrovirus↗

Most X;autosome translocations associated with premature ovarian failure do not interrupt X-linked genes.

Balanced translocations with breakpoints in a critical region of the X chromosome, Xq13-->q26, are associated with premature ovarian failure (POF). Translocations may cause POF either by affecting expression of specific X-linked genes essential for maintenance of normal ovarian function or by a chromosomal effect such as inhibition of meiotic pairing or altered X inactivation. We previously mapped seven Xq translocation breakpoints associated with POF to approximately 75-kb intervals. One translocation disrupted an aminopeptidase gene, XPNPEP2. We have now refined the map location of the remaining six breakpoints with respect to known genes and transcription units predicted from the draft human genome sequence. Only one of the six breakpoints disrupts a gene, DACH2, the human ortholog of a mouse gene expressed in embryonic nervous tissue, sensory organs, and limbs. DACH2 has no obvious relationship to ovarian function. The other five breakpoints fall in apparently intragenic regions. Our results are most consistent with models for POF associated with X;autosome translocations that involve generalized chromosome effects.

Amino Acid Sequence↗

Bases and spaces: resources on the web for accessing the draft human genome.

SUMMARY: Much is expected of the draft human genome sequence, and yet there is no central resource to host the plethora of sequence and mapping information available. Consequently, finding the most useful and reliable human genome data and resources currently available on the web can be challenging, but is not impossible.

Cloning, Molecular↗

Reconstruction of regulatory and metabolic pathways in metal-reducing delta-proteobacteria.

BACKGROUND: Relatively little is known about the genetic basis for the unique physiology of metal-reducing genera in the delta subgroup of the proteobacteria. The recent availability of complete finished or draft-quality genome sequences for seven representatives allowed us to investigate the genetic and regulatory factors in a number of key pathways involved in the biosynthesis of building blocks and cofactors, metal-ion homeostasis, stress response, and energy metabolism using a combination of regulatory sequence detection and analysis of genomic context. RESULTS: In the genomes of delta-proteobacteria, we identified candidate binding sites for four regulators of known specificity (BirA, CooA, HrcA, sigma-32), four types of metabolite-binding riboswitches (RFN-, THI-, B12-elements and S-box), and new binding sites for the FUR, ModE, NikR, PerR, and ZUR transcription factors, as well as for the previously uncharacterized factors HcpR and LysX. After reconstruction of the corresponding metabolic pathways and regulatory interactions, we identified possible functions for a large number of previously uncharacterized genes covering a wide range of cellular functions. CONCLUSIONS: Phylogenetically diverse delta-proteobacteria appear to have homologous regulatory components. This study for the first time demonstrates the adaptability of the comparative genomic approach to de novo reconstruction of a regulatory network in a poorly studied taxonomic group of bacteria. Recent efforts in large-scale functional genomic characterization of Desulfovibrio species will provide a unique opportunity to test and expand our predictions.

Bacterial Proteins↗

The hepatic stellate cell in the post-genomic era.

The draft human genome sequence was published on February 15, 2001, which will provide a huge amount of information on human genetics, human disease, and human cell biology. Now, medical scientists and cell biologists are turning their attention to illustrating gene expression pattern using gene microarray and to identifying the functions and the expression patterns of proteins encoded by the genes. Hepatic stellate cell is one of the sinusoid-constituent cells that play multiple roles in the liver pathophysiology. Transformation of stellate cells from the vitamin A-storing phenotype to the "myofibroblastic" one closely correlates to hepatic fibrosis during chronic liver trauma. Analyses of the molecular mechanisms of stellate cell activation have made a great progress, in particular, in the field of intracellular signal transduction of transforming growth factor-beta and platelet-derived growth factor, integrin signaling related to cell-adhesion, and cell motility-associated Rho and focal-adhesion kinase. Accumulation of the information on the stellate cell activation would shed light on the establishment of a novel therapeutic strategy against fibrosis of human liver disease.

Animals↗

Phylogeny and evolution of class-I helical cytokines.

The class-I helical cytokines constitute a large group of signalling molecules that play key roles in a plethora of physiological processes including host defence, immune regulation, somatic growth, reproduction, food intake and energy metabolism, regulation of neural growth and many more. Despite little primary amino acid sequence similarity, the view that all contemporary class-I helical cytokines have expanded from a single ancestor is widely accepted, as all class-I helical cytokines share a similar three-dimensional fold, signal via related class-I helical cytokine receptors and activate similar intracellular signalling cascades. Virtually all of our knowledge on class-I helical cytokine signalling derives from research on primate and rodent species. Information on the presence, structure and function of class-I helical cytokines in non-mammalian vertebrates and non-vertebrates is fragmentary. Consequently, our ideas about the evolution of this versatile multigene family are often based on a limited comparison of human and murine orthologs. In the last 5 years, whole genome sequencing projects have yielded draft genomes of the early vertebrates, pufferfish (Takifugu rubripes), spotted green pufferfish (Tetraodon nigroviridis) and zebrafish (Danio rerio). Fuelled by this development, fish orthologs of a number of mammalian class-I helical cytokines have recently been discovered. In this review, we have characterised the mammalian class-I helical cytokine family and compared it with the emerging class-I helical cytokine repertoire of teleost fish. This approach offers important insights into cytokine evolution as it identifies the helical cytokines shared by fish and mammals that, consequently, existed before the divergence of teleosts and tetrapods. A 'fish-mammalian' comparison will identify the class-I helical cytokines that still await discovery in fish or, alternatively, may have been evolutionarily recent additions to the mammalian cytokine repertoire.

Animals↗

The human genome and the future of medicine.

The draft human genome sequence (about 3 billion base pairs) was completed in 2001. Humans have fewer protein-coding genes than expected, and most of these are highly conserved among animals. Humans and other complex organisms produce massive amounts of non-coding RNAs, which may form another level of genetic output that controls differentiation and development. Aside from classical monogenic diseases and other differences caused by mutations and polymorphisms in protein-coding genes, much of the variation between individuals, including that which may affect our predispositions to common diseases, is probably due to differences in the non-coding regions of the genome (ie, the control architecture of the system). Within 10 years we can expect to see: increased penetration of DNA diagnostic tests to assess risk of disease, to diagnose pathogens, to determine the best treatment regimens, and for individual identification; a range of new pharmaceuticals as well as new gene and cell therapies to repair damage, to optimise health and to minimise future disease risk; and medicine become increasingly personalised, with the knowledge of individual genetic make-up and lifestyle influences.

Forecasting↗

Identification and mapping of nuclear matrix-attachment regions in a one megabase locus of human chromosome 19q13.12: long-range correlation of S/MARs and gene positions.

The first draft human genome sequence now available allowed the identification of an enormous number of gene coding areas of the genomic DNA. However, a great number of regulatory elements such as enhancers, promoters, transcription terminators, or replication origins can not be identified unequivocally by their nucleotide sequences in complex eukaryotic genomes. One important subclass of these type of sequences is scaffold/matrix attachment regions (S/MARs) that were hypothesized to anchor chromatin loops or domains to the nuclear matrix and/or chromosome scaffold. We developed an experimental selection procedure to identify S/MARs within a completely sequenced one megabase (1 Mb) long gene-rich D19S208-COX7A1 locus of human chromosome 19. A library of S/MAR elements from the locus was prepared and shown to contain -20 independent S/MARs. Sixteen of them were isolated, sequenced, and assigned to certain positions within the locus. A majority of the S/MARs identified (11 out of 16) lie in intergenic regions, suggesting their structural role, i.e., delimitation of chromatin domains. These 11 S/MARs subdivide the locus into 10 domains ranging from 6 to 272 kb with an average domain size of 88 kb. The remaining five S/MARs were found within intronic sequences of APLP1, HSPOX1, MAG, and NPHS1 genes, and can be tentatively characterized as regulatory S/MARs. The correspondence of the chromatin domains defined by the S/MARs to functional characteristics of the genes therein is discussed. The approach described can be a prototype of a similar search of long sequenced genomic stretches and/or whole chromosomes for various regulatory elements.

Chromosome Mapping↗

[Malaria kills over 1 million people every year. Genomic mapping of malaria parasite and mosquito raise hope for a vaccine as well as more effective drugs].

Every year, malaria kills between 1 and 2 million people. Another half billion get infected but survive. Most cases of malaria are found in sub-Saharan Africa. Because of drug and insecticide resistance and social and environmental changes the problems are still increasing. There is therefore a desperate need for vaccines and new drugs and insecticides. Several recently published research discoveries may help to speed up the development of new tools to fight malaria. Two years ago the draft human genome sequence was released. Now the sequencing of the genomes for the most common malaria parasite, Plasmodium falciparum, and the vector mosquito, Anopheles gambiae, have been completed. For the first time researchers have the genomic maps of all three organisms in an infectious disease available.

Animals↗

[Tandem and interspersed repeats contribute to the mosaic structure of segmented duplications in the human genome].

Intrachromosomal and interchromosomal segmental duplications account for more than 5% of the human genome. To analyze the processes resulting in the complex mosaic structure of duplicons, a draft human genome sequence was searched for duplicated segments of a genomic fragment of the pericentric region of the chromosome 21 short arm. The duplicons found consist of modules having paralogs in various genome regions. Module ends are flanked with various tandem or interspersed repeats, which are more unstable as compared with unique sequences. In most cases, the boundaries of duplicated segments exactly coincide with or are in close proximity to hot spots of various rearrangements within repeats or boundaries between repeats and unique sequences or between two different repeats. Homologous recombination between repetitive elements was assumed to be the major mechanism contributing to the mosaic structure of duplicons.

Chromosomes, Human, Pair 21↗

[Nutrition genomics].

The importance of nutrition for human health and its influence on the onset and course of many diseases are nowadays considered as proven. Only the recent development of molecular biology and biochemical methods allows the elucidation of the molecular mechanisms of diet constituent actions and their subsequent effect on homeostatic mechanisms in health and disease states. The availability of the draft human genome sequence as well as the genome sequences of model organisms, combined with the functional and integrative genomics approaches of systems biology, bring about the possibility to identify alleles and haplotypes responsible for specific reaction to the dietary challenge in susceptible individuals. Such complex interactions are studied within the newly conceived field, the nutrition genomics (nutrigenomics). Using the tools of highly parallel analyses of transcriptome, proteome and metabolome, the nutrition genomics pursues its ultimate goal, i.e. the individualized diet, respecting not only quantitative and qualitative nutritional needs and the actual health status, but also the genetic predispositions of an individual. This approach should lead to prevention of the onset of such diseases as obesity, hypertension or type 2 diabetes, or enhance the efficiency of their therapy.

Animals↗

Identification of nine human-specific frameshift mutations by comparative analysis of the human and the chimpanzee genome sequences.

MOTIVATION: The recent release of the draft sequence of the chimpanzee genome is an invaluable resource for finding genome-wide genetic differences that might explain phenotypic differences between humans and chimpanzees. AVAILABILITY: In this paper, we describe a simple procedure to identify potential human-specific frameshift mutations that occurred after the divergence of human and chimpanzee. The procedure involves collecting human coding exons bearing insertions or deletions compared with the chimpanzee genome and identification of homologs from other species, in support of the mutations being human-specific. Using this procedure, we identified nine genes, BASE, DNAJB3, FLJ33674, HEJ1, NTSR2, RPL13AP, SCGB1D4, WBSCR27 and ZCCHC13, that show human-specific alterations including truncations of the C-terminus. In some cases, the frameshift mutation results in gene inactivation or decay. In other cases, the altered protein seems to be functional. This study demonstrates that even the unfinished chimpanzee genome sequence can be useful in identifying modification of genes that are specific to the human lineage and, therefore, could potentially be relevant to the study of the acquisition of human-specific traits.

Amino Acid Sequence↗

Genomics of the human carnitine acyltransferase genes.

Five genes in the human genome are known to encode different active forms of related carnitine acyltransferases: CPT1A for liver-type carnitine palmitoyltransferase I, CPT1B for muscle-type carnitine palmitoyltransferase I, CPT2 for carnitine palmitoyltransferase II, CROT for carnitine octanoyltransferase, and CRAT for carnitine acetyltransferase. Only from two of these genes (CPT1B and CPT2) have full genomic structures been described. Data from the human genome sequencing efforts now reveal drafts of the genomic structure of CPT1A and CRAT, the latter not being known from any other mammal. Furthermore, cDNA sequences of human CROT were obtained recently, and database analysis revealed a completed bacterial artificial chromosome sequence that contains the entire CROT gene and several exons of the flanking genes P53TG and PGY3. The genomic location of CROT is at chromosome 7q21.1. There is a putative CPT1-like pseudogene in the carnitine/choline acyltransferase family at chromosome 19. Here we give a brief overview of the functional relations between the different carnitine acyltransferases and some of the common features of their genes. We will highlight the phylogenetics of the human carnitine acyltransferase genes in relation to the fungal genes YAT1 and CAT2, which encode cytosolic and mitochondrial/peroxisomal carnitine acetyltransferases, respectively.

Carnitine Acyltransferases↗