PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Pharmacogenetics, pharmacogenomics and airway disease.

The availability of a draft sequence for the human genome will revolutionise research into airway disease. This review deals with two of the most important areas impinging on the treatment of patients: pharmacogenetics and pharmacogenomics. Considerable inter-individual variation exists at the DNA level in targets for medication, and variability in response to treatment may, in part, be determined by this genetic variation. Increased knowledge about the human genome might also permit the identification of novel therapeutic targets by expression profiling at the RNA (genomics) or protein (proteomics) level. This review describes recent advances in pharmacogenetics and pharmacogenomics with regard to airway disease.

Animals↗

Human resistin gene: molecular scanning and evaluation of association with insulin sensitivity and type 2 diabetes in Caucasians.

Insulin resistance is strongly associated with obesity, but even among obese subjects insulin sensitivity varies widely. Recently, a new adipocyte hormone, resistin, was identified, shown to reduce insulin-mediated glucose uptake, and shown to be increased in obese mice. We used the chromosome 19 draft sequence to determine the genomic structure of human resistin and to screen the exons, introns, and flanking sequences for variation. We screened 44 subjects with type 2 diabetes and 20 nondiabetic family members who were at the extremes of insulin sensitivity. We identified eight noncoding single nucleotide polymorphisms (SNPs) and one GAT microsatellite repeat. Three SNPs, which were in incomplete linkage disequilibrium with each other and had allelic frequencies exceeding 5%, were selected for further study. No SNP was associated with type 2 diabetes, but the SNP in the promoter region was a significant determinant of insulin sensitivity index (P = 0.04) among nondiabetic family members who had undergone iv glucose tolerance tests. The three common SNPs showed statistical significance as determinants of insulin sensitivity index (P < 0.01) in interaction with body mass index. Noncoding SNPs in the resistin gene may influence insulin sensitivity in interaction with obesity, but this finding will need to be confirmed in other populations.

Base Sequence↗

Strategies to identify disease genes.

The correlation between genes and disease began in earnest in the early 1900s with the identification of Mendelian-like inheritance of "inborn errors of metabolism." Since then, the ever-broadening field of genetics has been established as one of the most important and groundbreaking branches of science and medicine to date. With the announcement of a "working draft" sequence of the human genome in 2001, the vast array of both genomic and expressed sequence information available in the public databases alone has meant that the concept of hunting for genes is evolving. Nowadays, researchers can substitute many labor-intensive hours in the lab for less time searching on the World Wide Web. Specialization within genetics has been continuously providing subsets of the genre such as genomics, pharmacogenetics, chemogenomics, gene therapy, proteomics and functional genomics, all of which are based on the fundamental starting block, the gene. This review aims to summarize both traditional and current strategies for identifying susceptibility and monogenetic disease genes and describes how these strategies have evolved in tune with the ever-expanding wealth of information now available at our fingertips.

Animals↗

Genomics, complexity and drug discovery: insights from Boolean network models of cellular regulation.

The completion of the first draft of the human genome sequence has revived the old notion that there is no one-to-one mapping between genotype and phenotype. It is now becoming clear that to elucidate the fundamental principles that govern how genomic information translates into organismal complexity, we must overcome the current habit of ad hoc explanations and instead embrace novel, formal concepts that will involve computer modelling. Most modelling approaches aim at recreating a living system via computer simulation, by including as much details as possible. In contrast, the Boolean network model reviewed here represents an abstraction and a coarse-graining, such that it can serve as a simple, efficient tool for the extraction of the very basic design principles of molecular regulatory networks, without having to deal with all the biochemical details. We demonstrate here that such a discrete network model can help to examine how genome-wide molecular interactions generate the coherent, rule-like behaviour of a cell - the first level of integration in the multi-scale complexity of the living organism. Hereby the various cell fates, such as differentiation, proliferation and apoptosis, are treated as attractor states of the network. This modelling language allows us to integrate qualitative gene and protein interaction data to explain a series of hitherto non-intuitive cell behaviours. As the human genome project starts to reveal the limits of the current simplistic 'one gene - one function - one target' paradigm, the development of conceptual tools to increase our understanding of how the intricate interplay of genes gives rise to a global 'biological observable' will open a new perspective for post-genomic drug target discovery.

Animals↗

Functional genomics approaches to understanding brain disorders.

The completed draft of the human genome sequence has facilitated a revolution in neuroscience research. This sequence information and the development of new technologies used to analyze gene expression on a genomic scale provides a new and powerful means to investigate brain disorders of unknown etiology and to isolate novel drug targets for these disorders. The term functional genomics broadly describes a set of technologies and strategies directed at the problem of determining the function of genes, and understanding how the genome works together to generate whole patterns of biological function. The most powerful of these functional genomics approaches, expression profiling or DNA microarrays, can be used to analyze the expression of thousands of genes simultaneously. The results to date from the application of DNA microarray methods to postmortem diseased human brain tissue, animal models and cell culture models of brain disorders provide an exciting glimpse into the future of this field.

Animals↗

Primer on medical genomics. Part VI: Genomics and molecular genetics in clinical practice.

An important milestone in medical science is the recent completion of a "working draft" of the human genome sequence. The identification of all human genes and their regulatory regions provides the framework to expedite our understanding of the molecular basis of disease. This advance has also formed the foundation for a broad range of genomic tools that can be applied to medical science. These developments in global gene and gene product analysis as well as targeted molecular genetic testing are destined to change the practice of modern medicine. Despite these exciting advances, many practicing clinicians perceive that the role of molecular genetics, especially that of genomics, is confined primarily to the research arena with little current clinical applicability. The aim of this article is to highlight advances in DNA/RNA-based methods of susceptibility screening, disease diagnosis and prognostication, and prediction of treatment outcome in regard to both drug toxicity and response as they apply to various areas of clinical medicine.

Cardiovascular Diseases↗

The genetic epidemiology of cancer: interpreting family and twin studies and their implications for molecular genetic approaches.

The recent completion of a rough draft of the human genome sequence has ushered in a new era of molecular genetics research into the inherited basis of a number of complex diseases such as cancer. At the same time, recent twin studies have suggested a limited role of genetic susceptibility to many neoplasms. A reappraisal of family and twin studies for many cancer sites suggests the following general conclusions: (a) all cancers are familial to approximately the same degree, with only a few exceptions (both high and low); (b) early age of diagnosis is generally associated with increased familiality; (c) familiality does not decrease with decreasing prevalence of the tumor-in fact, the trend is toward increasing familiality with decreasing prevalence; (d) a multifactorial (polygenic) threshold model fits the twin data for most cancers less well than single gene or genetic heterogeneity-type models; (e) recessive inheritance is less likely generally than dominant or additive models; (f) heritability decreases for rarer tumors only in the context of the polygenic model but not in the context of single-locus or heterogeneity models; (g) although the family and twin data do not account for gene-environment interactions or confounding, they are still consistent with genes contributing high attributable risks for most cancer sites. These results support continued search for genetic and environmental factors in cancer susceptibility for all tumor types. Suggestions are given for optimal study designs depending on the underlying architecture of genetic predisposition.

Environment↗

Functional proteomics: mapping protein-protein interactions and pathways.

The function of a protein is defined by its interactions with other proteins and molecules. Mapping of protein interactions can highlight new functionalities for a known protein or can even define the function of novel proteins. With the draft sequence of the human genome now available, it is possible to perform high-throughput mapping of protein-protein interactions in humans, which is termed as functional proteomics. The developments in functional proteomics are particularly timely since pharmaceutical companies are searching for technologies that will strengthen their genomic efforts and prioritize their drug discovery pipeline. In this article we review recent developments in functional proteomics.

Drug Design↗

[Molecular targets in original drug research].

Selection of the molecular target, a macromolecule mediating the beneficial effects of the drug substance and thereby determining the mechanism of drug action, is of utmost importance in original drug research. Publication of the draft sequence of the human genome offers a plethora of new possible molecular targets to scientists. That should deeply impact on drug research of the next decades.

Genome, Human↗

Radiation hybrid mapping of cataract genes in the dog.

PURPOSE: To facilitate the molecular characterization of naturally occurring cataracts in dogs by providing the radiation hybrid location of 21 cataract-associated genes along with their closely associated polymorphic markers. These can be used for segregation testing of the candidate genes in canine cataract pedigrees. METHODS: Twenty-one genes with known mutations causing hereditary cataracts in man and/or mouse were selected and mapped to canine chromosomes using a canine:hamster radiation hybrid RH5000 panel. Each cataract gene ortholog was mapped in relation to over 3,000 markers including microsatellites, ESTs, genes, and BAC clones. The resulting independently determined RH-map locations were compared with the corresponding gene locations from the draft sequence of the canine genome. RESULTS: Twenty-one cataract orthologs were mapped to canine chromosomes. The genetic locations and nearest polymorphic markers were determined for 20 of these orthologs. In addition, the resulting cataract gene locations, as determined experimentally by this study, were compared with those determined by the canine genome project. All genes mapped within or near chromosomal locations with previously established homology to the corresponding human gene locations based on canine:human chromosomal synteny. CONCLUSIONS: The location of selected cataract gene orthologs in the dog, along with their nearest polymorphic markers, serves as a resource for association and linkage testing in canine pedigrees segregating inherited cataracts. The recent development of canine genomic resources make canine models a practical and valuable resource for the study of human hereditary cataracts. Canine models can serve as large animal models intermediate between mouse and man for both gene discovery and the development of novel cataract therapies.

Animals↗

Computational function assignment for potential drug targets: from single genes to cellular systems.

Biomedical science is currently undergoing an epoch-marking transition from its classical phase to the post-genome era. The outstanding success of world-wide genome sequencing efforts, evidenced by the recent publication of the draft of the human genome, together with the completion of several genomes of eukaryotic model organisms and the availability of microbial genome sequences, is opening up data sources of unprecedented scale for drug discovery. Furthermore, the elucidation of genome expression states through transcriptomic and proteomic techniques is playing a crucial role in the characterisation of disease at the molecular level. At the same time, our still very limited knowledge of the biological functions of genes and proteins at different levels of cellular organisation is preventing full exploitation of the available data. This review will discuss current computational techniques for function prediction based on the sequence-structure-function paradigm. Newly emerging approaches aimed at gaining an expanded understanding of function through integration of data from various sources and modelling of complex cellular systems will also be highlighted.

Amino Acid Sequence↗

The genome sequence of the rice blast fungus Magnaporthe grisea.

Magnaporthe grisea is the most destructive pathogen of rice worldwide and the principal model organism for elucidating the molecular basis of fungal disease of plants. Here, we report the draft sequence of the M. grisea genome. Analysis of the gene set provides an insight into the adaptations required by a fungus to cause disease. The genome encodes a large and diverse set of secreted proteins, including those defined by unusual carbohydrate-binding domains. This fungus also possesses an expanded family of G-protein-coupled receptors, several new virulence-associated genes and large suites of enzymes involved in secondary metabolism. Consistent with a role in fungal pathogenesis, the expression of several of these genes is upregulated during the early stages of infection-related development. The M. grisea genome has been subject to invasion and proliferation of active transposable elements, reflecting the clonal nature of this fungus imposed by widespread rice cultivation.

Fungal Proteins↗

Multiple messenger ribonucleic acid transcripts and revised gene organization of the human TSH receptor.

Northern blot analysis of human TSH receptor (hTSHR) messenger ribonucleic acid (mRNA) expression has previously demonstrated multiple species of transcripts in the thyroid gland, suggesting the presence of multiple transcription initiation sites, alternatively spliced forms or alternate polyadenylation (poly(A)) sites. The first two have already been reported elsewhere. To clarify alternate poly(A) sites in the hTSHR gene, the present study was designed to characterize three full-length hTSHR cDNAs with distinct poly(A) signals that we have previously cloned. The comparison of the nucleotide sequencing data on the 3'UTR of these three clones to the Draft Human Genome in NCBI database revealed that the 3' segment of exon 10 of hTSHR gene contains three tandem repeats of the poly(A) sites, from which are expressed three full-length TSHR mRNAs with distinct 3'UTR length. The longest one appears to be a predominant transcript. From these data, together with (i) the previously reported organization of hTSHR genome and (ii) use of the Draft Human Genome to localize the unidentified sequence in the alternatively spliced form of truncated hTSHR, we propose the complete structure of hTSHR gene. Rather than 10 exons, our analysis suggests that hTSHR gene seems to contain 13 exons and 12 introns. At least three full-length TSHR mRNAs with distinct poly(A) sites and five alternatively spliced forms of TSHR mRNAs are expressed from the single hTSHR gene.

Alternative Splicing↗

Large multiple organism gene finding by collapsed Gibbs sampling.

The Gibbs sampling method has been widely used for sequence analysis after it was successfully applied to the problem of identifying regulatory motif sequences upstream of genes. Since then, numerous variants of the original idea have emerged: however, in all cases the application has been to finding short motifs in collections of short sequences (typically less than 100 nucleotides long). In this paper, we introduce a Gibbs sampling approach for identifying genes in multiple large genomic sequences up to hundreds of kilobases long. This approach leverages the evolutionary relationships between the sequences to improve the gene predictions, without explicitly aligning the sequences. We have applied our method to the analysis of genomic sequence from 14 genomic regions, totaling roughly 1.8 Mb of sequence in each organism. We show that our approach compares favorably with existing ab initio approaches to gene finding, including pairwise comparison based gene prediction methods which make explicit use of alignments. Furthermore, excellent performance can be obtained with as little as four organisms, and the method overcomes a number of difficulties of previous comparison based gene finding approaches: it is robust with respect to genomic rearrangements, can work with draft sequence, and is fast (linear in the number and length of the sequences). It can also be seamlessly integrated with Gibbs sampling motif detection methods.

Algorithms↗

Heterochromatic sequences in a Drosophila whole-genome shotgun assembly.

BACKGROUND: Most eukaryotic genomes include a substantial repeat-rich fraction termed heterochromatin, which is concentrated in centric and telomeric regions. The repetitive nature of heterochromatic sequence makes it difficult to assemble and analyze. To better understand the heterochromatic component of the Drosophila melanogaster genome, we characterized and annotated portions of a whole-genome shotgun sequence assembly. RESULTS: WGS3, an improved whole-genome shotgun assembly, includes 20.7 Mb of draft-quality sequence not represented in the Release 3 sequence spanning the euchromatin. We annotated this sequence using the methods employed in the re-annotation of the Release 3 euchromatic sequence. This analysis predicted 297 protein-coding genes and six non-protein-coding genes, including known heterochromatic genes, and regions of similarity to known transposable elements. Bacterial artificial chromosome (BAC)-based fluorescence in situ hybridization analysis was used to correlate the genomic sequence with the cytogenetic map in order to refine the genomic definition of the centric heterochromatin; on the basis of our cytological definition, the annotated Release 3 euchromatic sequence extends into the centric heterochromatin on each chromosome arm. CONCLUSIONS: Whole-genome shotgun assembly produced a reliable draft-quality sequence of a significant part of the Drosophila heterochromatin. Annotation of this sequence defined the intron-exon structures of 30 known protein-coding genes and 267 protein-coding gene models. The cytogenetic mapping suggests that an additional 150 predicted genes are located in heterochromatin at the base of the Release 3 euchromatic sequence. Our analysis suggests strategies for improving the sequence and annotation of the heterochromatic portions of the Drosophila and other complex genomes.

Algorithms↗

Annotated genome assemblies of two temperate North American dung beetles, Canthon chalcites and Phanaeus vindex.

Dung beetles serve as cultivators of their natural habitats, improving soil health and functions in both natural and anthropogenic environments. Despite their ecological importance, whole genome sequences for Scarabaeinae are limited. Here, we present the draft annotated genome assemblies for 2 temperate species of North American dung beetles collected from eastern Tennessee: Canthon chalcites and Phanaeus vindex. Both genome assemblies were generated from PacBio long reads and have high completeness, with BUSCO scores of 98.1% and 98.6% for C. chalcites and P. vindex, respectively. For C. chalcites, the BRAKER3 pipeline predicted 12,799 genes, and the gene set was 93.7% complete. For P. vindex, the BRAKER3 predicted 12,252 genes, and the gene set was 94.9% complete. From the annotated gene sets, orthologous protein sequence analyses among C. chalcites, P. vindex, the dung beetle species Onthophagus taurus, and the more evolutionarily distant beetle Tribolium castaneum indicated that there are 260 unique protein clusters for C. chalcites and 210 unique protein clusters for P. vindex. These 2 draft genomes provide valuable data for comparative genomics, evolution, and phylogenic studies for dung beetle species.

Animals↗

A computational scan for U12-dependent introns in the human genome sequence.

U12-dependent introns are found in small numbers in most eukaryotic genomes, but their scarcity makes accurate characterisation of their properties challenging. A computational search for U12-dependent introns was performed using the draft version of the human genome sequence. Human expressed sequences confirmed 404 U12-dependent introns within the human genome, a 6-fold increase over the total number of non-redundant U12-dependent introns previously identified in all genomes. Although most of these introns had AT-AC or GT-AG terminal dinucleotides, small numbers of introns with a surprising diversity of termini were found, suggesting that many of the non-canonical introns found in the human genome may be variants of U12-dependent introns and, thus, spliced by the minor spliceosome. Comparisons with U2-dependent introns revealed that the U12-dependent intron set lacks the 'short intron' peak characteristic of U2-dependent introns. Analysis of this U12-dependent intron set confirmed reports of a biased distribution of U12-dependent introns in the genome and allowed the identification of several alternative splicing events as well as a surprising number of apparent splicing errors. This new larger reference set of U12-dependent introns will serve as a resource for future studies of both the properties and evolution of the U12 spliceosome.

Alternative Splicing↗

New apolipoprotein A-V: comparative genomics meets metabolism.

The availability of the human genome sequence and the recently completed draft sequences of two major mammalian model species, the mouse (Mus musculus) and the rat (Rattus norvegicus), allow researchers to apply novel approaches for gene identification and characterization, using methods of comparative and functional genomics. Recently, a new gene coding for apolipoprotein A-V was identified in the vicinity of APOA-I/C-III/A-IV cluster on human chromosome 11q23 by comparative sequencing method. In a relatively short time, compelling evidence accumulated for the substantial role of APOA-V in lipid metabolism. Studies in knock-out and transgenic mice revealed that its expression pattern correlates negatively with triglyceride levels. This observation was verified in human population studies in variety of ethnic and age groups. Several single nucleotide polymorphisms were described and particular SNP alleles and haplotypes in the APO A-V gene region were shown to be associated with dyslipidemia. The discovery and characterization of the APO A-V demonstrates current possibilities of the integrative approaches in biology, boosted by the available bioinformatic tools.

Amino Acid Sequence↗