PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

In silico discovery of gene-coding variants in murine quantitative trait loci using strain-specific genome sequence databases.

BACKGROUND: The identification of genes underlying complex traits has been aided by quantitative trait locus (QTL) mapping approaches, which in turn have benefited from advances in mammalian genome research. Most recently, whole-genome draft sequences and assemblies have been generated for mouse strains that have been used for a large fraction of QTL mapping studies. Here we show how such strain-specific mouse genome sequence databases can be used as part of a high-throughput pipeline for the in silico discovery of gene-coding variations within murine QTLs. As a test of this approach we focused on two QTLs on mouse chromosomes 1 and 13 that are involved in physical dependence on alcohol. RESULTS: Interstrain alignment of sequences derived from the relevant mouse strain genome sequence databases for 199 QTL-localized genes spanning 210,020 base-pairs of coding sequence identified 21 genes with different coding sequences for the progenitor strains. Several of these genes, including four that exhibit strong phenotypic links to chronic alcohol withdrawal, are promising candidates to underlie these QTLs. CONCLUSIONS: This approach has wide general utility, and should be applicable to any of the several hundred mouse QTLs, encompassing over 60 different complex traits, that have been identified using strains for which relatively complete genome sequences are available.

Alcohol Withdrawal Seizures↗

Lessons from the human genome: transitions between euchromatin and heterochromatin.

The publication of the human genome draft sequence provides, for the first time, a global view of the structural properties of the human genome. Initial sequence analysis, in combination with previous published reports, reveals that more than half of the transition regions between euchromatin and centromeric heterochromatin contain duplicated segments. The individual duplications originate from diverse euchromatic regions of the human genome, often containing intron-exon structure of known genes. Multiple duplicons are concatenated together to form larger blocks of wall-to-wall duplications. For a single chromosome, these paralogous segments can span >1 Mb of sequence and define a buffer zone between unique sequence and tandemly repeated satellite sequences. Unusual pericentromeric interspersed repeat elements have been identified at the junctions of many of these duplications. Phylogenetic and comparative studies of pericentromeric sequences suggest that this peculiar genome organization has emerged within the last 30 million years of human evolution and is a source of considerable genomic variation between closely related primate species. Interestingly, not all human pericentromeric regions show this proclivity to duplicate and transpose genomic sequence, suggesting at least two different models for the organization of these regions.

Euchromatin↗

Genomic inventory and expression of Sox and Fox genes in the cnidarian Nematostella vectensis.

The Sox and Forkhead (Fox) gene families are comprised of transcription factors that play important roles in a variety of developmental processes, including germ layer specification, gastrulation, cell fate determination, and morphogenesis. Both the Sox and Fox gene families are divided into subgroups based on the amino acid sequence of their respective DNA-binding domains, the high-mobility group (HMG) box (Sox genes) or Forkhead domain (Fox genes). Utilizing the draft genome sequence of the cnidarian Nematostella vectensis, we examined the genomic complement of Sox and Fox genes in this organism to gain insight into the nature of these gene families in a basal metazoan. We identified 14 Sox genes and 15 Fox genes in Nematostella and conducted a Bayesian phylogenetic analysis comparing HMG box and Forkhead domain sequences from Nematostella with diverse taxa. We found that the majority of bilaterian Sox groups have clear Nematostella orthologs, while only a minority of Fox groups are represented, suggesting that the evolutionary pressures driving the diversification of these gene families may be distinct from one another. In addition, we examined the expression of a subset of these genes during development in Nematostella and found that some of these genes are expressed in patterns consistent with roles in germ layer specification and the regulation of cellular behaviors important for gastrulation. The diversity of expression patterns among members of these gene families in Nematostella reinforces the notion that despite their relatively simple morphology, cnidarians possess much of the molecular complexity observed in bilaterian taxa.

Amino Acid Sequence↗

Genomic and proteomic databases: large-scale analysis and integration of data.

With the completion of the human genome draft sequencing and assembly, an upcoming targeted goal will be identifying the proteins and changes in protein expression, which are derived from the genome template. There is important information already available on many proteins that can be a valuable resource to scientists in cardiovascular medicine as well as other fields. This article will summarize many of the different databases, their content and growth, and the differences between growth of nucleic acid and protein databases. Linkage between and integration of different protein and nuclei acid databases, along with related annotation, will greatly improve the information content and knowledge base with regards to protein data.

Databases, Factual↗

GeneSeqer@PlantGDB: Gene structure prediction in plant genomes.

The GeneSeqer@PlantGDB Web server (http://www.plantgdb.org/cgi-bin/GeneSeqer.cgi) provides a gene structure prediction tool tailored for applications to plant genomic sequences. Predictions are based on spliced alignment with source-native ESTs and full-length cDNAs or non-native probes derived from putative homologous genes. The tool is illustrated with applications to refinement of current gene structure annotation and de novo annotation of draft genomic sequences. The service should facilitate expert annotation as a community effort by providing convenient access to all public plant sequences via the PlantGDB database, a simple four-step protocol for spliced alignment and visually appealing displays of the predicted gene structures in addition to detailed sequence alignments.

Arabidopsis↗

Robust method for proteome analysis by MS/MS using an entire translated genome: demonstration on the ciliome of Tetrahymena thermophila.

To improve the utility of increasingly large numbers of available unannotated and initially poorly annotated genomic sequences for proteome analysis, we demonstrate that effective protein identification can be made on a large and unannotated genome. The strategy developed is to translate the unannotated genome sequence into amino acid sequence encoding putative proteins in all six reading frames, to identify peptides by tandem mass spectrometry (MS/MS), to localize them on the genome sequence, and to preliminarily annotate the protein via a similarity search by BLAST. These tasks have been optimized and automated. Optimization to obtain multiple peptide matches in effect extends the searchable region and results in more robust protein identification. The viability of this strategy is demonstrated with the identification of 223 cilia proteins in the unicellular eukaryotic model organism Tetrahymena thermophila, whose initial genomic sequence draft was released in November 2003. To the best of our knowledge, this is the first demonstration of large-scale protein identification based on such a large, unannotated genome. Of the 223 cilia proteins, 84 have no similarity to proteins in NCBI's nonredundant (nr) database. This methodology allows identifying the locations of the genes encoding these novel proteins, which is a necessary first step to downstream functional genomic experimentation.

Amino Acid Sequence↗

An enigmatic fourth runt domain gene in the fugu genome: ancestral gene loss versus accelerated evolution.

BACKGROUND: The runt domain transcription factors are key regulators of developmental processes in bilaterians, involved both in cell proliferation and differentiation, and their disruption usually leads to disease. Three runt domain genes have been described in each vertebrate genome (the RUNX gene family), but only one in other chordates. Therefore, the common ancestor of vertebrates has been thought to have had a single runt domain gene. RESULTS: Analysis of the genome draft of the fugu pufferfish (Takifugu rubripes) reveals the existence of a fourth runt domain gene, FrRUNT, in addition to the orthologs of human RUNX1, RUNX2 and RUNX3. The tiny FrRUNT packs six exons and two putative promoters in just 3 kb of genomic sequence. The first exon is located within an intron of FrSUPT3H, the ortholog of human SUPT3H, and the first exon of FrSUPT3H resides within the first intron of FrRUNT. The two gene structures are therefore "interlocked". In the human genome, SUPT3H is instead interlocked with RUNX2. FrRUNT has no detectable ortholog in the genomes of mammals, birds or amphibians. We consider alternative explanations for an apparent contradiction between the phylogenetic data and the comparison of the genomic neighborhoods of human and fugu runt domain genes. We hypothesize that an ancient RUNT locus was lost in the tetrapod lineage, together with FrFSTL6, a member of a novel family of follistatin-like genes. CONCLUSIONS: Our results suggest that the runt domain family may have started expanding in chordates much earlier than previously thought, and exemplify the importance of detailed analysis of whole-genome draft sequence to provide new insights into gene evolution.

Amino Acid Sequence↗

Computational comparison of two draft sequences of the human genome.

We are in the enviable position of having two distinct drafts of the human genome sequence. Although gaps, errors, redundancy and incomplete annotation mean that individually each falls short of the ideal, many of these problems can be assessed by comparison. Here we present some comparative analyses of these drafts. We look at a number of features of the sequences, including sequence gaps, continuity, consistency between the two sequences and patterns of DNA-binding protein motifs.

Algorithms↗

Human retroelements may introduce intragenic polyadenylation signals.

In the human genome, the insertion of LINE-1 and Alu elements can affect genes by sequence disruption, and by the introduction of elements that modulate the gene's expression. One of the modulating sequences retroelements may contribute is the canonical polyadenylation signal (pA), AATAAA. L1 elements include these within their own sequence and AATAAA sequences are commonly created in the A-rich tails of both SINEs and LINEs. Computational analysis of 34 genes randomly retrieved from the human genome draft sequence reveals an orientation bias, reflected as a lower number of L1s and Alus containing the pA in the same orientation as the gene. Experimental studies of Alu-based pA sequences when placed in pol II or pol III transcripts suggest that the signal is very weak, or often not used at all. Because the pA signal is highly affected by the surrounding sequence, it is likely that the Alu constructs evaluated did not provide the required recognition signals to the polyadenylation machinery. Although the effect of pA signals contributed by Alus is individually weak, the observed reduction of "sense" oriented pA-containing L1 and Alu elements within genes reflects that even a modest influence causes a change in evolutionary pressure, sufficient to create the biased distribution.

Base Sequence↗

Ureases of extreme halophiles of the genus Haloarcula with a unique structure of gene cluster.

We searched for urease activities in 71 strains of extreme halophiles by a urea-phenol red-agar plate method. Positive strains were further investigated by measuring the ammonia released from urea in cell-free extracts. Only 4 strains of the genus Haloarcula, Har. aidinensis, Har. hispanica, Har. japonica, and Har. marismortui were finally shown as the urease producers. A partially purified urease from Har. hispanica was a typical halophilic enzyme in that it showed maximum activity at 18-23% NaCl and lost the activity irreversibly in the absence of NaCl. Partial genes (1596 bp) of the urease encoding from upstream of the beta subunit down to the N-terminal 139 amino acids of the alpha subunit, were PCR amplified from the four strains, as well as from five urease-negative Haloarcula strains. Strains of other genera, which were urease-negative, did not yield PCR products. The deduced amino acid sequences of the beta subunit and partial alpha subunit were similar to each other (92-100% similarities) and to those from other organisms. Analysis of the draft genome sequence of Har. marismortui, however, suggested that the order of the genes encoding the three subunits (with the total number of amino acids of 834) and four accessory proteins was beta-alpha-gamma-UreG-UreD-UreE-UreF. This order is quite unique, since in other microorganisms the order is gamma-beta-alpha-UreE-UreF-UreG-UreD in most cases. No open reading frames were detected in the PCR-amplified upstream of the beta subunit, suggesting that all Haloarcula species have the same unique structure of the urease gene cluster.

Amino Acid Sequence↗

Decoding the rice genome.

Rice cultivation is one of the most important agricultural activities on earth, with nearly 90% of it being produced in Asia. It belongs to the family of crops that includes wheat, maize and barley, and it supplies more than 50% of calories consumed by the world population. Its immense economic value and a relatively small genome size makes it a focal point for scientific investigations, so much so that four whole genome sequence drafts with varying qualities have been generated by both public and privately funded ventures. The availability of a complete and high-quality map-based sequence has provided the opportunity to study genome organization and evolution. Most importantly, the order and identity of 37,544 genes of rice have been unraveled. The sequence provides the required ingredients for functional genomics and molecular breeding programs aimed at unraveling intricate cellular processes and improving rice productivity.

Animals↗

Identification and characterization of human TIPARP gene within the CCNL amplicon at human chromosome 3q25.31.

Array CGH combined with mRNA microarray analyses was successfully applied for genome-wide screening of proto-oncogenes as well as tumor suppressor genes in 2002. It has been reported that an uncharacterized gene, corresponding to a 5'-truncated partial cDNA DKFZp434J214, is amplified and up-regulated together with CCNL in head-and-neck squamous cell carcinoma (HNSCC). Here, we identified that the novel gene, corresponding to the 5'-truncated DKFZp434J214 cDNA, is the human ortholog of mouse Tiparp gene, by using bioinformatics. Complete coding sequence of human TIPARP mRNA was determined in silico by assembling nucleotide sequences of BC034397 and DKFZp434J214 cDNAs. Human TIPARP gene, consisting of six exons, was located between SSR3 and CCNL genes at human chromosome 3q25.31. Fugu tiparp gene was located around nucleotide position 14280-23355 of fugu genome draft sequence CAAB01003597.1. Human TIPARP showed 91.8% and 52.7% total amino-acid identity with mouse Tiparp and fugu tiparp, respectively. TPH domain (codon 242-296 of human TIPARP), WWE domain (codon 329-378) and PARP-like domain (codon 465-648) were evolutionarily conserved among TIPARP proteins. CCCH-type zinc finger was located within the N-terminal part of the TPH domain. Human TIPARP showed 27.5% and 26.0% total-amino-acid identity with human FLJ22693 and ZAP, respectively. TIPARP, FLJ22693 and ZAP, sharing the common structure with the TPH, WWE and PARP-like domains, were found to constitute the TIPARP family. This is the first report on human TIPARP and fugu tiparp genes as well as on the TIPARP family.

Amino Acid Sequence↗

Co-duplication of olfactory receptor and MHC class I genes in the mouse major histocompatibility complex.

We report the 897 kb sequence of a cluster of olfactory receptor (OR) genes located at the distal end of the major histocompatibility complex (MHC) class I region on mouse chromosome 17 of strain 129/SvJ (H2bc). With additional information from the mouse genome draft sequence, we identified 59 OR loci (approximately 20% pseudogenes) in contrast to only 25 OR loci (approximately 50% pseudogenes) in the corresponding centromeric OR cluster that is part of the 'extended MHC class I region' on human chromosome 6. Comparative analysis leads to three major observations: (i) most of the OR subfamilies have evolved independently in the two species, expanding more in the mouse, and resulting in co-orthologs--subfamilies of highly similar paralogs that keep orthologous relationships with their human counterparts; (ii) three of the mouse OR subfamilies have no orthologs in humans; and (iii) MHC class I loci are interspersed in the OR cluster in mouse but not in human, and were subjected to co-duplication with OR genes. Screening of our sequence against the available sequences of other strains/haplotypes revealed that most of the OR loci are polymorphic and that the number of OR loci may vary among strains/haplotypes. Our findings that MHC-linked OR loci share duplication with MHC class I loci, have duplicated extensively and are polymorphic revives questions about potential reciprocal influences acting on the dynamics and evolution of the H2 region and the H2-linked OR loci.

Alleles↗

Segmental polymorphisms in the proterminal regions of a subset of human chromosomes.

The subtelomeric domains of chromosomes are probably the most rapidly evolving structures of the human genome. The highly variable distribution of large duplicated subtelomeric segments has indicated that frequent exchanges between nonhomologous chromosomes may have been taking place during recent genome evolution. We have studied the extent and variability of such duplications using in situ hybridization techniques and a set of well-defined subtelomeric cosmid probes that identify discrete regions within the subtelomeric domain. In addition to reciprocal translocation and illegitimate recombination events that could explain the observed mosaic pattern of subtelomeric regions, it is likely that homology-based recombination mechanisms have also contributed to the spread of distal subtelomeric sequences among particular groups of nonhomologous chromosome arms. The frequency and distribution of large-scale subtelomeric polymorphisms may have direct implications for the design of chromosome-specific probes that are aimed at the identification of cryptic subtelomeric deletions. Furthermore, our results indicate that the relevance of some of the telomere closures proposed within the present Human Genome Sequence draft are restricted to specific allelic variants of unknown frequencies.

Black People↗

HGVbase: a human sequence variation database emphasizing data quality and a broad spectrum of data sources.

HGVbase (Human Genome Variation database; http://hgvbase.cgb.ki.se, formerly known as HGBASE) is an academic effort to provide a high quality and non-redundant database of available genomic variation data of all types, mostly comprising single nucleotide polymorphisms (SNPs). Records include neutral polymorphisms as well as disease-related mutations. Online search tools facilitate data interrogation by sequence similarity and keyword queries, and searching by genome coordinates is now being implemented. Downloads are freely available in XML, Fasta, SRS, SQL and tagged-text file formats. Each entry is presented in the context of its surrounding sequence and many records are related to neighboring human genes and affected features therein. Population allele frequencies are included wherever available. Thorough semi-automated data checking ensures internal consistency and addresses common errors in the source information. To keep pace with recent growth in the field, we have developed tools for fully automated annotation. All variants have been uniquely mapped to the draft genome sequence and are referenced to positions in EMBL/GenBank files. Data utility is enhanced by provision of genotyping assays and functional predictions. Recent data structure extensions allow the capture of haplotype and genotype information, and a new initiative (along with BiSC and HUGO-MDI) aims to create a central repository for the broad collection of clinical mutations and associated disease phenotypes of interest.

Base Sequence↗

A physical map of the chicken genome.

Strategies for assembling large, complex genomes have evolved to include a combination of whole-genome shotgun sequencing and hierarchal map-assisted sequencing. Whole-genome maps of all types can aid genome assemblies, generally starting with low-resolution cytogenetic maps and ending with the highest resolution of sequence. Fingerprint clone maps are based upon complete restriction enzyme digests of clones representative of the target genome, and ultimately comprise a near-contiguous path of clones across the genome. Such clone-based maps are used to validate sequence assembly order, supply long-range linking information for assembled sequences, anchor sequences to the genetic map and provide templates for closing gaps. Fingerprint maps are also a critical resource for subsequent functional genomic studies, because they provide a redundant and ordered sampling of the genome with clones. In an accompanying paper we describe the draft genome sequence of the chicken, Gallus gallus, the first species sequenced that is both a model organism and a global food source. Here we present a clone-based physical map of the chicken genome at 20-fold coverage, containing 260 contigs of overlapping clones. This map represents approximately 91% of the chicken genome and enables identification of chicken clones aligned to positions in other sequenced genomes.

Animals↗

Keratin K6irs is specific to the inner root sheath of hair follicles in mice and humans.

BACKGROUND: Keratins are a multigene family of intermediate filament proteins that are differentially expressed in specific epithelial tissues. To date, no type II keratins specific for the inner root sheath of the human hair follicle have been identified. OBJECTIVES: To characterize a novel type II keratin in mice and humans. METHODS: Gene sequences were aligned and compared by BLAST analysis. Genomic DNA and mRNA sequences were amplified by polymerase chain reaction (PCR) and confirmed by direct sequencing. Gene expression was analysed by reverse transcription (RT)-PCR in mouse and human tissues. A rabbit polyclonal antiserum was raised against a C-terminal peptide derived from the mouse K6irs protein. Protein expression in murine tissues was examined by immunoblotting and immunofluorescence. RESULTS: Analysis of human expressed sequence tag (EST) data generated by the Human Genome Project revealed a fragment of a novel cytokeratin mRNA with characteristic amino acid substitutions in the 2B domain. No further human ESTs were found in the database; however, the complete human gene was identified in the draft genome sequence and several mouse ESTs were identified, allowing assembly of the murine mRNA. Both species' mRNA sequences and the human gene were confirmed experimentally by PCR and direct sequencing. The human gene spans more than 16 kb of genomic DNA and is located in the type II keratin cluster on chromosome 12q. A comprehensive immunohistochemical survey of expression in the adult mouse by immunofluorescence revealed that this novel keratin is expressed only in the inner root sheath of the hair follicle. Immunoblotting of murine epidermal keratin extracts revealed that this protein is specific to the anagen phase of the hair cycle, as one would expect of an inner root sheath marker. In humans, expression of this keratin was confirmed by RT-PCR using mRNA derived from plucked anagen hairs and epidermal biopsy material. By this means, strong expression was detected in human hair follicles from scalp and eyebrow. Expression was also readily detected in human palmoplantar epidermis; however, no expression was detected in face skin despite the presence of fine hairs histologically. CONCLUSIONS: This new keratin, designated K6irs, is a valuable histological marker for the inner root sheath of hair follicles in mice and humans. In addition, this keratin represents a new candidate gene for inherited structural hair defects such as loose anagen syndrome.

Amino Acid Sequence↗

Prediction of the coding sequences of mouse homologues of KIAA gene: II. The complete nucleotide sequences of 400 mouse KIAA-homologous cDNAs identified by screening of terminal sequences of cDNA clones randomly sampled from size-fractionated libraries.

We have accumulated information of the coding sequences of uncharacterized human genes, which are known as KIAA genes, and the number of these genes exceeds 2000 at present. As an extension of this sequencing project, we recently have begun to accumulate mouse KIAA-homologous cDNAs, because it would be useful to prepare a set of human and mouse homologous cDNA pairs for further functional analysis of the KIAA genes. We herein present the entire sequences of 400 mouse KIAA cDNA clones and 4 novel cDNA clones which were incidentally identified during this project. Most of clones entirely sequenced in this study were selected by computer-assisted analysis of terminal sequences of the cDNAs. The average size of the 404 cDNA sequences reached 5.3 kb and that of the deduced amino acid sequences from these cDNAs was 868 amino acid residues. The results of sequence analyses of these clones showed that single mouse KIAA cDNAs bridged two different human KIAA cDNAs in some cases, which indicated that these two human KIAA cDNAs were derived from single genes although they had been supposed to originate from different genes. Furthermore, we successfully mapped all the mouse KIAA cDNAs along the genome using a recently published mouse genome draft sequence.

Animals↗