PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Identification of SmtB/ArsR cis elements and proteins in archaea using the Prokaryotic InterGenic Exploration Database (PIGED).

Microbial genome sequencing projects have revealed an apparently wide distribution of SmtB/ArsR metal-responsive transcriptional regulators among prokaryotes. Using a position-dependent weight matrix approach, prokaryotic genome sequences were screened for SmtB/ArsR DNA binding sites using data derived from intergenic sequences upstream of orthologous genes encoding these regulators. Sixty SmtB/ArsR operators linked to metal detoxification genes, including nine among various archaeal species, are predicted among 230 annotated and draft prokaryotic genome sequences. Independent multiple sequence alignments of putative operator sites and corresponding winged helix-turn-helix motifs define sequence signatures for the DNA binding activity of this SmtB/ArsR subfamily. Prediction of an archaeal SmtB/ArsR based upon these signature sequences is confirmed using purified Methanosarcina acetivorans C2A protein and electrophoretic mobility shift assays. Tools used in this study have been incorporated into a web application, the Prokaryotic InterGenic Exploration Database (PIGED; http://bioinformatics.uwp.edu/~PIGED/home.htm), facilitating comparable studies. Use of this tool and establishment of orthology based on DNA binding signatures holds promise for deciphering potential cellular roles of various archaeal winged helix-turn-helix transcriptional regulators.

Archaea↗

A general approach to single-nucleotide polymorphism discovery.

Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence as a template on which to layer often unmapped, fragmentary sequence data and to use base quality values to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.

Algorithms↗

Genomics, the cytoskeleton and motility.

The draft human genome sequence is an important step in cataloguing the molecular hardware that supports the processes of life. Here I look at what we have learned from the draft sequence about our cytoskeletal and motility systems. Most cytoskeletal and motility proteins were discovered previously by biochemical isolation, traditional cloning methods or random sequences of complementary DNAs. The ongoing challenges of assembling and annotating genes for motor proteins with long, fragmented coding sequences emphasize the importance of expert knowledge of related proteins and confirmatory evidence from cDNA sequences.

Actins↗

Gene structure for adenosine kinase in Chinese hamster and human: high-frequency mutants of CHO cells involve deletions of several introns and exons.

The structure for the adenosine kinase (AK) gene has been determined from Chinese hamster (CH) and human cells. The AK gene in CH is comprised of 11 exons ranging in length from 36 to 765 nt, with the majority <100 nt. The exact lengths of the intervening introns have not been determined, but most of them are indicated to be very large (>15 kb). A 6.6-kb fragment from human cells was also sequenced, and it contained only a single exon corresponding to exon 10 in CH. The BLAST searches of the subsequently released draft human genome sequence have revealed that the AK gene structure in human is identical to that in CH. In the human genome, the AK exons are distributed over four genomic clones totaling 752 kb, providing direct evidence that the AK gene in mammalian species is unusually large. In contrast to CH and human, the AK genes from several other eukaryotic organisms whose complete genomes are now known are quite small (between 1.2 and 2.5 kb) and either contain no introns (Saccharomyces cerevisiae and Schizosaccharomyces pombe) or various numbers of introns (Drosophila melanogaster [2], Caenorhabditis elegans [4], Arabidopsis thaliana [10]). Some of the intron-exon junctions in these species are in the same positions as in mammals. The AK gene in CH and human, as well as mouse, is linked upstream in a head-to-head fashion with the gene for the clathrin adaptor mu3 protein (or beta 3A subunit of the AP-3 protein complex), which is affected in type 2 Hermansky-Pudlak syndrome. These two genes are separated by <200 nt, and it is possible that they have a common or overlapping promoter(s). We have also determined the nature of the genetic alterations in two of the class A AK(-) mutants of CHO cells, which are obtained at a very high spontaneous frequency (10(-3)-10(-4)) in this cell line. Both mutants contained large deletions within the AK gene and greatly shortened AK transcripts. The cloning and sequencing of the transcripts from these mutants showed that the deletion in one of them led to the loss of exons 5 through 8, whereas in the other, all exons from 2 through 8 are deleted. The endpoints of these deletions lie in the large introns within the AK gene.

Adenosine Kinase↗

Transcript map and complete genomic sequence for the 310 kb region of minimal allele loss on chromosome segment 11p15.5 in non-small-cell lung cancer.

Molecular, functional, and clinical analyses strongly suggest that chromosome segment 11p15.5 contains a gene involved in lung cancer pathogenesis. The critical region of allele loss is 310 kb in size. We used our contig of P1-phage artificial chromosome (PAC) clones together with newly identified bacterial artificial chromosome (BAC) clones and the draft human genome sequence to complete a contiguous string of 380 407 bp. Three PAC clones that span the region were used to identify transcripts by exon trapping. Computational gene prediction algorithms were used to query the sequence for potential genes and exons. Screening for expression was performed with tissue-specific and cell line derived mRNA arrays. The region contains the complete SSA/Ro52 and RRM1 genes, exons 7-12 of the GOK gene, and the psirad pseudo-gene. A cluster of six nearly identical genes with an intact open reading frame (ORF) of 585 bp that share 75% identity with the HSPC182 gene was found. In addition, five putative novel genes were identified. Sequence tagged sites (STS) and polymorphic markers were used to screen 117 lung cancer cell lines for homozygous deletions and none were identified. These data provide the basis for the identification of a lung cancer suppressor gene on 11p15.5.

Base Sequence↗

A comprehensive analysis of recently integrated human Ta L1 elements.

The Ta (transcribed, subset a) subfamily of L1 LINEs (long interspersed elements) is characterized by a 3-bp ACA sequence in the 3' untranslated region and contains approximately 520 members in the human genome. Here, we have extracted 468 Ta L1Hs (L1 human specific) elements from the draft human genomic sequence and screened individual elements using polymerase-chain-reaction (PCR) assays to determine their phylogenetic origin and levels of human genomic diversity. One hundred twenty-four of the elements amenable to complete sequence analysis were full length ( approximately 6 kb) and have apparently escaped any 5' truncation. Forty-four of these full-length elements have two intact open reading frames and may be capable of retrotransposition. Sequence analysis of the Ta L1 elements showed a low level of nucleotide divergence with an estimated age of 1.99 million years, suggesting that expansion of the L1 Ta subfamily occurred after the divergence of humans and African apes. A total of 262 Ta L1 elements were screened with PCR-based assays to determine their phylogenetic origin and the level of human genomic variation associated with each element. All of the Ta L1 elements analyzed by PCR were absent from the orthologous positions in nonhuman primate genomes, except for a single element (L1HS72) that was also present in the common (Pan troglodytes) and pygmy (P. paniscus) chimpanzee genomes. Sequence analysis revealed that this single exception is the product of a gene conversion event involving an older preexisting L1 element. One hundred fifteen (45%) of the Ta L1 elements were polymorphic with respect to insertion presence or absence and will serve as identical-by-descent markers for the study of human evolution.

Animals↗

Identification of six novel genes by experimental validation of GeneMachine predicted genes.

In silico gene identification from finished and unfinished human genome sequence has become critically important in many projects seeking to gain insights into the gene content of genomic regions implicated in diseases. To establish limitations and criteria for in silico gene identification, and to identify novel genes of potential relevance to human prostate cancer and melanoma, 3 Mb of chromosome 1 sequence have been analyzed using GeneMachine. This program is a software suite comprising of sequence similarity programs and four gene identification programs. A total of 49 potential transcripts were selected and 37 of them were selected for experimental validation. We verified 16 of the predicted genes by experimental analysis. The comparison of the predicted transcripts with their cloned forms helped to refine predicted gene models as well as to identify splice variants for several of them. Although sequences matching with ten of our verified genes have been recently deposited in the GenBank, six of them remain novel. Our studies support the feasibility of identifying novel genes from regions of interest using draft human genome sequence.

Chromosomes, Human, Pair 1↗

Conservation of human alternative splice events in mouse.

Human and mouse genomes share similar long-range sequence organization, and have most of their genes being homologous. As alternative splicing is a frequent and important aspect of gene regulation, it is of interest to assess the level of conservation of alternative splicing. We examined mouse transcript data sets (EST and mRNA) for the presence of transcripts that both make spliced-alignment with the draft mouse genome sequence and demonstrate conservation of human transcript-confirmed alternative and constitutive splice junctions. This revealed 15% of alternative and 67% of constitutive splice junctions as conserved; however, these numbers are patently dependent on the extent of transcript coverage. Transcript coverage of conserved splice patterns is found to correlate well between human and mouse. A model, which extrapolates from observed levels of conservation at increasing levels of transcript support, estimates overall conservation of 61% of alternative and 74% of constitutive splice junctions, albeit with broad confidence intervals. Observed numbers of conserved alternative splicing events agreed with those expected on the basis of the model. Thus, it is apparent that many, and probably most, alternative splicing events are conserved between human and mouse. This, combined with the preservation of alternative frame stop codons in conserved frame breaking events, indicates a high level of commonality in patterns of gene expression between these two species.

Alternative Splicing↗

Two CD1 genes map to the chicken MHC, indicating that CD1 genes are ancient and likely to have been present in the primordial MHC.

CD1 molecules play an important role in the immune system, presenting lipid-containing antigens to T and NKT cells. CD1 genes have long been thought to be as ancient as MHC class I and II genes, based on various arguments, but thus far they have been described only in mammals. Here we describe two CD1 genes in chickens, demonstrating that the CD1 system was present in the last common ancestor of mammals and birds at least 300 million years ago. In phylogenetic analysis, these sequences cluster with CD1 sequences from other species but are not obviously like any particular CD1 isotype. Sequence analysis suggests that the expressed proteins bind hydrophobic molecules and are recycled through intracellular vesicles. RNA expression is strong in lymphoid tissues but weaker to undetectable in some nonlymphoid tissues. Flow cytometry confirms expression from one gene on B cells. Based on Southern blotting and cloning, only two such CD1 genes are detected, located approximately 800 nucleotides apart and in the same transcriptional orientation. The sequence of one gene is nearly identical in six chicken lines. By mapping with a backcross family, this gene could not be separated from the chicken MHC on chromosome 16. Mining the draft chicken genome sequence shows that chicken has only these two CD1 genes located approximately 50 kb from the classical class I genes. The unexpected location of these genes in the chicken MHC suggests the CD1 system was present in the primordial MHC and is thus approximately 600 million years old.

Amino Acid Sequence↗

FlyBase: genes and gene models.

FlyBase (http://flybase.org) is the primary repository of genetic and molecular data of the insect family Drosophilidae. For the most extensively studied species, Drosophila melanogaster, a wide range of data are presented in integrated formats. Data types include mutant phenotypes, molecular characterization of mutant alleles and aberrations, cytological maps, wild-type expression patterns, anatomical images, transgenic constructs and insertions, sequence-level gene models and molecular classification of gene product functions. There is a growing body of data for other Drosophila species; this is expected to increase dramatically over the next year, with the completion of draft-quality genomic sequences of an additional 11 Drosophila species.

Animals↗

Genome sequence data of the chitinase-producing bacterium Paenibacillus mucilaginosus YWY-5.1.

Paenibacillus mucilaginosus is a beneficial bacterium widely applied as a biofertilizer in agriculture. To date, genomic information on this species remains limited; however, no genome assemblies from Vietnam have been reported. This work presented the draft genome of P. mucilaginosus YWY-5.1, a promising strain with strong chitin-degrading capability and agricultural potential, isolated from Yok Don National Park, Vietnam, using Illumina technology. Results showed that the assembled genome comprised 48 contigs with 4,076,146 bp and 73.8% GC-content. Genome annotation identified 3,611 protein-coding genes, 2 rRNA genes, and 53 tRNA genes. A total of 150 carbohydrate-active enzyme-related genes were predicted from the genome; among them, seven putative chitinolytic genes were identified, including 4 genes related to family 18 chitinase, 2 genes to family 20 &#x3b2;-N-acetylglucosaminidase, and one gene to auxiliary activity family 10. In addition, at least 32 genes related to plant growth-promoting functions were identified, including those associated with indole-3-acetic acid production, phosphate and potassium solubilization, siderophore biosynthesis, iron uptake, ACC metabolism, and nitrate transport and reduction. Furthermore, genome mining identified 4 biosynthetic gene clusters probably involved in secondary metabolite production, of which 3 displayed no similarity to previously reported clusters, indicating potential for novel bioactive compounds. These genomic data improved our understanding of the biodegradation capacity and agricultural potential of P. mucilaginosus YWY-5.1 isolated from Vietnam, and provided a valuable genomic resource for future functional and biotechnological investigations toward crop production and related fields.

Chitinases↗

The 2R hypothesis and the human genome sequence.

One theory formalised in 1970 proposes that the complexity of vertebrate genomes originated by means of genome duplication at the base of the vertebrate lineage. Since then, the theory has remained both popular and controversial. Here we review the theory, and present preliminary results from our analysis of duplications in the draft human genome sequence. We find evidence for extensive duplication of parts of the genome. We also question the validity of the 'parsimony test' that has been used in other analyses.

Animals↗

Shotgun sequencing of the human transcriptome with ORF expressed sequence tags.

Theoretical considerations predict that amplification of expressed gene transcripts by reverse transcription-PCR using arbitrarily chosen primers will result in the preferential amplification of the central portion of the transcript. Systematic, high-throughput sequencing of such products would result in an expressed sequence tag (EST) database consisting of central, generally coding regions of expressed genes. Such a database would add significant value to existing public EST databases, which consist mostly of sequences derived from the extremities of cDNAs, and facilitate the construction of contigs of transcript sequences. We tested our predictions, creating a database of 10,000 sequences from human breast tumors. The data confirmed the central distribution of the sequences, the significant normalization of the sequence population, the frequent extension of contigs composed of existing human ESTs, and the identification of a series of potentially important homologues of known genes. This approach should make a significant contribution to the early identification of important human genes, the deciphering of the draft human genome sequence currently being compiled, and the shotgun sequencing of the human transcriptome.

Animals↗

LINE-1 preTa elements in the human genome.

The preTa subfamily of long interspersed elements (LINEs) is characterized by a three base-pair "ACG" sequence in the 3' untranslated region, contains approximately 400 members in the human genome, and has low level of nucleotide divergence with an estimated average age of 2.34 million years old suggesting that expansion of the L1 preTa subfamily occurred just after the divergence of humans and African apes. We have identified 362 preTa L1 elements from the draft human genomic sequence, investigated the genomic characteristics of preTa L1 insertions, and screened individual elements across diverse human populations and various non-human primate species using polymerase chain reaction (PCR) assays to determine the phylogenetic origin and levels of human genomic diversity associated with the L1 elements. All of the preTa L1 elements analyzed by PCR were absent from the orthologous positions in non-human primate genomes with 33 (14%) of the L1 elements being polymorphic with respect to insertion presence or absence in the human genome. The newly identified L1 insertion polymorphisms will prove useful as identical by descent genetic markers for the study of human population genetics. We provide evidence that preTa L1 elements show an integration site preference for genomic regions with low GC content. Computational analysis of the preTa L1 elements revealed that 29% of the elements amenable to complete sequence analysis have apparently escaped 5' truncation and are essentially full-length (approximately 6kb). In all, 29 have two intact open reading frames and may be capable of retrotransposition.

Animals↗

Whole-Genome Sequence Dataset of Rhodococcus qingshengii IEGM 267-Terpenoid Biotransformer Toward Genetic Functional Annotation.

Background/Objectives: Microbial biotransformation of monoterpenoids is a promising approach for obtaining bioactive compounds. Rhodococcus species are attractive biocatalysts due to their metabolic versatility and ability to transform hydrophobic substrates. In this study, we investigated the catalytic potential of Rhodococcus qingshengii IEGM 267 toward carveol isomers and explored genomic features that may underlie this activity. Methods: The strain was cultivated in mineral medium supplemented with (-)-trans-carveol. Biotransformation products were analyzed by TLC and GC-MS. The draft genome was sequenced, assembled, taxonomically assigned, and annotated using standard bioinformatics tools. Results: Rhodococcus qingshengii IEGM 267 efficiently converted (-)-trans-carveol to carvone. Genome analysis confirmed the taxonomic assignment of the strain and revealed a large repertoire of oxidoreductases, including monooxygenases, hydroxylases, and dehydrogenases. Seven genes encoding cytochrome P450-dependent oxygenases were identified as candidate enzymes potentially involved in carveol oxidation. Conclusions: R. qingshengii IEGM 267 is an efficient and stereoselective biocatalyst for (-)-trans-carveol oxidation. The results of bioinformatics analysis suggest an alternative enzymatic basis for this transformation and provide a foundation for future functional characterization.

Rhodococcus↗

Human genome. Storm erupts over terms for publishing Celera's sequence.

A dispute has been raging behind the scenes for weeks over the conditions under which Celera Genomics is prepared to make its human genome sequence data publicly available. The argument went public on 6 December, when geneticist Michael Ashburner e-mailed an open letter to Science's board of reviewing editors and members of the press slamming an agreement on data release that Science had reached with Celera as a condition for accepting its paper for review. This spat is the latest round in an intense rivalry between Celera president J. Craig Venter and leaders of the Human Genome Project, a publicly funded consortium that has produced its own draft human genome sequence.

Biotechnology↗

Comparative genomics of nematodes.

Recent transcriptome and genome projects have dramatically expanded the biological data available across the phylum Nematoda. Here we summarize analyses of these sequences, which have revealed multiple unexpected results. Despite a uniform body plan, nematodes are more diverse at the molecular level than was previously recognized, with many species- and group-specific novel genes. In the genus Caenorhabditis, changes in chromosome arrangement, particularly local inversions, are also rapid, with breakpoints occurring at 50-fold the rate in vertebrates. Tylenchid plant parasitic nematode genomes contain several genes closely related to genes in bacteria, implicating horizontal gene transfer events in the origins of plant parasitism. Functional genomics techniques are also moving from Caenorhabditis elegans to application throughout the phylum. Soon, eight more draft nematode genome sequences will be available. This unique resource will underpin both molecular understanding of these most abundant metazoan organisms and aid in the examination of the dynamics of genome evolution in animals.

Animals↗