PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Computational methods for the identification of differential and coordinated gene expression.

With the first complete 'draft' of the human genome sequence expected for Spring 2000, the three basic challenges for today's bioinformatics are more than ever: (i) finding the genes; (ii) locating their coding regions; and (iii) predicting their functions. However, our capacity for interpreting vertebrate genomic and transcript (cDNA) sequences using experimental or computational means very much lags behind our raw sequencing power. If the performances of current programs in identifying internal coding exons are good, the precise 5'-->3' delineation of transcription units (and promoters) still requires additional experiments. Similarly, functional predictions made with reference to previously characterized homologues are leaving >50% of human genes unannotated or classified in uninformative categories ('kinase', 'ATP-binding', etc.). In the context of functional genomics, large-scale gene expression studies using massive cDNA tag sequencing, two-dimensional gel proteome analysis or microarray technologies are the only approaches providing genome-scale experimental information at a pace consistent with the progress of sequencing. Given the difficulty and cost of characterizing genes one by one, academic and industrial researchers are increasingly relying on those methods to prioritize their studies and choose their targets. The study of expression patterns can also provide some insight into the function, reveal regulatory pathways, indicate side effects of drugs or serve as a diagnostic tool. In this article, I review the theoretical and computational approaches used to: (i) identify genes differentially expressed (across cell types, developmental stages, pathological conditions, etc.); (ii) identify genes expressed in a coordinated manner across a set of conditions; and (iii) delineate clusters of genes sharing coherent expression features, eventually defining global biological pathways.

Animals↗

Adaptive diversification of bitter taste receptor genes in Mammalian evolution.

The diversity and evolution of bitter taste perception in mammals is not well understood. Recent discoveries of bitter taste receptor (T2R) genes provide an opportunity for a genetic approach to this question. We here report the identification of 10 and 30 putative T2R genes from the draft human and mouse genome sequences, respectively, in addition to the 23 and 6 previously known T2R genes from the two species. A phylogenetic analysis of the T2R genes suggests that they can be classified into three main groups, which are designated A, B, and C. Interestingly, while the one-to-one gene orthology between the human and mouse is common to group B and C genes, group A genes show a pattern of species- or lineage-specific duplication. It is possible that group B and C genes are necessary for detecting bitter tastants common to both humans and mice, whereas group A genes are used for species-specific bitter tastants. The analysis also reveals that phylogenetically closely related T2R genes are close in their chromosomal locations, demonstrating tandem gene duplication as the primary source of new T2Rs. For closely related paralogous genes, a rate of nonsynonymous nucleotide substitution significantly higher than the rate of synonymous substitution was observed in the extracellular regions of T2Rs, which are presumably involved in tastant-binding. This suggests the role of positive selection in the diversification of newly duplicated T2R genes. Because many natural poisonous substances are bitter, we conjecture that the mammalian T2R genes are under diversifying selection for the ability to recognize a diverse array of poisons that the organisms may encounter in exploring new habitats and diets.

Amino Acid Sequence↗

Insertion events of CR1 retrotransposable elements elucidate the phylogenetic branching order in galliform birds.

Using standard phylogenetic methods, it can be hard to resolve the order in which speciation events took place when new lineages evolved in the distant past and within a short time frame. As an example, phylogenies of galliform birds (including well-known species such as chicken, turkey, and quail) usually show low bootstrap support values at short internal branches, reflecting the rapid diversification of these birds in the Eocene. However, given the key role of chicken and related poultry species in agricultural, evolutionary, general biological and disease studies, it is important to know their internal relationships. Recently, insertion patterns of transposable elements such as long and short interspersed nuclear element markers have proved powerful in revealing branching orders of difficult phylogenies. Here we decipher the order of speciation events in a group of 27 galliform species based on insertion events of chicken repeat 1 (CR1) transposable elements. Forty-four CR1 marker loci were identified from the draft sequence of the chicken genome, and from turkey BAC clone sequence, and the presence or absence of markers across species was investigated via electrophoretic size separation of amplification products and subsequent confirmation by DNA sequencing. Thirty markers proved possible to type with electrophoresis of which 20 were phylogenetically informative. The distribution of these repeat elements supported a single homoplasy-free cladogram, which confirmed that megapodes, cracids, New World quail, and guinea fowl form outgroups to Phasianidae and that quails, pheasants, and partridges are each polyphyletic groups. Importantly, we show that chicken is an outgroup to turkey and quail, an observation which does not have significant support from previous DNA sequence- and DNA-DNA hybridization-based trees and has important implications for evolutionary studies based on sequence or karyotype data from galliforms. We discuss the potential and limitations of using a genome-based retrotransposon approach in resolving problematic phylogenies among birds.

Animals↗

The high diversity of snoRNAs in plants: identification and comparative study of 120 snoRNA genes from Oryza sativa.

Using a powerful computer-assisted analysis strategy, a large-scale search of small nucleolar RNA (snoRNA) genes in the recently released draft sequence of the rice genome was carried out. This analysis identified 120 different box C/D snoRNA genes with a total of 346 gene variants, which were predicted to guide 135 2'-O-ribose methylation sites in rice rRNAs. Though not exhaustive, this analysis has revealed that rice has the highest number of known box C/D snoRNAs among eukaryotes. Interestingly, although many snoRNA genes are conserved between rice and Arabidopsis, almost half of the identified snoRNA genes are rice specific, which may highlight further the differences in rRNA methylation patterns between monocotyledons and dicotyledons. In addition to 76 singletons, 70 clusters involving 270 snoRNA genes were also found in rice. The large number of the novel snoRNA polycistrons found in the introns of rice protein-coding genes is in contrast to the one-snoRNA-per-intron organization of vertebrates and yeast, and of Arabidopsis in which only a few intronic snoRNA gene clusters were identified. Furthermore, due to a high degree of gene duplication, rice snoRNA genes are clearly redundant and exhibit great sequence variation among isoforms, allowing generation of new snoRNAs for selection. Thus, the large snoRNA gene family in plants can serve as an excellent model for a rapid and functional evolution.

Base Sequence↗

Reporting, appraising, and integrating data on genotype prevalence and gene-disease associations.

The recent completion of the first draft of the human genome sequence and advances in technologies for genomic analysis are generating tremendous opportunities for epidemiologic studies to evaluate the role of genetic variants in human disease. Many methodological issues apply to the investigation of variation in the frequency of allelic variants of human genes, of the possibility that these influence disease risk, and of assessment of the magnitude of the associated risk. Based on a Human Genome Epidemiology workshop, a checklist for reporting and appraising studies of genotype prevalence and studies of gene-disease associations was developed. This focuses on selection of study subjects, analytic validity of genotyping, population stratification, and statistical issues. Use of the checklist should facilitate the integration of evidence from these studies. The relation between the checklist and grading schemes that have been proposed for the evaluation of observational studies is discussed. Although the limitations of grading schemes are recognized, a robust approach is proposed. Other issues in the synthesis of evidence that are particularly relevant to studies of genotype prevalence and gene-disease association are discussed, notably identification of studies, publication bias, criteria for causal inference, and the appropriateness of quantitative synthesis.

Case-Control Studies↗

Characterization of the intragenomic spread of the human endogenous retrovirus family HERV-W.

This study examines the intragenomic spread of the human endogenous retrovirus family HERV-W from insertions present within the draft sequence of the human genome. Identification of shared diagnostic differences and phylogenetic analyses revealed the existence of three main subfamilies. The average divergence between sequences for each of the subfamilies suggests that most of the HERV-W elements were inserted within the genome during a short period of evolutionary time. Each one of the subfamilies consists of two types of insertions, the expected proviral sequences and other sequences resembling the structure of processed retrogenes. These HERV-W retrosequences extend from the R region of the 5' long-terminal repeat (LTR) to the R region of the 3' LTR (as viral genomic RNAs), end in poly(A) 3' tails, and are flanked by direct repeats longer than the proviral integrations. Furthermore, several of the HERV-W retrosequences are 5'-truncated at different sites. I suggest the involvement of the L1 machinery in these integrations and discuss the characteristic features of the evolutionary history of HERV-W, with emphasis on the putative impact of HERV-W retrosequence integrations on the mammalian genome.

Base Sequence↗

Comparative genetics and evolution of annexin A13 as the founder gene of vertebrate annexins.

Annexin A13 (ANXA13) is believed to be the original founder gene of the 12-member vertebrate annexin A family, and it has acquired an intestine-specific expression associated with a highly differentiated intracellular transport function. Molecular characterization of this subfamily in a range of vertebrate species was undertaken to assess coding region conservation, gene organization, chromosomal linkage, and phylogenetic relationships relevant to its progenitor role in the structure-function evolution of the annexin gene superfamily. Protein diagnostic features peculiar to this subfamily include an alternate isoform containing a KGD motif, an elevated basic amino acid content with polyhistidine expansion in the 5'-translated region, and the conservation of 15% core tetrad residues specific to annexin A13 members. The 12 coding exons comprising the 58-kb human ANXA13 gene were deduced from BAC clone sequencing, whereas internal repetitive elements and neighboring genes in chromosome 8q24.12 were identified by contig analysis of the draft sequence from the human genome project. A unique exon splicing pattern in the annexin A13 gene was corroborated by coanalysis of mouse, rat, zebrafish, and pufferfish genomic DNA and determined to be the most distinct of all vertebrate annexins. The putative promoter region was identified by phylogenetic footprinting of potential binding sites for intestine-specific transcription factors. Mouse annexin A13 cDNA was used to map the gene to an orthologous linkage group in mouse chromosome 15 (between Sdc2 and Myc by backcross analysis), and the zebrafish cDNA permitted its localization to linkage group 24. Comparative analysis of annexin A13 from nine species traced this gene's speciation history and assessed coding region variation, whereas phylogenetic analysis showed it to be the deepest-branching vertebrate annexin, and computational analysis estimated the gene age and divergence rate. The unique, conserved aspects of annexin A13 primary structure, gene organization, and genetic maps identify it as the probable common ancestor of all vertebrate annexins, beginning with the sequential duplication to annexins A7 and A11 approximately 700 MYA, before the emergence of chordates.

Alternative Splicing↗

The sucrose transporter gene family in rice.

In this paper we report the identification, cloning and expression analysis of four putative sucrose transporter (SUT) genes from rice, designated OsSUT2, 3, 4 and 5. Three of the four genes were identified through extensive searches of the recently published draft sequence of the rice genome. Along with the previously reported OsSUT1 we propose that these five genes comprise the rice SUT gene family. Complementary DNA clones were isolated for the four newly identified genes. The deduced proteins of all five SUT genes were predicted to contain 12 membrane-spanning helices and a domain highly conserved throughout all known plant SUTs, suggesting the four additional OsSUT genes encode functional SUTs. Reverse transcription-PCR analysis was performed in order to investigate the expression pattern of each member of the SUT family in rice. A differing but overlapping expression pattern was observed for each member of the SUT family at different stages through plant development. These results, together with the structural variations apparent from the deduced protein sequences, suggest that the five SUTs possess diverse roles in both sink and source tissues. We also discuss the classification and evolution of the rice SUT gene family, using a comparison of the gene structures and deduced amino acid sequences with other known plant SUT genes.

Amino Acid Sequence↗

Recent advances in breast cancer biology.

Within the past year, the draft sequence of the human genome was completed and made available to researchers worldwide. Recent advances in technology along with the vast amount of sequence data on the human genome now provide a previously unimagined means of defining the genetic architecture of cancer cells. Implicit in this approach is the ability to describe the evolution of that architecture as normal breast cells progress toward the malignant phenotype. Ongoing experiments involving the simultaneous analysis of the entire genome in a high-throughput manner are expected to reveal those genes and regulatory mechanisms that are critical at each step of progression toward malignancy, including (1) providing a growth advantage over normal cells, (2) maintaining the malignant state, (3) modulating response to therapy, and (4) developing metastatic potential. Once these data are available, the ability to design preventive, diagnostic, prognostic, and therapeutic tools directed at those targets will be within reach.

Antineoplastic Agents↗

Minos transposon causes germline transgenesis of the ascidian Ciona savignyi.

An ascidian, Ciona savignyi, is regarded as a good experimental animal for genetics because of its small and compact genome for which a draft sequence is available, its short generation time and its interesting phylogenic position. ENU-based mutagenesis has been carried out using this animal. However, insertional mutagenesis using transposable elements (transposons) has not yet been introduced. Recently, one of the Tc1/mariner superfamily transposons, Minos, was demonstrated to cause germline transgenesis in the related species Ciona intestinalis. In this report, we show that Minos has the ability to transpose from DNA to DNA in Ciona savignyi in transposition assays. Although the activity was slightly weaker than in Ciona intestinalis, Minos still caused germline transgenesis in Ciona savignyi. In addition, one insertion seemed to have caused an enhancer trapping. These results indicate that Minos provides a potential tool for transgenic techniques such as insertional mutagenesis in Ciona savignyi.

Animals↗

Structure, organization, and transcriptional regulation of a family of copper radical oxidase genes in the lignin-degrading basidiomycete Phanerochaete chrysosporium.

The white rot basidiomycete Phanerochaete chrysosporium produces an array of nonspecific extracellular enzymes thought to be involved in lignin degradation, including lignin peroxidases, manganese peroxidases, and the H2O2-generating copper radical oxidase, glyoxal oxidase (GLX). Preliminary analysis of the P. chrysosporium draft genome had identified six sequences with significant similarity to GLX and designated cro1 through cro6. The predicted mature protein sequences diverge substantially from one another, but the residues coordinating copper and constituting the radical redox site are conserved. Transcript profiles, microscopic examination, and lignin analysis of inoculated thin wood sections are consistent with differential regulation as decay advances. The cro2-encoded protein was detected by liquid chromatography-tandem mass spectrometry in defined medium. The cro2 cDNA was successfully expressed in Aspergillus nidulans under the control of the A. niger glucoamylase promoter and secretion signal. The recombinant CRO2 protein had a substantially different substrate preference than GLX. The role of structurally and functionally diverse cro genes in lignocellulose degradation remains to be established.

Alcohol Oxidoreductases↗

Structure and expression of mobile ETnII retroelements and their coding-competent MusD relatives in the mouse.

ETnII elements are mobile members of the repetitive early transposon family of mouse long terminal repeat (LTR) retroelements and have caused a number of mutations by inserting into genes. ETnII sequences lack retroviral genes, but the recent discovery of related MusD retroviral elements with regions similar to gag, pro, and pol suggests that MusD provides the proteins necessary for ETnII transposition in trans. For this study, we analyzed all ETnII elements in the draft sequence of the C57BL/6J genome and classified them into three subtypes (alpha, beta, and gamma) based on structural differences. We then used database searches and quantitative real-time PCR to determine the copy number and expression of ETnII and MusD elements in various mouse strains. In 7.5-day-old embryos of a mouse strain in which two mutations due to ETnII-beta insertions have been identified (SELH/Bc), we detected a three- to sixfold higher level of ETnII-beta and MusD transcripts than in control strains (C57BL/6J and LM/Bc). The increased ETnII transcription level can in part be attributed to a higher number of ETnII-beta elements, but 70% of the MusD transcripts appear to have been derived from one or a few MusD elements that are not detectable in C57BL/6J mice. This element belongs to a young MusD subgroup with intact open reading frames and identical LTRs, suggesting that the overexpressed element(s) in SELH/Bc mice might provide the proteins for the retrotransposition of ETnII and MusD elements. We also show that ETnII is expressed up to 30-fold more than MusD, which could explain why only ETnII, but not MusD, elements have been positively identified as new insertions.

Animals↗

The genetics of tobacco use: methods, findings and policy implications.

Research on the genetics of smoking has increased our understanding of nicotine dependence, and it is likely to illuminate the mechanisms by which cigarette smoking adversely effects the health of smokers. Given recent advances in molecular biology, including the completion of the draft sequence of the human genome, interest has now turned to identifying gene markers that predict a heightened risk of using tobacco and developing nicotine dependence

Adoption↗

Comparative mapping of Homo sapiens chromosome 4 (HSA4) and Sus scrofa chromosome 8 (SSC8) using orthologous genes representing different cytogenetic bands as landmarks.

The recently published draft sequence of the human genome will provide a basic reference for the comparative mapping of genomes among mammals. In this study, we selected 214 genes with complete coding sequences on Homo sapiens chromosome 4 (HSA4) to search for orthologs and expressed sequence tag (EST) sequences in eight other mammalian species (cattle, pig, sheep, goat, horse, dog, cat, and rabbit). In particular, 46 of these genes were used as landmarks for comparative mapping of HSA4 and Sus scrofa chromosome 8 (SSC8); most of HSA4 is homologous to SSC8, which is of particular interest because of its association with genes affecting the reproductive performance of pigs. As a reference framework, the 46 genes were selected to represent different cytogenetic bands on HSA4. Polymerase chain reaction (PCR) products amplified from pig DNA were directly sequenced and their orthologous status was confirmed by a BLAST search. These 46 genes, plus 11 microsatellite markers for SSC8, were typed against DNA from a pig-mouse radiation hybrid (RH) panel with 110 lines. RHMAP analysis assigned these 57 markers to 3 linkage groups in the porcine genome, 52 to SSC8, 4 to SSC15, and 1 to SSC17. By comparing the order and orientation of orthologous landmark genes on the porcine RH maps with those on the human sequence map, HSA4 was recognized as being split into nine conserved segments with respect to the porcine genome, seven with SSC8, one with SSC15, and one with SSC17. With 41 orthologous gene loci mapped, this report provides the largest functional gene map of SSC8, with 30 of these loci representing new single-gene assignments to SSC8.

Animals↗

Molecular archeology of L1 insertions in the human genome.

BACKGROUND: As the rough draft of the human genome sequence nears a finished product and other genome-sequencing projects accumulate sequence data exponentially, bioinformatics is emerging as an important tool for studies of transposon biology. In particular, L1 elements exhibit a variety of sequence structures after insertion into the human genome that are amenable to computational analysis. We carried out a detailed analysis of the anatomy and distribution of L1 elements in the human genome using a new computer program, TSDfinder, designed to identify transposon boundaries precisely. RESULTS: Structural variants of L1 elements shared similar trends in the length and quality of their target site duplications (TSDs) and poly(A) tails. Furthermore, we found no correlation between the composition and genomic location of the pre-insertion locus and the resulting anatomy of the L1 insertion. We verified that L1 insertions with TSDs have the 5'-TTAAAA-3' cleavage site associated with L1 endonuclease activity. In addition, the second target DNA cut required for L1 insertion weakly matches the consensus pattern TTAAAA. On the other hand, the L1-internal breakpoints of deleted and inverted L1 elements do not resemble L1 endonuclease cleavage sites. Finally, the genome sequence data indicate that whereas singly inverted elements are common, doubly inverted elements are almost never found. CONCLUSIONS: The sequence data give no indication that the creation of L1 structural variants depends on characteristics of the insertion locus. In addition, the formation of 5' truncated and 5' inverted L1s are probably not due to the action of the L1 endonuclease.

Algorithms↗

Pharmacogenetics, pharmacogenomics and airway disease.

The availability of a draft sequence for the human genome will revolutionise research into airway disease. This review deals with two of the most important areas impinging on the treatment of patients: pharmacogenetics and pharmacogenomics. Considerable inter-individual variation exists at the DNA level in targets for medication, and variability in response to treatment may, in part, be determined by this genetic variation. Increased knowledge about the human genome might also permit the identification of novel therapeutic targets by expression profiling at the RNA (genomics) or protein (proteomics) level. This review describes recent advances in pharmacogenetics and pharmacogenomics with regard to airway disease.

Animals↗