PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Evidence that rice and other cereals are ancient aneuploids.

Detailed analyses of the genomes of several model organisms revealed that large-scale gene or even entire-genome duplications have played prominent roles in the evolutionary history of many eukaryotes. Recently, strong evidence has been presented that the genomic structure of the dicotyledonous model plant species Arabidopsis is the result of multiple rounds of entire-genome duplications. Here, we analyze the genome of the monocotyledonous model plant species rice, for which a draft of the genomic sequence was published recently. We show that a substantial fraction of all rice genes ( approximately 15%) are found in duplicated segments. Dating of these block duplications, their nonuniform distribution over the different rice chromosomes, and comparison with the duplication history of Arabidopsis suggest that rice is not an ancient polyploid, as suggested previously, but an ancient aneuploid that has experienced the duplication of one-or a large part of one-chromosome in its evolutionary past, approximately 70 million years ago. This date predates the divergence of most of the cereals, and relative dating by phylogenetic analysis shows that this duplication event is shared by most if not all of them.

Aneuploidy↗

Epistemological issues in omics and high-dimensional biology: give the people what they want.

Gene expression microarrays have been the vanguard of new analytic approaches in high-dimensional biology. Draft sequences of several genomes coupled with new technologies allow study of the influences and responses of entire genomes rather than isolated genes. This has opened a new realm of highly dimensional biology where questions involve multiplicity at unprecedented scales: thousands of genetic polymorphisms, gene expression levels, protein measurements, genetic sequences, or any combination of these and their interactions. Such situations demand creative approaches to the processes of inference, estimation, prediction, classification, and study design. Although bench scientists intuitively grasp the need for flexibility in the inferential process, the elaboration of formal supporting statistical frameworks is just at the very start. Here, we will discuss some of the unique statistical challenges facing investigators studying high-dimensional biology, describe some approaches being developed by statistical scientists, and offer an epistemological framework for the validation of proffered statistical procedures. A key theme will be the challenge in providing methods that a statistician judges to be sound and a biologist finds informative. The shift from family-wise error rate control to false discovery rate estimation and to assessment of ranking and other forms of stability will be portrayed as illustrative of approaches to this challenge.

Computational Biology↗

Fish tale.

The draft sequence of the genome of the Japanese pufferfish has just been announced and the temptation to humor is great.

Animals↗

Comprehensive search for cysteine cathepsins in the human genome.

Our study was aimed at examinating whether or not the human genome encodes for previously unreported cysteine cathepsins. To this end, we used analyses of the genome sequence and mRNA expression levels. The program TBLASTN was employed to scan the draft sequence of the human genome for the 11 known cysteine cathepsins. The cathepsin-like segments in the genome were inspected, filtered, and annotated. In addition to the known cysteine cathepsins, the scan identified three pseudogenes, closely related to cathepsin L, on chromosome 10, as well as two remote homologs, tubulointerstitial protein antigen and tubulointerstitial protein antigen-related protein. No new members of the family were identified. mRNA expression profiles for 10 known human cysteine cathepsins showed varying expression levels in 46 different human tissues and cell lines. No expression of any of the three cathepsin L-like pseudogenes was found. Based on these results, it is likely that to date all human cysteine cathepsins are known.

Cathepsins↗

Genomic organisation of the approximately 1.5 Mb Smith-Magenis syndrome critical interval: transcription map, genomic contig, and candidate gene analysis.

Smith-Magenis syndrome (SMS) is a multiple congenital anomalies/mental retardation syndrome associated with an interstitial deletion of chromosome 17 involving band p11.2. SMS is hypothesised to be a contiguous gene syndrome in which the phenotype arises from the haploinsufficiency of multiple, functionally-unrelated genes in close physical proximity, although the true molecular basis of SMS is not yet known. In this study, we have generated the first overlapping and contiguous transcription map of the SMS critical interval, linking the proximal 17p11.2 region near the SMS-REPM and the distal region near D17S740 in a minimum tiling path of 16 BACs and two PACs. Additional clones provide greater coverage throughout the critical region. Not including the repetitive sequences that flank the critical interval, the map is comprised of 13 known genes, 14 ESTs, and six genomic markers, and is a synthesis of Southern hybridisation and polymerase chain reaction data from gene and marker localisation to BACs and PACs and database sequence analysis from the human genome project high-throughput draft sequence. In order to identify possible candidate genes, we performed sequence analysis and determined the tissue expression pattern analysis of 10 novel ESTs that are deleted in all SMS patients. We also present a detailed review of six promising candidate genes that map to the SMS critical region.

Abnormalities, Multiple↗

In silco mapping of ESTs from the turkey (Meleagris gallopavo).

Sequence similarity was used to predict the position of expressed sequence tags (ESTs) in the genome of the turkey (Meleagris gallopavo). Turkey EST sequences were compared with the draft assembly of the chicken whole-genome sequence and the chicken EST database by BLASTN. Among the 877 ESTs examined, 788 had significant matches in the chicken genome sequence. Position of orthologous sequences in the chicken genome and the predicted position of the EST loci in the turkey genome are presented Genetic assignments suggest a high level of accuracy for the COMPASS predictions.

Animals↗

Towards the molecular dissection of fertilization signaling: Our functional genomic/proteomic strategies.

Recent advances in DNA sequencing techniques and automated informatics has led to clarification of all genome sequence of some model organisms in a very short period. The demonstration of the first draft sequence of the human genome has prompted us to elaborate new approaches in biology, pharmacology and medicine. Such new research will focus on high throughput methods to function on collections of genes, and hopefully, on a genome-wide, quantitative modeling of the cell system as a whole. In this review article, we discuss the present status of "post genome sequencing" approaches in line with our strategies for understanding the molecular mechanism of fertilization and activation of development using the African clawed frog, Xenopus laevis, as a model system.

Animals↗

'Gotta pick a megabase or two': in silico routes to gene regulation.

A number of vertebrate genome sequences are now available in draft or high-quality form. By comparing genomes from related species, conserved elements that are located outside the coding regions can be identified, many of which represent regulatory elements. The design of the sequence comparisons, taking into account the extent of the evolutionary divergence, is crucial to the outcome. Clearly, investigations of these conserved regulatory elements are important in understanding mechanisms underlying both vertebrate evolution and human disease.

Animals↗

A genome-wide survey of Major Histocompatibility Complex (MHC) genes and their paralogues in zebrafish.

BACKGROUND: The genomic organisation of the Major Histocompatibility Complex (MHC) varies greatly between different vertebrates. In mammals, the classical MHC consists of a large number of linked genes (e.g. greater than 200 in humans) with predominantly immune function. In some birds, it consists of only a small number of linked MHC core genes (e.g. smaller than 20 in chickens) forming a minimal essential MHC and, in fish, the MHC consists of a so far unknown number of genes including non-linked MHC core genes. Here we report a survey of MHC genes and their paralogues in the zebrafish genome. RESULTS: Using sequence similarity searches against the zebrafish draft genome assembly (Zv4, September 2004), 149 putative MHC gene loci and their paralogues have been identified. Of these, 41 map to chromosome 19 while the remaining loci are spread across essentially all chromosomes. Despite the fragmentation, a set of MHC core genes involved in peptide transport, loading and presentation are still found in a single linkage group. CONCLUSION: The results extend the linkage information of MHC core genes on zebrafish chromosome 19 and show the distribution of the remaining MHC genes and their paralogues to be genome-wide. Although based on a draft genome assembly, this survey demonstrates an essentially fragmented MHC in zebrafish.

Animals↗

A comparative analysis of HGSC and Celera human genome assemblies and gene sets.

MOTIVATION: Since the simultaneous publication of the human genome assembly by the International Human Genome Sequencing Consortium (HGSC) and Celera Genomics, several comparisons have been made of various aspects of these two assemblies. In this work, we set out to provide a more comprehensive comparative analysis of the two assemblies and their associated gene sets. RESULTS: The local sequence content for both draft genome assemblies has been similar since the early releases, however it took a year for the quality of the Celera assembly to approach that of HGSC, suggesting an advantage of HGSC's hierarchical shotgun (HS) sequencing strategy over Celera's whole genome shotgun (WGS) approach. While similar numbers of ab initio predicted genes can be derived from both assemblies, Celera's Otto approach consistently generated larger, more varied gene sets than the Ensembl gene build system. The presence of a non-overlapping gene set has persisted with successive data releases from both groups. Since most of the unique genes from either genome assembly could be mapped back to the other assembly, we conclude that the gene set discrepancies do not reflect differences in local sequence content but rather in the assemblies and especially the different gene-prediction methodologies.

Databases, Protein↗

Fugu ESTs: new resources for transcription analysis and genome annotation.

The draft Fugu rubripes genome was released in 2002, at which time relatively few cDNAs were available to aid in the annotation of genes. The data presented here describe the sequencing and analysis of 24,398 expressed sequence tags (ESTs) generated from 15 different adult and juvenile Fugu tissues, 74% of which matched protein database entries. Analysis of the EST data compared with the Fugu genome data predicts that approximately 10,116 gene tags have been generated, covering almost one-third of Fugu predicted genes. This represents a remarkable economy of effort. Comparison with the Washington University zebrafish EST assemblies indicates strong conservation within fish species, but significant differences remain. This potentially represents divergence of sequence in the 5' terminal exons and UTRs between these two fish species, although clearly, complete EST data sets are not available for either species. This project provides new Fugu resources, and the analysis adds significant weight to the argument that EST programs remain an essential resource for genome exploitation and annotation. This is particularly timely with the increasing availability of draft genome sequence from different organisms and the mounting emphasis on gene function and regulation.

Animals↗

DNA-based diagnosis of isolated sulfite oxidase deficiency by denaturing high-performance liquid chromatography.

Isolated sulfite oxidase deficiency is a rare autosomal recessive disease, characterized by severe neurological abnormalities, seizures, mental retardation, and dislocation of the ocular lenses, that often leads to death in infancy. There is a special demand for prenatal diagnosis, since no effective treatment is available for isolated sulfite oxidase deficiency. Until now, the cDNA sequence of the sulfite oxidase (SUOX) gene has been available, but the genomic sequence of the SUOX gene has not been published. In this study, we have performed a DNA-based diagnosis of isolated sulfite oxidase deficiency in a Chinese patient. To do so, we designed oligonucleotide primers for amplification of the predicted exons and intron-exon boundaries of the SUOX gene obtained from the completed draft version of the human genome. Using overlapping PCR products, we confirmed the flanking intronic sequences of the coding exons and that the entire 466-residue mature peptide is encoded by the last exon of the gene. We then performed mutation detection using denaturing high-performance liquid chromatography (DHPLC). The DHPLC chromatogram of exon 2b showed the presence of heteroduplex peaks only after mixing of the mutant DNA with the wild-type DNA, indicating the presence of a homozygous mutation. Direct DNA sequencing showed a homozygous base substitution at codon 160, changing the codon from CGG to CAG, which changes the amino acid from arginine to glutamine, i.e., R160Q. The DNA-based diagnosis of isolated sulfite oxidase deficiency will enable us to make an accurate determination of carrier status and to perform prenatal diagnosis of this disease. The availability of the genomic sequences of human genes from the completed draft human genome sequence will simplify the development of molecular genetic diagnoses of human diseases from peripheral blood DNA.

Child, Preschool↗

Identification of a tissue-specific putative transcription factor in breast tissue by serological screening of a breast cancer library.

Application of SEREX (serological analysis of recombinant tumor cDNA expression libraries) to different tumor types has led to the identification of several categories of human tumor antigens. In this study, the analysis of a breast cancer library with autologous patient serum led to the isolation of seven genes, designated NY-BR-1 through NY-BR-7. NY-BR-1, representing 6 of 14 clones isolated, showed tissue-restricted mRNA expression in breast and testis but not in 13 other normal tissues tested. Among tumor specimens, NY-BR-1 mRNA expression was found in 21 of 25 breast cancers but in only 2 of 82 nonmammary tumors. Structural analysis of NY-BR-1 cDNA and the corresponding genomic sequences in the recently released working draft of human genome indicated that NY-BR-1 is composed of 37 exons and has an open reading frame of 4.0-4.2 kb, encoding a peptide of Mr 150,000-160,000. A bipartite nuclear localization signal motif indicates a nuclear site for NY-BR-1, and the presence of a bZIP site (DNA-binding site followed by leucine zipper motif) suggests that NY-BR-1 is a transcription factor. Additional structural features include five tandem ankyrin repeats, implying a role for NY-BR-1 in protein-protein interactions. NY-BR-1 thus represents a breast tissue-specific putative transcription factor with autoimmunogenicity in breast cancer patients. In addition to NY-BR-1, a homologous gene, NY-BR-1.1, was identified in this study. NY-BR-1.1 shares 54% amino acid homology with NY-BR-1 and also shows tissue-restricted mRNA expression. However, unlike NY-BR-1, NY-BR-1.1 mRNA is expressed in brain, in addition to breast and testis. The exon structure of NY-BR-1.1 remains to be defined. Using human genome database, NY-BR-1 was localized to chromosome 10p11-p12, and NY-BR-1.1 was tentatively localized to chromosome 9.

Alternative Splicing↗

CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigation.

SUMMARY: Next Generation Sequencing is widely deployed in cholera-endemic regions, yet an end-to-end reproducible pipeline that unifies read QC, filtering, reference mapping, variant calling/annotation, recombination screening, and extraction of parsimony informative sites/variant codons, phylogenetic inference for downstream phylodynamic and epidemiological analyses have been lacking, slowing outbreak investigation and public health response. CholeraSeq is a high-throughput genomics pipeline for cholera genomic surveillance. It ingests consensus genomes, short read sequence data, draft assemblies, and scales seamlessly from local to cloud environments. To accelerate epidemiological context placement of new outbreak strains, we provide a curated ready-to-use core genome alignment compiled from public data, enabling flexible, fast, integration of new samples for outbreak investigations. AVAILABILITY AND IMPLEMENTATION: CholeraSeq is freely available on the GitHub platform https://github.com/CERI-KRISP/CholeraSeq. CholeraSeq is implemented in Nextflow with a modular design building upon the nf-core community standards.

Cholera↗

Strategies for the systematic sequencing of complex genomes.

Recent spectacular advances in the technologies and strategies for DNA sequencing have profoundly accelerated the detailed analysis of genomes from myriad organisms. The past few years alone have seen the publication of near-complete or draft versions of the genome sequence of several well-studied, multicellular organisms - most notably, the human. As well as providing data of fundamental biological significance, these landmark accomplishments have yielded important strategic insights that are guiding current and future genome-sequencing projects.

Animals↗

Analysis of a human brain transcriptome map.

BACKGROUND: Genome wide transcriptome maps can provide tools to identify candidate genes that are over-expressed or silenced in certain disease tissue and increase our understanding of the structure and organization of the genome. Expressed Sequence Tags (ESTs) from the public dbEST and proprietary Incyte LifeSeq databases were used to derive a transcript map in conjunction with the working draft assembly of the human genome sequence. RESULTS: Examination of ESTs derived from brain tissues (excluding brain tumor tissues) suggests that these genes are distributed on chromosomes in a non-random fashion. Some regions on the genome are dense with brain-enriched genes while some regions lack brain-enriched genes, suggesting a significant correlation between distribution of genes along the chromosome and tissue type. ESTs from brain tumor tissues have also been mapped to the human genome working draft. We reveal that some regions enriched in brain genes show a significant decrease in gene expression in brain tumors, and, conversely that some regions lacking in brain genes show an increased level of gene expression in brain tumors. CONCLUSIONS: This report demonstrates a novel approach for tissue specific transcriptome mapping using EST-based quantitative assessment.

Journal Article↗

Draft versus finished sequence data for DNA and protein diagnostic signature development.

Sequencing pathogen genomes is costly, demanding careful allocation of limited sequencing resources. We built a computational Sequencing Analysis Pipeline (SAP) to guide decisions regarding the amount of genomic sequencing necessary to develop high-quality diagnostic DNA and protein signatures. SAP uses simulations to estimate the number of target genomes and close phylogenetic relatives (near neighbors or NNs) to sequence. We use SAP to assess whether draft data are sufficient or finished sequencing is required using Marburg and variola virus sequences. Simulations indicate that intermediate to high-quality draft with error rates of 10(-3)-10(-5) (approximately 8x coverage) of target organisms is suitable for DNA signature prediction. Low-quality draft with error rates of approximately 1% (3x to 6x coverage) of target isolates is inadequate for DNA signature prediction, although low-quality draft of NNs is sufficient, as long as the target genomes are of high quality. For protein signature prediction, sequencing errors in target genomes substantially reduce the detection of amino acid sequence conservation, even if the draft is of high quality. In summary, high-quality draft of target and low-quality draft of NNs appears to be a cost-effective investment for DNA signature prediction, but may lead to underestimation of predicted protein signatures.

Computational Biology↗

Multiple gene organization of pufferfish Fugu rubripes tropomyosin isoforms and tissue distribution of their transcripts.

The Japanese pufferfish, torafugu (Fugu rubripes), has a haploid genome of about 400 Mb in size, which has been sequenced to approximately 90% coverage. Here we identified six Fugu tropomyosin (TPM) gene sequences by using the BLASTN program and the sequence of the white croaker TPM1 gene in our collection against the draft assembly of the Fugu genomic sequence database. TPM2, TPM3 and TPM4 genes were identified together with a set of two potentially duplicated genes of TPM1 (TPM1-1 and TPM1-2) as described in our previous report and TPM4 (TPM4-1 and TPM4-2) newly found in this study. The expression patterns of these Fugu TPM genes were determined by reverse transcription polymerase chain reaction (RT-PCR). A phylogenetic tree was constructed using the deduced amino acid sequences, which were encoded by the exons common to all vertebrate TPM genes. This indicated that the Fugu TPM1 and TPM4 genes had resulted from a gene duplication in the fish evolutionary lineage.

Alternative Splicing↗