PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Comparative DNA sequence analysis of mapped wheat ESTs reveals the complexity of genome relationships between rice and wheat.

The use of DNA sequence-based comparative genomics for evolutionary studies and for transferring information from model species to related large-genome species has revolutionized molecular genetics and breeding strategies for improving those crops. Comparative sequence analysis methods can be used to cross-reference genes between species maps, enhance the resolution of comparative maps, study patterns of gene evolution, identify conserved regions of the genomes, and facilitate interspecies gene cloning. In this study, 5,780 Triticeae ESTs that have been physically mapped using wheat ( Triticum aestivum L.) deletion lines and segregating populations were compared using NCBI BLASTN to the first draft of the public rice ( Oryza sativa L.) genome sequence data from 3,280 ordered BAC/PAC clones. A rice genome view of the homoeologous wheat genome locations based on sequence analysis shows general similarity to the previously published comparative maps based on Southern analysis of RFLP. For most rice chromosomes there is a preponderance of wheat genes from one or two wheat chromosomes. The physical locations of non-conserved regions were not consistent across rice chromosomes. Some wheat ESTs with multiple wheat genome locations are associated with the non-conserved regions of similarity between rice and wheat. The inverse view, showing the relationship between the wheat deletion map and rice genomic sequence, revealed the breakdown of gene content and order at the resolution conferred by the physical chromosome deletions in the wheat genome. An average of 35% of the putative single copy genes that were mapped to the most conserved bins matched rice chromosomes other than the one that was most similar. This suggests that there has been an abundance of rearrangements, insertions, deletions, and duplications eroding the wheat-rice genome relationship that may complicate the use of rice as a model for cross-species transfer of information in non-conserved regions.

Chromosome Mapping↗

HTS in the new millennium: the role of pharmacology and flexibility.

Over the past decade, high throughput screening (HTS) has become the focal point for discovery programs within the pharmaceutical industry. The role of this discipline has been and remains the rapid and efficient identification of lead chemical matter within chemical libraries for therapeutics development. Recent advances in molecular and computational biology, i.e., genomic sequencing and bioinformatics, have resulted in the announcement of publication of the first draft of the human genome. While much work remains before a complete and accurate genomic map will be available, there can be no doubt that the number of potential therapeutic intervention points will increase dramatically, thereby increasing the workload of early discovery groups. One current drug discovery paradigm integrates genomics, protein biosciences and HTS in establishing what the authors refer to as the "gene-to-screen" process. Adoption of the "gene-to-screen" paradigm results in a dramatic increase in the efficiency of the process of converting a novel gene coding for a putative enzymatic or receptor function into a robust and pharmacologically relevant high throughput screen. This article details aspects of the identification of lead chemical matter from HTS. Topics discussed include portfolio composition (molecular targets amenable to small molecule drug discovery), screening file content, assay formats and plating densities, and the impact of instrumentation on the ability of HTS to identify lead chemical matter.

Animals↗

Sequence information can be obtained from single DNA molecules.

The completion of the human genome draft has taken several years and is only the beginning of a period in which large amounts of DNA and RNA sequence information will be required from many individuals and species. Conventional sequencing technology has limitations in cost, speed, and sensitivity, with the result that the demand for sequence information far outstrips current capacity. There have been several proposals to address these issues by developing the ability to sequence single DNA molecules, but none have been experimentally demonstrated. Here we report the use of DNA polymerase to obtain sequence information from single DNA molecules by using fluorescence microscopy. We monitored repeated incorporation of fluorescently labeled nucleotides into individual DNA strands with single base resolution, allowing the determination of sequence fingerprints up to 5 bp in length. These experiments show that one can study the activity of DNA polymerase at the single molecule level with single base resolution and a high degree of parallelization, thus providing the foundation for a practical single molecule sequencing technology.

Base Sequence↗

Whole-genome sequence of Streptococcus agalactiae strain GIFTS31 isolated from streptococcosis-infected Nile tilapia in Bangladesh.

Streptococcus agalactiae strain GIFTS31 was isolated from a Nile tilapia infected with streptococcosis in Gazipur, Bangladesh. The draft genome of GIFTS31 comprises 2,039,674 bp with a GC content of 35% and encodes 1,957 predicted protein-coding sequences. The genome sequence provides valuable insights into the pathogenic potential of fish-associated S. agalactiae.

Streptococcus agalactiae↗

Interrogating the human genome using uninterpreted mass spectrometry data.

The public availability of a draft assembly of the human genome has enabled us to demonstrate, for the first time, the feasibility of searching a complete, unmasked eukaryotic genome using uninterpreted mass spectrometry data. A complex LC-MS/MS data set, containing peptides from at least 22 human proteins, was searched against a comprehensive, nonidentical protein database, an expressed sequence tag (EST) database, and the International Human Genome Project draft assembly of the human genome. The results from the three searches are compared in detail, and the merits of the different databases for this application are discussed. In the case of the EST database, the UniGene index provided a method of simplifying and summarising the search results. In the case of the genomic DNA, the presence of introns prevented matching of roughly one quarter of the spectra, but the technique can provide primary experimental verification of predicted coding sequences, and has the potential to identify novel coding sequences.

Algorithms↗

The human genome and public policy: a nursing perspective.

The Human Genome Project has completed a rough draft of the sequence that comprises human DNA. For the first time, the scientific discoveries have been conducted in tandem with research exploring the social, legal, and ethical implications of the research. Several committees have studied the policy implications of the genetics information explosion. The Secretary's Advisory Committee on Genetic Testing's report was received by the Secretary of the U.S. Department of Health and Human Services in November 2000 and contains recommendations for the oversight of genetic testing. Public policy issues that affect nursing practice in the area of genetics must be explored by individual nurses and professional nursing organizations.

Genetic Engineering↗

Genomic mining type III secretion system effectors in Pseudomonas syringae yields new picks for all TTSS prospectors.

Many bacterial pathogens of plants and animals use a type III secretion system (TTSS) to deliver virulence effector proteins into host cells. Because effectors are heterogeneous in sequence and function, there has not been a systematic way to identify the genes encoding them in pathogen genomes, and our current inventories are probably incomplete. A pre-closure draft sequence of Pseudomonas syringae pv. tomato DC3000, a pathogen of tomato and Arabidopsis, has recently supported five complementary studies which, collectively, identify 36 TTSS-secreted proteins and many more candidate effectors in this strain. These studies demonstrate the advantages of combining experimental and computational approaches, and they yield new insights into TTSS effectors and virulence regulation in P. syringae, potential effector targeting signals in all TTSS-dependent pathogens, and strategies for finding TTSS effectors in other bacteria that have sequenced genomes.

Arabidopsis↗

Computational analysis of full-length mouse cDNAs compared with human genome sequences.

Although the sequencing of the human genome is complete, identification of encoded genes and determination of their structures remain a major challenge. In this report, we introduce a method that effectively uses full-length mouse cDNAs to complement efforts in carrying out these difficult tasks. A total of 61,227 RIKEN mouse cDNAs (21,076 full-length and 40,151 EST sequences containing certain redundancies) were aligned with the draft human sequences. We found 35,141 non-redundant genomic regions that showed a significant alignment with the mouse cDNAs. We analyzed the structures and compositional properties of the regions detected by the full-length cDNAs, including cross-species comparisons, and noted a systematic bias of GENSCAN against exons of small size and/or low GC-content. Of the cDNAs locating the 35,141 genomic regions, 3,217 did not match any sequences of the known human genes or ESTs. Among those 3,217 cDNAs, 1,141 did not show any significant similarity to any protein sequence in the GenBank non-redundant protein database and thus are candidates for novel genes.

Algorithms↗

Genome-derived vaccines.

Vaccine research entered a new era when the complete genome of a pathogenic bacterium was published in 1995. Since then, more than 97 bacterial pathogens have been sequenced and at least 110 additional projects are now in progress. Genome sequencing has also dramatically accelerated: high-throughput facilities can draft the sequence of an entire microbe (two to four megabases) in 1 to 2 days. Vaccine developers are using microarrays, immunoinformatics, proteomics and high-throughput immunology assays to reduce the truly unmanageable volume of information available in genome databases to a manageable size. Vaccines composed by novel antigens discovered from genome mining are already in clinical trials. Within 5 years we can expect to see a novel class of vaccines composed by genome-predicted, assembled and engineered T- and Bcell epitopes. This article addresses the convergence of three forces--microbial genome sequencing, computational immunology and new vaccine technologies--that are shifting genome mining for vaccines onto the forefront of immunology research.

Animals↗

Molecular diversity of phospholipase D in angiosperms.

BACKGROUND: The phospholipase D (PLD) family has been identified in plants by recent molecular studies, fostered by the emerging importance of plant PLDs in stress physiology and signal transduction. However, the presence of multiple isoforms limits the power of conventional biochemical and pharmacological approaches, and calls for a wider application of genetic methodology. RESULTS: Taking advantage of sequence data available in public databases, we attempted to provide a prerequisite for such an approach. We made a complete inventory of the Arabidopsis thaliana PLD family, which was found to comprise 12 distinct genes. The current nomenclature of Arabidopsis PLDs was refined and expanded to include five newly described genes. To assess the degree of plant PLD diversity beyond Arabidopsis we explored data from rice (including the genome draft by Monsanto) as well as cDNA and EST sequences from several other plants. Our analysis revealed two major PLD subfamilies in plants. The first, designated C2-PLD, is characterised by presence of the C2 domain and comprises previously known plant PLDs as well as new isoforms with possibly unusual features catalytically inactive or independent on Ca2+. The second subfamily (denoted PXPH-PLD) is novel in plants but is related to animal and fungal enzymes possessing the PX and PH domains. CONCLUSIONS: The evolutionary dynamics, and inter-specific diversity, of plant PLDs inferred from our phylogenetic analysis, call for more plant species to be employed in PLD research. This will enable us to obtain generally valid conclusions.

Journal Article↗

Assembly, annotation, and integration of UNIGENE clusters into the human genome draft.

The recent release of the first draft of the human genome provides an unprecedented opportunity to integrate human genes and their functions in a complete positional context. However, at least three significant technical hurdles remain: first, to assemble a complete and nonredundant human transcript index; second, to accurately place the individual transcript indices on the human genome; and third, to functionally annotate all human genes. Here, we report the extension of the UNIGENE database through the assembly of its sequence clusters into nonredundant sequence contigs. Each resulting consensus was aligned to the human genome draft. A unique location for each transcript within the human genome was determined by the integration of the restriction fingerprint, assembled genomic contig, and radiation hybrid (RH) maps. A total of 59,500 UNIGENE clusters were mapped on the basis of at least three independent criteria as compared with the 30,000 human genes/ESTs currently mapped in Genemap'99. Finally, the extension of the human transcript consensus in this study enabled a greater number of putative functional assignments than the 11,000 annotated entries in UNIGENE. This study reports a draft physical map with annotations for a majority of the human transcripts, called the Human Index of Nonredundant Transcripts (HINT). Such information can be immediately applied to the discovery of new genes and the identification of candidate genes for positional cloning.

Alleles↗

WindowMasker: window-based masker for sequenced genomes.

MOTIVATION: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes. RESULTS: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis. AVAILABILITY: WM is included in the NCBI C++ toolkit. The source code for the entire toolkit is available at ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools++/CURRENT/. Once the toolkit source is unpacked, the instructions for building WindowMasker application in the UNIX environment can be found in file src/app/winmasker/README.build. SUPPLEMENTARY INFORMATION: Supplementary data are available at ftp://ftp.ncbi.nlm.nih.gov/pub/agarwala/windowmasker/windowmasker_suppl.pdf

Algorithms↗

Canonical TTAGG-repeat telomeres and telomerase in the honey bee, Apis mellifera.

The draft assembly of the honey bee Apis mellifera genome sequence reveals that the 17 centromeric-distal telomeres are of a simple, shared, and canonical structure, with 3-4 kb of a unique subtelomeric sequence, followed by several kilobases of TTAGG or variant telomeric repeats. This simple subtelomeric structure differs from the centromeric-proximal telomeres on the short arms of the 15 acrocentric chromosomes, which are apparently composed primarily of the 176-bp AluI tandem repeat. This dichotomy between the distal and proximal telomeres may involve differential participation of the telomeres of the 15 acrocentric chromosomes in the Rabl configuration after mitosis and the chromosome bouquet in meiotic prophase I. As expected from the presence of canonical TTAGG telomeric repeats, we identified a candidate telomerase gene in the bee, as well as the silkmoth Bombyx mori and the flour beetle Tribolium castaneum.

Animals↗

Reference based annotation with GeneMapper.

We introduce GeneMapper, a program for transferring annotations from a well annotated genome to other genomes. Drawing on high quality curated annotations, GeneMapper enables rapid and accurate annotation of newly sequenced genomes and is suitable for both finished and draft genomes. GeneMapper uses a profile based approach for mapping genes into multiple species, improving upon the standard pairwise approach. GeneMapper is freely available for academic use.

Algorithms↗

Transposable elements create distinct genomic niches for effector evolution among Magnaporthe oryzae lineages.

BACKGROUND: Plant-pathogen interactions are characterized by evolutionary arms races. At the molecular level, fungal effectors can target important plant functions, while plants evolve to improve effector recognition. Rapid evolution in genes encoding effectors can be facilitated by transposable elements (TEs). In Magnaporthe oryzae, the causal agent of blast disease in several cereals and grasses, TEs play important roles in chromosomal evolution as well as the gain or loss of effector genes in host specialized lineages. However, a global understanding of TE dynamics driving effector evolution at population scale and across lineages is lacking. RESULTS: Here, we focus on 16 AVR effector loci assessed across a global sampling of 11 reference genomes and 447 newly generated draft genome assemblies from publicly available short-read sequencing data across all major M. oryzae lineages and outgroups. We classified each effector based on evidence for duplication, deletion and translocation processes among lineages. Next, we determined AVR gain and loss dynamics across lineages allowing for a broad categorization of effector dynamics. Each AVR was integrated in a distinct genomic niche determined by the TE activity profile contributing to the diversification at the locus. We quantified TE contributions to effector niches and found that TE identity helped diversify AVR loci. We used the large genomic dataset to recapitulate the evolution of the rice blast AVR1-CO39 locus. CONCLUSIONS: Taken together, our work demonstrates how TE dynamics are an integral component of M. oryzae effector evolution, likely facilitating escape from host recognition. In-depth tracking of effector loci is a valuable tool to predict the durability of host resistance.

Ascomycota↗

Fluorescent in situ hybridization to ascidian chromosomes.

The draft genome of the ascidian Ciona intestinalis has been sequenced. Mapping of the genome sequence to the Ciona 14 haploid chromosomes is essential for future studies of the genome-wide control of gene expression in this basal chordate. Here we describe an efficient protocol for fluorescent in situ hybridization for mapping genes to the Ciona chromosomes. We demonstrate how the locations of two BAC clones can be mapped relative to each other. We also show that this method is efficient for coupling two so-far independent scaffolds into one longer scaffold when two BAC clones represent sequences located at either end of the two scaffolds.

Animals↗

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗