PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

454 sequencing put to the test using the complex genome of barley.

BACKGROUND: During the past decade, Sanger sequencing has been used to completely sequence hundreds of microbial and a few higher eukaryote genomes. In recent years, a number of alternative technologies became available, among them adaptations of the pyrosequencing procedure (i.e. "454 sequencing"), promising an approximately 100-fold increase in throughput over Sanger technology--an advancement which is needed to make large and complex genomes more amenable to full genome sequencing at affordable costs. Although several studies have demonstrated its potential usefulness for sequencing small and compact microbial genomes, it was unclear how the new technology would perform in large and highly repetitive genomes such as those of wheat or barley. RESULTS: To study its performance in complex genomes, we used 454 technology to sequence four barley Bacterial Artificial Chromosome (BAC) clones and compared the results to those from ABI-Sanger sequencing. All gene containing regions were covered efficiently and at high quality with 454 sequencing whereas repetitive sequences were more problematic with 454 sequencing than with ABI-Sanger sequencing. 454 sequencing provided a much more even coverage of the BAC clones than ABI-Sanger sequencing, resulting in almost complete assembly of all genic sequences even at only 9 to 10-fold coverage. To obtain highly advanced working draft sequences for the BACs, we developed a strategy to assemble large parts of the BAC sequences by combining comparative genomics, detailed repeat analysis and use of low-quality reads from 454 sequencing. Additionally, we describe an approach of including small numbers of ABI-Sanger sequences to produce hybrid assemblies to partly compensate the short read length of 454 sequences. CONCLUSION: Our data indicate that 454 pyrosequencing allows rapid and cost-effective sequencing of the gene-containing portions of large and complex genomes and that its combination with ABI-Sanger sequencing and targeted sequence analysis can result in large regions of high-quality finished genomic sequences.

Base Pairing↗

The Ensembl genome database project.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of the human genome sequence, with confirmed gene predictions that have been integrated with external data sources, and is available as either an interactive web site or as flat files. It is also an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements from sequence analysis to data storage and visualisation. The Ensembl site is one of the leading sources of human genome sequence annotation and provided much of the analysis for publication by the international human genome project of the draft genome. The Ensembl system is being installed around the world in both companies and academic sites on machines ranging from supercomputers to laptops.

Computational Biology↗

Comparative genomic analysis links karyotypic evolution with genomic evolution in the Indian muntjac (Muntiacus muntjak vaginalis).

The karyotype of Indian muntjacs (Muntiacus muntjak vaginalis) has been greatly shaped by chromosomal fusion, which leads to its lowest diploid number among the extant known mammals. We present, here, comparative results based on draft sequences of 37 bacterial artificial clones (BAC) clones selected by chromosome painting for this special muntjac species. Sequence comparison on these BAC clones uncovered sequence syntenic relationships between the muntjac genome and those of other mammals. We found that the muntjac genome has peculiar features with respect to intron size and evolutionary rates of genes. Inspection of more than 80 pairs of orthologous introns from 15 genes reveals a significant reduction in intron size in the Indian muntjac compared to that of human, mouse, and dog. Evolutionary analysis using 19 genes indicates that the muntjac genes have evolved rapidly compared to other mammals. In addition, we identified and characterized sequence composition of the first BAC clone containing a chromosomal fusion site. Our results shed new light on the genome architecture of the Indian muntjac and suggest that chromosomal rearrangements have been accompanied by other salient genomic changes.

Animals↗

Putative evolution of Myxococcus fulvus 124B02 plasmid pMF1 from a chromosomal segment in another Myxococcus species.

Myxobacteria or order Myxococcales (old nomenclature) or phylum Myxococcota (new terminology) are fascinating organisms well known for their diverse peculiar physiological, taxonomic, and genomic properties. Researchers have long sought to identify plasmids within these organisms, yet thus far, only two organisms from different families have been found to harbor a plasmid. This study delves into the putative evolution of one of these plasmids, i.e., pMF1 present in Myxococcus fulvus 124B02 in the suborder Cystobacterineae and family Myxococcaceae. Here, we first reannotated the pMF1 plasmid genome sequence and identified two additional open reading frames or putative genes which were not annotated until now. We further reported that all pMF1 plasmid genes depict homology with Myxococcus stipitatus CYD1 draft genome (contig 28) and a chromosomal segment of M. stipitatus DSM14675 in a syntenic manner, implying the presence of plasmid-like structure in M. stipitatus CYD1, integrated into its chromosome. To comprehend the relationship among these three species, we conducted phylogenetic analyses using 16S and concatenated housekeeping genes and genome-to-genome distance calculator (GGDC) analysis, which confirmed that M. stipitatus CYD1 is a distinct and novel species within the genus Myxococcus. Overall, this comparative genomic study sheds light on the putative emergence of the pMF1 plasmid from a common ancestor of closely related yet distinct species, M. stipitatus CYD1, possibly through the partition from its chromosome as a segment.IMPORTANCEMyxobacteria are not well known to have plasmids. Until now, only two organisms have been shown to have plasmids, raising a pertinent question about how these plasmids evolved randomly within the phylum Myxococcota. The study presented in this manuscript delves into the emergence of the pMF1 plasmid found in Myxococcus fulvus 124B02, a member of the suborder Cystobacterineae and family Myxococcaceae. Our research addresses this intriguing topic of plasmid identification and evolution within myxobacteria, which are a group of fascinating organisms that have garnered significant interest due to their diverse physiological, taxonomic, and genomic properties.

Plasmids↗

Genome-wide profiling of gene amplification and deletion in cancer.

Accumulations of genetic changes in somatic cells induce phenotypic transformations leading to cancer. Among these genetic changes, gene amplification and deletion are most frequently observed in several kinds of cancers. Amplification of oncogene and/or deletion of tumor suppressor gene, together with dysfunction of the gene by point mutation, are the main causes of cancer. Genome-wide analysis of amplification and deletion of genes in cancers is basic to resolving the mechanisms of carcinogenesis. Comparative genomic hybridization (CGH) developed in 1992 has been utilized to identify DNA copy number abnormalities in various kind of cancers and several reports have shown its usefulness in screening of the genes involved in carcinogenesis, and also in the identification of prognostic factors in cancer. We have shown that 1q23 gain is associated with neuroblastomas that are resistant to aggressive treatment, and have poor prognosis, and 1q and 13q gains are possibly related to drug resistance in ovarian cancers. Recently, the "rough draft" of the human genome was reported and we are ready to utilize the vast information on genomic sequences in cancer research. Moreover, microarray technology enables us to analyze more than ten thousand genes at a time and revealed genetic abnormalities in cancers at a genome-wide level. By combination of microarray and CGH, a powerful screening method for oncogenes and tumor suppressor genes in cancers, called array-CGH, has been developed by several groups. In this article, we overview these genome-wide analytical methods, CGH and array-CGH, and discuss their potential in molecular characterization of cancers.

Gene Amplification↗

An efficient algorithm for large-scale detection of protein families.

Detection of protein families in large databases is one of the principal research objectives in structural and functional genomics. Protein family classification can significantly contribute to the delineation of functional diversity of homologous proteins, the prediction of function based on domain architecture or the presence of sequence motifs as well as comparative genomics, providing valuable evolutionary insights. We present a novel approach called TRIBE-MCL for rapid and accurate clustering of protein sequences into families. The method relies on the Markov cluster (MCL) algorithm for the assignment of proteins into families based on precomputed sequence similarity information. This novel approach does not suffer from the problems that normally hinder other protein sequence clustering algorithms, such as the presence of multi-domain proteins, promiscuous domains and fragmented proteins. The method has been rigorously tested and validated on a number of very large databases, including SwissProt, InterPro, SCOP and the draft human genome. Our results indicate that the method is ideally suited to the rapid and accurate detection of protein families on a large scale. The method has been used to detect and categorise protein families within the draft human genome and the resulting families have been used to annotate a large proportion of human proteins.

Algorithms↗

[The human genome project in the year 2000].

The human genome project was officially launched in 1990. This program started with a mapping phase which led to the development of a genetic map, a physical map based on large DNA fragments and more recently, a map of genes. Since 1996, the programme has progressively shifted to massive sequencing. A spectacular acceleration has occurred during the last 12 months and about 90% of the sequence is at present available in a draft format. This will be soon followed by a more complete version and by the progressive completion of each of the 24 chromosomes, a few of which being already in a "finished" state. It is to be hoped that the genome sequence, which can be used to efficiently identify genes involved in Mendelian phenotypes and which will lead to a better understanding of the evolution process, will also allow us to address other questions, such as those involving multifactorial inheritance.

Human Genome Project↗

The Genomes of Oryza sativa: a history of duplications.

We report improved whole-genome shotgun sequences for the genomes of indica and japonica rice, both with multimegabase contiguity, or almost 1,000-fold improvement over the drafts of 2002. Tested against a nonredundant collection of 19,079 full-length cDNAs, 97.7% of the genes are aligned, without fragmentation, to the mapped super-scaffolds of one or the other genome. We introduce a gene identification procedure for plants that does not rely on similarity to known genes to remove erroneous predictions resulting from transposable elements. Using the available EST data to adjust for residual errors in the predictions, the estimated gene count is at least 38,000-40,000. Only 2%-3% of the genes are unique to any one subspecies, comparable to the amount of sequence that might still be missing. Despite this lack of variation in gene content, there is enormous variation in the intergenic regions. At least a quarter of the two sequences could not be aligned, and where they could be aligned, single nucleotide polymorphism (SNP) rates varied from as little as 3.0 SNP/kb in the coding regions to 27.6 SNP/kb in the transposable elements. A more inclusive new approach for analyzing duplication history is introduced here. It reveals an ancient whole-genome duplication, a recent segmental duplication on Chromosomes 11 and 12, and massive ongoing individual gene duplications. We find 18 distinct pairs of duplicated segments that cover 65.7% of the genome; 17 of these pairs date back to a common time before the divergence of the grasses. More important, ongoing individual gene duplications provide a never-ending source of raw material for gene genesis and are major contributors to the differences between members of the grass family.

Base Sequence↗

Chromosome localization analysis of genes strongly expressed in human visceral adipose tissue.

To understand fully the physiologic functions of visceral adipose tissue and to provide a basis for the identification of novel genes related to obesity and insulin resistance, the gene expression profiling of human visceral adipose tissue was established by using cDNA array. The characterization and chromosome localization of 400 expressed sequence tags (ESTs) strongly expressed in visceral adipose tissue were analyzed by searching PubMed, UniGene, the Human Genome Draft Database, and Location Data Base. Two hundred eighty-nine clones were classified into known genes among the 400 ESTs strongly expressed in the tissue. Among them, <20% have been previously reported to be expressed in adipose tissue. The chromosome localization of 389 ESTs strongly expressed in visceral adipose tissue showed that their relative abundance was significantly increased on chromosomes 1, 16, 19, 20, and 22 compared with the expected distribution of the same number of random genes. The intrachromosome distribution of the genes strongly expressed in visceral adipose tissue was concentrated in certain regions, such as 1p36.2-1p36.3, 6p21.3-6p22.1, 19p13.3 and 19q13.1. Among them, the region of 1p36.2-1p36.3 appeared to be specific for visceral adipose tissue. Interestingly, some genes playing an important role in the pathogenesis of insulin signal transduction and adipocyte differentiation, such as tumor necrosis factor-alpha and its receptors; CCAAT/enhancer-binding proteina; and phosphoinositide-3-kinase, regulatory subunit, polypeptide 2 (p85beta), were also localized in the concentrated regions, which may provide clues to identifying novel genes closely related to adipocyte function with potential pathophysiologic implications.

Adipocytes↗

New goals for the U.S. Human Genome Project: 1998-2003.

The Human Genome Project has successfully completed all the major goals in its current 5-year plan, covering the period 1993-98. A new plan, for 1998-2003, is presented, in which human DNA sequencing will be the major emphasis. An ambitious schedule has been set to complete the full sequence by the end of 2003, 2 years ahead of previous projections. In the course of completing the sequence, a "working draft" of the human sequence will be produced by the end of 2001. The plan also includes goals for sequencing technology development; for studying human genome sequence variation; for developing technology for functional genomics; for completing the sequence of Caenorhabditis elegans and Drosophila melanogaster and starting the mouse genome; for studying the ethical, legal, and social implications of genome research; for bioinformatics and computational studies; and for training of genome scientists.

Animals↗

Applications of the double-barreled data in whole-genome shotgun sequence assembly and analysis.

Double-barreled (DB) data have been widely used for the assembly of large genomes. Based on the experience of building the whole-genome working draft of Oryza sativa L. ssp. Indica, we present here the prevailing and improved uses of DB data in the assembly procedure and report on novel applications during the following data-mining processes such as acquiring precise insert fragment information of each clone across the genome, and a new kind of low-cost whole-genome microarray. With the increasing number of organisms being sequenced, we believe that DB data will play an important role both in other assembly procedures and in future genomic studies.

Cloning, Molecular↗

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high &#x3b2;-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50&#xa0;mM sodium acetate, 6.0 pH, and 60&#xa0;&#xb0;C temperature on the seventh day of production with a value of 155.77&#xa0;&#xb1;&#xa0;3.21&#xa0;U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately &#xb1;34.8-49.1&#xa0;kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862&#xa0;Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting↗

Use of bovine EST data and human genomic sequences to map 100 gene-specific bovine markers.

A system to use bovine EST data in conjunction with human genomic sequence to improve the bovine linkage map over the entire genome or on specific chromosomes was evaluated. Bovine EST sequence was used to provide primer sequences corresponding to bovine genes, while human genomic sequence directed primer design to flank introns and produce amplicons of appropriate size for efficient direct sequencing. The sequence tagged sites (STS) produced in this way from the four sires of the MARC reference families were examined for single nucleotide polymorphisms (SNPs) that could be used to map the corresponding genes. With this approach, along with a primer/extension mass spectrometry SNP genotyping assay, 100 ESTs were placed on the bovine genetic linkage map. The first 70 were chosen at random from bovine EST-human genomic comparisons. An additional 30 ESTs were successfully mapped to bovine Chromosome 19 (BTA19), and comparison of the resulting BTA19 map to the position of the corresponding human orthologs on the HSA17 draft sequences revealed differences in the spacing and order of genes. Over 80% of successful amplicons contained SNPs, indicating that this is an efficient approach to generating EST-associated genetic markers. We have demonstrated the feasibility of constructing a linkage map based on SNPs associated with ESTs and the plausibility of utilizing EST, comparative mapping information, and human sequence data to target regions of the bovine genome for SNP marker development.

Animals↗

Characterization of a draft chromosome-scale genome assembly for the mutton snapper, Lutjanus analis.

BACKGROUND: The mutton snapper (Lutjanus analis) is a reef fish commonly found in tropical waters of the Western Atlantic Ocean. Genomic studies of this species are needed to support conservation efforts and breeding programs. OBJECTIVE: Here, we report the development of a chromosome-scale reference assembly for the mutton snapper and conduct an initial comparative genomic analysis with other lutjanids. METHODS: The genome of one mutton snapper specimen was sequenced using PAC-Bio HiFi long reads and Illumina short reads. Contigs and scaffolds were assembled in the Flye pipeline and anchored using Hi-C proximity guided assembly. Gene prediction and functional annotations were obtained in AUGUSTUS and eggNOG-mapper, respectively. The mutton snapper genome was compared to those of other lutjanids to infer gene family evolution and chromosome synteny conservation. RESULTS: Assembly and polishing yielded 946 contigs and 926 scaffolds (N50 of 3.16&#xa0;Mb, complete BUSCO score 98.1%) that were anchored using Hi-C scaffolding in 24 draft chromosomes. The anchored assembly featured a N50 of 42.47&#xa0;Mb and contained 97.6% of the unanchored assembly length. The 24 mutton snapper chromosomes showed a one-to-one syntenic relationship with their counterparts in medaka, and other Lutjanids. AUGUSTUS predicted 29,023 genes, 24,335 of which (83.85%) could be functionally annotated. Gene family evolution analysis revealed 1,014 significantly expanded or contracted hierarchical ortholog groups in mutton snapper. Expansions and contractions were linked to several biological functions including growth, oocyte maturation, and response to exogenous stressors. CONCLUSION: The draft genome will be a valuable tool for forthcoming applied genomic studies of mutton snapper.

Animals↗

pp-Blast: a "pseudo-parallel" Blast.

We have developed a software called pp-Blast that uses the publicly available Blast package and PVM (parallel virtual machine) to partition a multi-sequence query across a set of nodes with replicated or shared databases. Benchmark tests show that pp-Blast running in a cluster of 14 PCs outperformed conventional Blast running in large servers. In addition, using pp-Blast and the cluster we were able to map all human cDNAs onto the draft of the human genome in less than 6 days. We propose here that the cost/benefit ratio of pp-Blast makes it appropriate for large-scale sequence analysis. The source code and configuration files for pp-Blast are available at http://www.ludwig.org.br/biocomp/tools/pp-blast.

Computing Methodologies↗

Rapid expansion of the Ly49 gene cluster in rat.

The cytotoxic activity of mouse natural killer cells is regulated in part through cell surface molecules belonging to the Ly49 multigene family. In mice, the genomic sequence of the Ly49 gene cluster has been examined in detail and this analysis provided a model of the expansion of this multigene family. In the present study, we have analyzed a 1.8-Mb region of the draft rat genome revealing surprising differences in size and gene content between the mouse and the rat Ly49 clusters. The rat cluster contains at least 36 Ly49 genes, including pseudogenes, while dot-plot analysis of the cluster reveals an equidistant spacing of genes, suggesting that duplication of genes in the cluster occurred through a mechanism similar to that in the mouse. Phylogenetic analysis of the predicted rat genes reveals a number of distinct gene clusters and indicates that the majority of gene duplication events occurred after the divergence of mice and rats. Thus, the rodent Ly49 locus is subject to extremely rapid gene amplification and diversification.

Animals↗

Phylogenomic approaches to common problems encountered in the analysis of low copy repeats: the sulfotransferase 1A gene family example.

BACKGROUND: Blocks of duplicated genomic DNA sequence longer than 1000 base pairs are known as low copy repeats (LCRs). Identified by their sequence similarity, LCRs are abundant in the human genome, and are interesting because they may represent recent adaptive events, or potential future adaptive opportunities within the human lineage. Sequence analysis tools are needed, however, to decide whether these interpretations are likely, whether a particular set of LCRs represents nearly neutral drift creating junk DNA, or whether the appearance of LCRs reflects assembly error. Here we investigate an LCR family containing the sulfotransferase (SULT) 1A genes involved in drug metabolism, cancer, hormone regulation, and neurotransmitter biology as a first step for defining the problems that those tools must manage. RESULTS: Sequence analysis here identified a fourth sulfotransferase gene, which may be transcriptionally active, located on human chromosome 16. Four regions of genomic sequence containing the four human SULT1A paralogs defined a new LCR family. The stem hominoid SULT1A progenitor locus was identified by comparative genomics involving complete human and rodent genomes, and a draft chimpanzee genome. SULT1A expansion in hominoid genomes was followed by positive selection acting on specific protein sites. This episode of adaptive evolution appears to be responsible for the dopamine sulfonation function of some SULT enzymes. Each of the conclusions that this bioinformatic analysis generated using data that has uncertain reliability (such as that from the chimpanzee genome sequencing project) has been confirmed experimentally or by a "finished" chromosome 16 assembly, both of which were published after the submission of this manuscript. CONCLUSION: SULT1A genes expanded from one to four copies in hominoids during intra-chromosomal LCR duplications, including (apparently) one after the divergence of chimpanzees and humans. Thus, LCRs may provide a means for amplifying genes (and other genetic elements) that are adaptively useful. Being located on and among LCRs, however, could make the human SULT1A genes susceptible to further duplications or deletions resulting in 'genomic diseases' for some individuals. Pharmacogenomic studies of SULT1Asingle nucleotide polymorphisms, therefore, should also consider examining SULT1A copy number variability when searching for genotype-phenotype associations. The latest duplication is, however, only a substantiated hypothesis; an alternative explanation, disfavored by the majority of evidence, is that the duplication is an artifact of incorrect genome assembly.

Animals↗

Hierarchical scaffolding with Bambus.

The output of a genome assembler generally comprises a collection of contiguous DNA sequences (contigs) whose relative placement along the genome is not defined. A procedure called scaffolding is commonly used to order and orient these contigs using paired read information. This ordering of contigs is an essential step when finishing and analyzing the data from a whole-genome shotgun project. Most recent assemblers include a scaffolding module; however, users have little control over the scaffolding algorithm or the information produced. We thus developed a general-purpose scaffolder, called Bambus, which affords users significant flexibility in controlling the scaffolding parameters. Bambus was used recently to scaffold the low-coverage draft dog genome data. Most significantly, Bambus enables the use of linking data other than that inferred from mate-pair information. For example, the sequence of a completed genome can be used to guide the scaffolding of a related organism. We present several applications of Bambus: support for finishing, comparative genomics, analysis of the haplotype structure of genomes, and scaffolding of a mammalian genome at low coverage. Bambus is available as an open-source package from our Web site.

Algorithms↗