PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

The Anopheles gambiae genome: an update.

As a result of an international collaborative effort, the first draft of the Anopheles gambiae genome sequence and its preliminary annotation were published in October 2002. Since then, the assembly, annotation and means of accession of the An. gambiae genome have been under continuous development. This article reviews progress and considers limitations in the current sequence assembly and gene annotation, as well as approaches to address these problems and outstanding issues that users of the data must bear in mind.

Animals↗

Congenic mice: cutting tools for complex immune disorders.

Autoimmune diseases are, in general, under complex genetic control and subject to strong interactions between genetics and the environment. Greater knowledge of the underlying genetics will provide immunologists with a framework for study of the immune dysregulation that occurs in such diseases. Ascertaining the number of genes that are involved and their characterization have, however, proven to be difficult. Improved methods of genetic analysis and the availability of a draft sequence of the complete mouse genome have markedly improved the outlook for such research, and they have emphasized the advantages of mice as a model system. In this review, we provide an overview of the genetic analysis of autoimmune diseases and of the crucial role of congenic and consomic mouse strains in such research.

Animals↗

Identification of rat genes by TWINSCAN gene prediction, RT-PCR, and direct sequencing.

The publication of a draft sequence of a third mammalian genome--that of the rat--suggests a need to rethink genome annotation. New mammalian sequences will not receive the kind of labor-intensive annotation efforts that are currently being devoted to human. In this paper, we demonstrate an alternative approach: reverse transcription-polymerase chain reaction (RT-PCR) and direct sequencing based on dual-genome de novo predictions from TWINSCAN. We tested 444 TWINSCAN-predicted rat genes that showed significant homology to known human genes implicated in disease but that were partially or completely missed by methods based on protein-to-genome mapping. Using primers in exons flanking a single predicted intron, we were able to verify the existence of 59% of these predicted genes. We then attempted to amplify the complete predicted open reading frames of 136 genes that were verified in the single-intron experiment. Spliced sequences were amplified in 46 cases (34%). We conclude that this procedure for elucidating gene structures with native cDNA sequences is cost-effective and will become even more so as it is further optimized.

Animals↗

Exceptionally high levels of recombination across the honey bee genome.

The first draft of the honey bee genome sequence and improved genetic maps are utilized to analyze a genome displaying 10 times higher levels of recombination (19 cM/Mb) than previously analyzed genomes of higher eukaryotes. The exceptionally high recombination rate is distributed genome-wide, but varies by two orders of magnitude. Analysis of chromosome, sequence, and gene parameters with respect to recombination showed that local recombination rate is associated with distance to the telomere, GC content, and the number of simple repeats as described for low-recombining genomes. Recombination rate does not decrease with chromosome size. On average 5.7 recombination events per chromosome pair per meiosis are found in the honey bee genome. This contrasts with a wide range of taxa that have a uniform recombination frequency of about 1.6 per chromosome pair. The excess of recombination activity does not support a mechanistic role of recombination in stabilizing pairs of homologous chromosome during chromosome pairing. Recombination rate is associated with gene size, suggesting that introns are larger in regions of low recombination and may improve the efficacy of selection in these regions. Very few transposons and no retrotransposons are present in the high-recombining genome. We propose evolutionary explanations for the exceptionally high genome-wide recombination rate.

Animals↗

Novel architecture of family-9 glycoside hydrolases identified in cellulosomal enzymes of Acetivibrio cellulolyticus and Clostridium thermocellum.

We have sequenced a new gene, cel9B, encoding a family-9 cellulase from a cellulosome-producing bacterium, Acetivibrio cellulolyticus. The gene includes a signal peptide, a family-9 glycoside hydrolases (GH9) catalytic module, two family-3 carbohydrate-binding modules (CBM3c-CBM3b tandem dyad) and a C-terminal dockerin module. An identical modular arrangement exists in two putative GH9 genes from the draft sequence of the Clostridium thermocellum genome. The three homologous CBM3b modules from A. cellulolyticus and C. thermocellum were overexpressed, but, surprisingly, none bound cellulosic substrates. The results raise fundamental questions concerning the possible role(s) of the newly described CBMs. Phylogenetic analysis and preliminary site-directed mutagenesis studies suggest that the catalytic module and the CBM3 dyad are distinctive in their sequences and are proposed to constitute a new GH9 architectural theme.

Amino Acid Sequence↗

Discovery and profiling of bovine microRNAs from immune-related and embryonic tissues.

MicroRNAs are small approximately 22 nucleotide-long noncoding RNAs capable of controlling gene expression by inhibiting translation. Alignment of human microRNA stem-loop sequences (mir) against a recent draft sequence assembly of the bovine genome resulted in identification of 334 predicted bovine mir. We sequenced five tissue-specific cDNA libraries derived from the small RNA fractions of bovine embryo, thymus, small intestine, and lymph node to validate these predictions and identify new mir. This strategy combined with comparative sequence analysis identified 129 sequences that corresponded to mature microRNAs (miR). A total of 107 sequences aligned to known human mir, and 100 of these matched expressed miR. The other seven sequences represented novel miR expressed from the complementary strand of previously characterized human mir. The 22 sequences without matches displayed characteristic mir secondary structures when folded in silico, and 10 of these retained sequence conservation with other vertebrate species. Expression analysis based on sequence identity counts revealed that some miR were preferentially expressed in certain tissues, while bta-miR-26a and bta-miR-103 were prevalent in all tissues examined. These results support the premise that species differences in regulation of gene expression by miR occur primarily at the level of expression and processing.

Animals↗

Human mast cell transcriptome project.

After draft reading of the human genome sequence, systemic analysis of the transcriptome (the whole transcripts present in a cell) is progressing especially in commonly available cell types. Until recently, human mast cells were not commonly available. We have succeeded to generate a substantial number of human mast cells from umbilical cord blood and from adult peripheral blood progenitors. Then, we have examined messenger RNA selectively transcribed in these mast cells using high-density oligonucleotide probe arrays. Many unexpected but important transcripts were selectively expressed in human mast cells. We discuss the results obtained from transcriptome screening by introducing our data regarding mast-cell-specific genes.

Blood Cells↗

Bioinformatics in microbial biotechnology--a mini review.

The revolutionary growth in the computation speed and memory storage capability has fueled a new era in the analysis of biological data. Hundreds of microbial genomes and many eukaryotic genomes including a cleaner draft of human genome have been sequenced raising the expectation of better control of microorganisms. The goals are as lofty as the development of rational drugs and antimicrobial agents, development of new enhanced bacterial strains for bioremediation and pollution control, development of better and easy to administer vaccines, the development of protein biomarkers for various bacterial diseases, and better understanding of host-bacteria interaction to prevent bacterial infections. In the last decade the development of many new bioinformatics techniques and integrated databases has facilitated the realization of these goals. Current research in bioinformatics can be classified into: (i) genomics--sequencing and comparative study of genomes to identify gene and genome functionality, (ii) proteomics--identification and characterization of protein related properties and reconstruction of metabolic and regulatory pathways, (iii) cell visualization and simulation to study and model cell behavior, and (iv) application to the development of drugs and anti-microbial agents. In this article, we will focus on the techniques and their limitations in genomics and proteomics. Bioinformatics research can be classified under three major approaches: (1) analysis based upon the available experimental wet-lab data, (2) the use of mathematical modeling to derive new information, and (3) an integrated approach that integrates search techniques with mathematical modeling. The major impact of bioinformatics research has been to automate the genome sequencing, automated development of integrated genomics and proteomics databases, automated genome comparisons to identify the genome function, automated derivation of metabolic pathways, gene expression analysis to derive regulatory pathways, the development of statistical techniques, clustering techniques and data mining techniques to derive protein-protein and protein-DNA interactions, and modeling of 3D structure of proteins and 3D docking between proteins and biochemicals for rational drug design, difference analysis between pathogenic and non-pathogenic strains to identify candidate genes for vaccines and anti-microbial agents, and the whole genome comparison to understand the microbial evolution. The development of bioinformatics techniques has enhanced the pace of biological discovery by automated analysis of large number of microbial genomes. We are on the verge of using all this knowledge to understand cellular mechanisms at the systemic level. The developed bioinformatics techniques have potential to facilitate (i) the discovery of causes of diseases, (ii) vaccine and rational drug design, and (iii) improved cost effective agents for bioremediation by pruning out the dead ends. Despite the fast paced global effort, the current analysis is limited by the lack of available gene-functionality from the wet-lab data, the lack of computer algorithms to explore vast amount of data with unknown functionality, limited availability of protein-protein and protein-DNA interactions, and the lack of knowledge of temporal and transient behavior of genes and pathways.

Journal Article↗

Quantitative DNA fiber mapping in genome research and construction of physical maps.

Efforts to prepare a first draft of the human DNA genomic sequence forced multidisciplinary teams of researchers to face unique challenges. At the same time, these unprecedented obstacles stimulated the development of many highly innovative approaches to biomedical problem solving, robotics, and bioinformatics. High-resolution physical maps are required for ordering individual segments of information for the construction of a comprehensive map of the entire genome. This chapter describes a novel way to identify, delineate, and characterize selected, often small DNA sequences along a larger piece of the human genome. The technology is based on immobilization of high molecular weight DNA molecules on a solid substrate (such as a glass slide) followed by uniform stretching of the DNA molecule by the force of a receding meniscus. The hydrodynamic force stretches the DNA molecules homogeneously to approximately 2.3 kb/microm, so that distances measured after probe binding in microm can be converted directly into kb distances. Out of a large number of applications, this article focuses on mapping of genomic sequences relative to one another, the assembly of physical maps with near kb resolution, and, finally, quality control during physical map assembly and sequencing.

Chromosomes, Artificial↗

Genetic variation in immune function and susceptibility to human filariasis.

The generation of a draft sequence of a the human genome has provided the opportunity to characterize human diversity, even as it pertains to differences in host response to parasitic infection with organisms that cause lymphatic filariasis, malaria and schistosomiasis. Worldwide, human infection with filarial pathogens represents a significant cause of morbidity throughout the tropics. In particular, epidemiologic evidence suggests that a genetic component contributes to susceptibility and possibly the outcomes of filarial infection. Different approaches can be applied in population-based studies in areas where filarial infection is endemic, such as genome linkage scans and candidate gene analysis for the purpose of identifying genetic risk factors. This review summarizes recent advances in our understanding of genetic contributions to human lymphatic filariasis and addresses the immediate questions facing the field. It is anticipated that the identification of susceptibility genes in filarial infection could provide new insights into therapeutic strategies, including pharmacological intervention and vaccine development, and influence public health measures to control or avert infection.

Animals↗

Cloning and functional expression of a novel human connexin-25 gene.

Gap junctions are intercellular, water-filled channels composed of transmembrane proteins called connexins, six of which are arranged radially and dock with six homologous proteins in an adjacent cell to form an approximate 16 A pore. Through this pore cell-to-cell transfer of small water-soluble molecules up to about 1000 daltons occurs along concentration gradients. Connexins comprise a multigene family that share consensus sequences in the trans-membrane domains and the first and second extracellular loops. Comparison of the protein sequences of known human connexins with the draft nucleotide sequence of the human genome revealed two clones from chromosome 6 which showed strong similarity to highly conserved connexin sequences. Detailed analysis revealed the presence of a 672 nt open reading frame in these clones, encoding a 223 amino acid polypeptide with a predicted molecular weight of about 25 kD. This is smaller than other known human connexins. The ORF of the potential connexin25 was amplified by semi-nested PCR using human genomic DNA as a template. To confirm that this new gene encodes a connexin, Cx25 was transfected into a gap junction deficient subclone of the human HeLa cell line. After selection of transformants, cells were microinjected with the fluorescent dye Lucifer yellow. Transfectants but not controls successfully transferred dye, demonstrating that this new gene encodes a functional connexin.

Amino Acid Sequence↗

PlasmoDB: the Plasmodium genome resource. A database integrating experimental and computational data.

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates the recently completed P. falciparum genome sequence and annotation, as well as draft sequence and annotation emerging from other Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for intra- and inter-species comparisons. Sequence information is integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects and proteomics studies. The relational schema used to build PlasmoDB, GUS (Genomics Unified Schema) employs a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically-based, queries of the database. A stand-alone version of the database is also available on CD-ROM (P. falciparum GenePlot), facilitating access to the data in situations where internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to facilitate utilization of the vast quantities of genomic-scale data produced by the global malaria research community. The software used to develop PlasmoDB has been used to create a second Apicomplexan parasite genome database, ToxoDB (http://ToxoDB.org).

Animals↗

The 200-kb segmental duplication on human chromosome 21 originates from a pericentromeric dissemination involving human chromosomes 2, 18 and 13.

Regions close to human centromeres contain DNA fragments spanning hundreds of kilobases that exhibit a high degree of sequence identity (>95%). Here we report the genomic structure and evolution of a family of four paralogous regions related to a 220-kb genomic fragment present on the long arm of human chromosome 21 (21q22.1). Phylogenetic classification of the paralogous sequences obtained from the draft of the Human Genome Project are in agreement with results from comparative fluorescence in situ hybridization on metaphase chromosomes from human and great apes. The original copy present in 21q22.1 in human was duplicated in great apes after the divergence of the orang-utan and inserted in a pericentromeric region, most likely the ancestor of HSA2q, then disseminated by transposition of a larger fragment to other pericentromeric locations: HSA18p11, HSA13q11 and HSA21q11.1. The degree of dissemination varies among species.

Animals↗

Splice variation in mouse full-length cDNAs identified by mapping to the mouse genome.

We mapped the collection of The Institute of Physical and Chemical Research (Japan) (RIKEN) 21,076 full-length mouse cDNA clone sequences and the mouse RefSeq sequences to the recently completed draft of the mouse genome. Using this mapping, we identified 3674 mouse genes with multiple transcripts, of which 1098 have splice variants. All but 532 of 21,076 clones (97.5%) mapped to the genome assembly. Alignments of cDNA clone sequences with proteins show that much of the detected splice variation alters coding regions and affects the translated protein. We developed novel analytical techniques to classify observed splice variation and to assess the relation between splice variation and alternative transcription. This analysis indicates that an alternative choice of transcription start or polyadenylation signal frequently induces splice variation.

Alternative Splicing↗

Human-zebrafish non-coding conserved elements act in vivo to regulate transcription.

Whole genome comparisons of distantly related species effectively predict biologically important sequences--core genes and cis-acting regulatory elements (REs)--but require experimentation to verify biological activity. To examine the efficacy of comparative genomics in identification of active REs from anonymous, non-coding (NC) sequences, we generated a novel alignment of the human and draft zebrafish genomes, and contrasted this set to existing human and fugu datasets. We tested the transcriptional regulatory potential of candidate sequences using two in vivo assays. Strict selection of non-genic elements which are deeply conserved in vertebrate evolution identifies 1744 core vertebrate REs in human and two fish genomes. We tested 16 elements in vivo for cis-acting gene regulatory properties using zebrafish transient transgenesis and found that 10 (63%) strongly modulate tissue-specific expression of a green fluorescent protein reporter vector. We also report a novel quantitative enhancer assay with potential for increased throughput based on normalized luciferase activity in vivo. This complementary system identified 11 (69%; including 9 of 10 GFP-confirmed elements) with cis-acting function. Together, these data support the utility of comparative genomics of distantly related vertebrate species to identify REs and provide a scaleable, in vivo quantitative assay to define functional activity of candidate REs.

Animals↗

Composition and evolution of the V2r vomeronasal receptor gene repertoire in mice and rats.

Pheromones are chemicals produced and detected by conspecifics to elicit social/sexual physiological and behavioral responses, and they are perceived primarily by the vomeronasal organ (VNO) in terrestrial vertebrates. Two large superfamilies of G protein-coupled receptors, V1rs and V2rs, have been identified as pheromone receptors in vomeronasal sensory neurons. Based on a computational analysis of the mouse and rat genome sequences, we report the first global draft of the V2r gene repertoire, composed of approximately 200 genes and pseudogenes. Rodent V2rs are subject to rapid gene births/deaths and accelerated amino acid substitutions, likely reflecting the species-specific nature of pheromones. Vertebrate V2rs appear to have originated twice prior to the emergence of the VNO in ancestral tetrapods, explaining seemingly inconsistent observations among different V2rs. The identification of the entire V2r repertoire opens the door to genomic-level studies of the structure, function, and evolution of this diverse group of sensory receptors.

Amino Acid Sequence↗

A novel gene family NBPF: intricate structure generated by gene duplications during primate evolution.

Partial and complete genome duplications occurred during evolution and resulted in the creation of new genes and gene families. We identified a novel and intricate human gene family located primarily in regions of segmental duplications on human chromosome 1. We named it NBPF, for neuroblastoma breakpoint family, because one of its members is disrupted by a chromosomal translocation in a neuroblastoma patient. The NBPF genes have a repetitive structure with high intragenic and intergenic sequence similarity in both coding and noncoding regions. These similarities might expose these genomic regions to illegitimate recombination, resulting in structural variation in the NBPF genes. The encoded proteins contain a highly conserved domain of unknown function, which we have named the NBPF repeat. In silico analysis combined with the isolation of multiple full-length cDNA clones showed that several members of this gene family are abundantly expressed in a large variety of tissues and cell lines. Strikingly, no discernable orthologues could be identified in the completed genomes of fruit fly, nematode, mouse, or rat, but sequences with low homology could be isolated from the draft canine and bovine genomes. Interestingly, this gene family shows primate-specific duplications that result in species-specific arrays of NBPF homologous sequences. Overall, this novel NBPF family reflects the continuous evolution of primate genomes that resulted in large physiological differences, and its potential role in this process is discussed.

Amino Acid Sequence↗