PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Annotated genome assemblies of two temperate North American dung beetles, Canthon chalcites and Phanaeus vindex.

Dung beetles serve as cultivators of their natural habitats, improving soil health and functions in both natural and anthropogenic environments. Despite their ecological importance, whole genome sequences for Scarabaeinae are limited. Here, we present the draft annotated genome assemblies for 2 temperate species of North American dung beetles collected from eastern Tennessee: Canthon chalcites and Phanaeus vindex. Both genome assemblies were generated from PacBio long reads and have high completeness, with BUSCO scores of 98.1% and 98.6% for C. chalcites and P. vindex, respectively. For C. chalcites, the BRAKER3 pipeline predicted 12,799 genes, and the gene set was 93.7% complete. For P. vindex, the BRAKER3 predicted 12,252 genes, and the gene set was 94.9% complete. From the annotated gene sets, orthologous protein sequence analyses among C. chalcites, P. vindex, the dung beetle species Onthophagus taurus, and the more evolutionarily distant beetle Tribolium castaneum indicated that there are 260 unique protein clusters for C. chalcites and 210 unique protein clusters for P. vindex. These 2 draft genomes provide valuable data for comparative genomics, evolution, and phylogenic studies for dung beetle species.

Animals↗

A computational scan for U12-dependent introns in the human genome sequence.

U12-dependent introns are found in small numbers in most eukaryotic genomes, but their scarcity makes accurate characterisation of their properties challenging. A computational search for U12-dependent introns was performed using the draft version of the human genome sequence. Human expressed sequences confirmed 404 U12-dependent introns within the human genome, a 6-fold increase over the total number of non-redundant U12-dependent introns previously identified in all genomes. Although most of these introns had AT-AC or GT-AG terminal dinucleotides, small numbers of introns with a surprising diversity of termini were found, suggesting that many of the non-canonical introns found in the human genome may be variants of U12-dependent introns and, thus, spliced by the minor spliceosome. Comparisons with U2-dependent introns revealed that the U12-dependent intron set lacks the 'short intron' peak characteristic of U2-dependent introns. Analysis of this U12-dependent intron set confirmed reports of a biased distribution of U12-dependent introns in the genome and allowed the identification of several alternative splicing events as well as a surprising number of apparent splicing errors. This new larger reference set of U12-dependent introns will serve as a resource for future studies of both the properties and evolution of the U12 spliceosome.

Alternative Splicing↗

New apolipoprotein A-V: comparative genomics meets metabolism.

The availability of the human genome sequence and the recently completed draft sequences of two major mammalian model species, the mouse (Mus musculus) and the rat (Rattus norvegicus), allow researchers to apply novel approaches for gene identification and characterization, using methods of comparative and functional genomics. Recently, a new gene coding for apolipoprotein A-V was identified in the vicinity of APOA-I/C-III/A-IV cluster on human chromosome 11q23 by comparative sequencing method. In a relatively short time, compelling evidence accumulated for the substantial role of APOA-V in lipid metabolism. Studies in knock-out and transgenic mice revealed that its expression pattern correlates negatively with triglyceride levels. This observation was verified in human population studies in variety of ethnic and age groups. Several single nucleotide polymorphisms were described and particular SNP alleles and haplotypes in the APO A-V gene region were shown to be associated with dyslipidemia. The discovery and characterization of the APO A-V demonstrates current possibilities of the integrative approaches in biology, boosted by the available bioinformatic tools.

Amino Acid Sequence↗

Critically assessing the state-of-the-art in protein structure prediction.

One of the most tantalising 'grand challenges' in structural biology is to solve the problem of predicting the structure of a protein from its amino acid sequence alone. Although this problem appeals to many researchers on a purely academic level, the practical importance of protein structure prediction has become particularly clear with the release of the first draft of the complete human genome sequence last year. This moved modern biology into the new so-called 'post genome' era, and for the foreseeable future, one of the main issues in modern biology will be the characterisation of the many 'unknown' gene sequences which are now sitting waiting in DNA and protein data banks. Protein structure can provide a great deal of insight into the evolutionary origins, function and mechanism of a protein, and so any means for determining the 3-D structure of a novel protein will likely be of critical importance.

Animals↗

Genome sequence of the lignocellulose degrading fungus Phanerochaete chrysosporium strain RP78.

White rot fungi efficiently degrade lignin, a complex aromatic polymer in wood that is among the most abundant natural materials on earth. These fungi use extracellular oxidative enzymes that are also able to transform related aromatic compounds found in explosive contaminants, pesticides and toxic waste. We have sequenced the 30-million base-pair genome of Phanerochaete chrysosporium strain RP78 using a whole genome shotgun approach. The P. chrysosporium genome reveals an impressive array of genes encoding secreted oxidases, peroxidases and hydrolytic enzymes that cooperate in wood decay. Analysis of the genome data will enhance our understanding of lignocellulose degradation, a pivotal process in the global carbon cycle, and provide a framework for further development of bioprocesses for biomass utilization, organopollutant degradation and fiber bleaching. This genome provides a high quality draft sequence of a basidiomycete, a major fungal phylum that includes important plant and animal pathogens.

Base Composition↗

Sequencing a new target genome: the Boophilus microplus (Acari: Ixodidae) genome project.

The southern cattle tick, Boophilus microplus (Canestrini), causes annual economic losses in the hundreds of millions of dollars to cattle producers throughout the world, and ranks as the most economically important tick from a global perspective. Control failures attributable to the development of pesticide resistance have become commonplace, and novel control technologies are needed. The availability of the genome sequence will facilitate the development of these new technologies, and we are proposing sequencing to a 4-6X draft coverage. Many existing biological resources are available to facilitate a genome sequencing project, including several inbred laboratory tick strains, a database of approximately 45,000 expressed sequence tags compiled into a B. microplus Gene Index, a bacterial artificial chromosome (BAC) library, an established B. microplus cell line, and genomic DNA suitable for library synthesis. Collaborative projects are underway to map BACs and cDNAs to specific chromosomes and to sequence selected BAC clones. When completed, the genome sequences from the cow, B. microplus, and the B. microplus-borne pathogens Babesia bovis and Anaplasma marginale will enhance studies of host-vector-pathogen systems. Genes involved in the regeneration of amputated tick limbs and transitions through developmental stages are largely unknown. Studies of these and other interesting biological questions will be advanced by tick genome sequence data. Comparative genomics offers the prospect of new insight into many, perhaps all, aspects of the biology of ticks and the pathogens they transmit to farm animals and people. The B. microplus genome sequence will fill a major gap in comparative genomics: a sequence from the Metastriata lineage of ticks. The purpose of the article is to synergize interest in and provide rationales for sequencing the genome of B. microplus and for publicizing currently available genomic resources for this tick.

Animals↗

Characterization of the aryl hydrocarbon receptor repressor gene and association of its Pro185Ala polymorphism with micropenis.

BACKGROUND: Genetic background of a fetus contributes to the abnormal development after teratogen exposure. In rodents, in utero exposure to dioxins affects male external genital development. The effects of dioxins are mediated via the aryl hydrocarbon receptor (AHR) and its binding protein, aryl hydrocarbon receptor nuclear translocator (ARNT). In mice, aryl hydrocarbon receptor repressor (AHRR), which binds to ARNT in competition with AHR, plays a critical negative regulatory role in AHR signaling. We attempt to characterize the human AHRR gene and investigate the relationship between AHRR polymorphisms and the incidence of micropenis, a phenotype of undermasculinization. METHODS: We identified and characterized the human homolog of mouse AHRR, taking advantage of the publicly available draft version of the human genome sequence. After detecting an AHRR protein polymorphism by the direct sequencing of pooled human genomic DNA, we evaluated the association between the polymorphism and the presence or absence of micropenis (< -2.5 SD) in patients with micropenis and control subjects. RESULTS: The deduced sequence for human AHRR (715 residues) and the mouse AHRR protein exhibited 81% sequence homology to each other. The Pro185Ala polymorphism was identified between the PAS-A region and the highly conserved arginine/cysteine-rich RCFRCRL/VRC region. Forty-six percent (27/59) of patients with micropenis and 27% (22/80) of the controls were homozygous for 185Pro; this difference in frequencies was significant (P = 0.03). CONCLUSIONS: Homozygosity for the 185Pro allele of AHRR may increase the susceptibility of a fetus to the undermasculinizing effects of dioxin exposure in utero, presumably through the diminished inhibition of AHR-mediated signaling.

Alanine↗

BAC contig from a 3-cM region of mouse chromosome 11 surrounding Brca1.

Even with the completion of a draft version of the human genome sequence only a fraction of the genes identified from this sequence have known functions. Chromosomal engineering in mouse cells, in concert with gene replacement assays to prove the functional significance of a given genomic region or gene, represents a rapid and productive means for understanding the role of a given set of genes. Both techniques rely heavily on detailed maps of chromosomal regions, initially to understand the scope of the regions being modified and finally to provide the cloned resources necessary to allow both finished sequencing and large insert complementation. This report describes the creation of a BAC clone contig on mouse chromosome 11 in a region showing conservation of synteny with sequences on human chromosome 17. We have created a detailed map of an approximately 3-cM region containing at least 33 genes through the use of multiple BAC mapping strategies, including chromosome walking and multiplex oligonucleotide hybridization and gap filling. The region described is one of the targets of a large effort to create a series of mice with regional deletions on mouse chromosome 11 (33-80 cM) that can subsequently be subjected to further mutagenesis.

Animals↗

The misidentification of bacterial genes as human cDNAs: was the human D-1 tumor antigen gene acquired from bacteria?

The initial analysis of the draft copy of the human genome sequence revealed the presence of several genes that were proposed to have been directly transferred from bacteria. We investigated the human D-1 antigen as a potential lateral transfer event. We report that although the human D-1 antigen seems to be an excellent candidate for lateral transfer, it is a contaminating bacterial sequence present in a human cDNA library that was included in the human genome analysis. Furthermore, several other genes present in the publicly available databases that were included in the analysis of the human genome are also likely contaminating bacterial sequences present in cDNA libraries.

Antigens, Neoplasm↗

Structural and functional features of fructansucrases present in Leuconostoc mesenteroides ATCC 8293.

Glycosyltransferases produced by Leuconostoc mesenteroides subsp. mesenteroides ATCC 8293 (equivalent to NRRL B-1118) were identified. Two glucansucrases and one fructansucrases were observed in batch culture while levC and levL genes, corresponding to two fructansucrases, were isolated from information obtained from the released draft sequence of this Leuconostoc strain genome and cloned in Escherichia coli. The recombinant enzymes were shown to be fructansucrases producing a polymer identified by NMR as levan, confirming our recent report stating that these are also mosaic levansucrases bearing structural features of glucansucrases in the amino and carboxy terminal regions, as is also the case of inulosucrase (IslA) from Leuconostoc citreum CW28 and levansucrase (LevS) from L. mesenteroides NRRL B-512F. The recombinant levansucrase LevC was purified and characterized in terms of pH, temperature, and kinetic properties. The enzyme exhibits Michaelis-Menten kinetic properties with a K(m) = 27.3 mM and a k(cat) = 282.9 s(-1). This levansucrase behaves mainly as a transferase as only 30% of the substrate is hydrolyzed in a wide range of sucrose concentrations, with higher hydrolytic activities at low substrate concentrations. With this report we experimentally confirm the unusual structural pattern displayed by fructansucrases present in Leuconostoc species that group as a novel sub family of fructansucrases.

Amino Acid Sequence↗

The chick; a great model system becomes even greater.

The chick embryo has a long and distinguished history as a major model system in developmental biology and has also contributed major concepts to immunology, genetics, virology, cancer, and cell biology. Now, it has become even more powerful thanks to several new technologies: in vivo electroporation (allowing gain- and loss-of-function in vivo in a time- and space-controlled way), embryonic stem (ES) cells, novel methods for transgenesis, and the completion of the first draft of the sequence of its genome along with many new resources to access this information. In combination with classical techniques such as grafting and lineage tracing, the chicken is now one of the most versatile experimental systems available.

Animals↗

The Anopheles gambiae genome: an update.

As a result of an international collaborative effort, the first draft of the Anopheles gambiae genome sequence and its preliminary annotation were published in October 2002. Since then, the assembly, annotation and means of accession of the An. gambiae genome have been under continuous development. This article reviews progress and considers limitations in the current sequence assembly and gene annotation, as well as approaches to address these problems and outstanding issues that users of the data must bear in mind.

Animals↗

Congenic mice: cutting tools for complex immune disorders.

Autoimmune diseases are, in general, under complex genetic control and subject to strong interactions between genetics and the environment. Greater knowledge of the underlying genetics will provide immunologists with a framework for study of the immune dysregulation that occurs in such diseases. Ascertaining the number of genes that are involved and their characterization have, however, proven to be difficult. Improved methods of genetic analysis and the availability of a draft sequence of the complete mouse genome have markedly improved the outlook for such research, and they have emphasized the advantages of mice as a model system. In this review, we provide an overview of the genetic analysis of autoimmune diseases and of the crucial role of congenic and consomic mouse strains in such research.

Animals↗

Identification of rat genes by TWINSCAN gene prediction, RT-PCR, and direct sequencing.

The publication of a draft sequence of a third mammalian genome--that of the rat--suggests a need to rethink genome annotation. New mammalian sequences will not receive the kind of labor-intensive annotation efforts that are currently being devoted to human. In this paper, we demonstrate an alternative approach: reverse transcription-polymerase chain reaction (RT-PCR) and direct sequencing based on dual-genome de novo predictions from TWINSCAN. We tested 444 TWINSCAN-predicted rat genes that showed significant homology to known human genes implicated in disease but that were partially or completely missed by methods based on protein-to-genome mapping. Using primers in exons flanking a single predicted intron, we were able to verify the existence of 59% of these predicted genes. We then attempted to amplify the complete predicted open reading frames of 136 genes that were verified in the single-intron experiment. Spliced sequences were amplified in 46 cases (34%). We conclude that this procedure for elucidating gene structures with native cDNA sequences is cost-effective and will become even more so as it is further optimized.

Animals↗

Exceptionally high levels of recombination across the honey bee genome.

The first draft of the honey bee genome sequence and improved genetic maps are utilized to analyze a genome displaying 10 times higher levels of recombination (19 cM/Mb) than previously analyzed genomes of higher eukaryotes. The exceptionally high recombination rate is distributed genome-wide, but varies by two orders of magnitude. Analysis of chromosome, sequence, and gene parameters with respect to recombination showed that local recombination rate is associated with distance to the telomere, GC content, and the number of simple repeats as described for low-recombining genomes. Recombination rate does not decrease with chromosome size. On average 5.7 recombination events per chromosome pair per meiosis are found in the honey bee genome. This contrasts with a wide range of taxa that have a uniform recombination frequency of about 1.6 per chromosome pair. The excess of recombination activity does not support a mechanistic role of recombination in stabilizing pairs of homologous chromosome during chromosome pairing. Recombination rate is associated with gene size, suggesting that introns are larger in regions of low recombination and may improve the efficacy of selection in these regions. Very few transposons and no retrotransposons are present in the high-recombining genome. We propose evolutionary explanations for the exceptionally high genome-wide recombination rate.

Animals↗

Novel architecture of family-9 glycoside hydrolases identified in cellulosomal enzymes of Acetivibrio cellulolyticus and Clostridium thermocellum.

We have sequenced a new gene, cel9B, encoding a family-9 cellulase from a cellulosome-producing bacterium, Acetivibrio cellulolyticus. The gene includes a signal peptide, a family-9 glycoside hydrolases (GH9) catalytic module, two family-3 carbohydrate-binding modules (CBM3c-CBM3b tandem dyad) and a C-terminal dockerin module. An identical modular arrangement exists in two putative GH9 genes from the draft sequence of the Clostridium thermocellum genome. The three homologous CBM3b modules from A. cellulolyticus and C. thermocellum were overexpressed, but, surprisingly, none bound cellulosic substrates. The results raise fundamental questions concerning the possible role(s) of the newly described CBMs. Phylogenetic analysis and preliminary site-directed mutagenesis studies suggest that the catalytic module and the CBM3 dyad are distinctive in their sequences and are proposed to constitute a new GH9 architectural theme.

Amino Acid Sequence↗

Discovery and profiling of bovine microRNAs from immune-related and embryonic tissues.

MicroRNAs are small approximately 22 nucleotide-long noncoding RNAs capable of controlling gene expression by inhibiting translation. Alignment of human microRNA stem-loop sequences (mir) against a recent draft sequence assembly of the bovine genome resulted in identification of 334 predicted bovine mir. We sequenced five tissue-specific cDNA libraries derived from the small RNA fractions of bovine embryo, thymus, small intestine, and lymph node to validate these predictions and identify new mir. This strategy combined with comparative sequence analysis identified 129 sequences that corresponded to mature microRNAs (miR). A total of 107 sequences aligned to known human mir, and 100 of these matched expressed miR. The other seven sequences represented novel miR expressed from the complementary strand of previously characterized human mir. The 22 sequences without matches displayed characteristic mir secondary structures when folded in silico, and 10 of these retained sequence conservation with other vertebrate species. Expression analysis based on sequence identity counts revealed that some miR were preferentially expressed in certain tissues, while bta-miR-26a and bta-miR-103 were prevalent in all tissues examined. These results support the premise that species differences in regulation of gene expression by miR occur primarily at the level of expression and processing.

Animals↗