PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “WGS sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

A preprocessor for shotgun assembly of large genomes.

The whole-genome shotgun (WGS) assembly technique has been remarkably successful in efforts to determine the sequence of bases that make up a genome. WGS assembly begins with a large collection of short fragments that have been selected at random from a genome. The sequence of bases at each end of the fragment is determined, albeit imprecisely, resulting in a sequence of letters called a "read." Each letter in a read is assigned a quality value, which estimates the probability that a sequencing error occurred in determining that letter. Reads are typically cut off after about 500 letters, where sequencing errors become endemic. We report on a set of procedures that (1) corrects most of the sequencing errors, (2) changes quality values accordingly, and (3) produces a list of "overlaps," i.e., pairs of reads that plausibly come from overlapping parts of the genome. Our procedures, which we call collectively the "UMD Overlapper," can be run iteratively and as a preprocessor for other assemblers. We tested the UMD Overlapper on Celera's Drosophila reads. When we replaced Celera's overlap procedures in the front end of their assembler, it was able to produce a significantly improved genome.

Animals↗

On the sequencing of the human genome.

Two recent papers using different approaches reported draft sequences of the human genome. The international Human Genome Project (HGP) used the hierarchical shotgun approach, whereas Celera Genomics adopted the whole-genome shotgun (WGS) approach. Here, we analyze whether the latter paper provides a meaningful test of the WGS approach on a mammalian genome. In the Celera paper, the authors did not analyze their own WGS data. Instead, they decomposed the HGP's assembled sequence into a "perfect tiling path", combined it with their WGS data, and assembled the merged data set. To study the implications of this approach, we perform computational analysis and find that a perfect tiling path with 2-fold coverage is sufficient to recover virtually the entirety of a genome assembly. We also examine the manner in which the assembly was anchored to the human genome and conclude that the process primarily depended on the HGP's sequence-tagged site maps, BAC maps, and clone-based sequences. Our analysis indicates that the Celera paper provides neither a meaningful test of the WGS approach nor an independent sequence of the human genome. Our analysis does not imply that a WGS approach could not be successfully applied to assemble a draft sequence of a large mammalian genome, but merely that the Celera paper does not provide such evidence.

Chromosomes, Artificial, Bacterial↗

Water-gas-shift reaction on molybdenum carbide surfaces: essential role of the oxycarbide.

Density functional theory (DFT) was employed to investigate the behavior of Mo carbides in the water-gas-shift reaction (WGS, CO + H(2)O --> H(2) +CO(2)). The kinetics of the WGS reaction was studied on the surfaces of Mo-terminated Mo(2)C(001) (Mo-Mo(2)C), C-terminated Mo(2)C(001) (C-Mo(2)C), and Cu(111) as a known active catalyst. Our results show that the WGS activity decreases in a sequence: Cu > C-Mo(2)C > Mo-Mo(2)C. The slow kinetics on C-Mo(2)C and Mo-Mo(2)C is due to the fact that the C or Mo sites bond oxygen too strongly to allow the facile removal of this species. In fact, due to the strong O-Mo and O-C interactions, the carbide surfaces are likely to be covered by O produced from the H(2)O dissociation. It is shown that the O-covered Mo-terminated Mo(2)C(001) (O_Mo-Mo(2)C) surface displays the lowest WGS activity of all. With the Mo oxide in the surface, O_Mo-Mo(2)C is too inert to adsorb CO or to dissociate H(2)O. In contrast, the same amount of O on the C-Mo(2)C surface (O_C-Mo(2)C) does not lead to deactivation, but enhances the rate of the WGS reaction and makes this system even more active than Cu. The good behavior of O_C-Mo(2)C is attributed to the formation of a Mo oxycarbide in the surface. The C atoms destabilize O-poisoning by forming CO species, which shift away from the Mo hollow sites when the surface reacts with other adsorbates. In this way, the Mo sites are able to provide a moderate bond to the reaction intermediates. In addition, both C and O atoms are not spectators and directly participate in the WGS reaction.

Journal Article↗

The phusion assembler.

The Phusion assembler has assembled the mouse genome from the whole-genome shotgun (WGS) dataset collected by the Mouse Genome Sequencing Consortium, at ~7.5x sequence coverage, producing a high-quality draft assembly 2.6 gigabases in size, of which 90% of these bases are in 479 scaffolds. For the mouse genome, which is a large and repeat-rich genome, the input dataset was designed to include a high proportion of paired end sequences of various size selected inserts, from 2-200 kbp lengths, into various host vector templates. Phusion uses sequence data, called reads, and information about reads that share common templates, called read pairs, to drive the assembly of this large genome to highly accurate results. The preassembly stage, which clusters the reads into sensible groups, is a key element of the entire assembler, because it permits a simple approach to parallelization of the assembly stage, as each cluster can be treated independent of the others. In addition to the application of Phusion to the mouse genome, we will also present results from the WGS assembly of Caenorhabditis briggsae sequenced to about 11x coverage. The C. briggsae assembly was accessioned through EMBL, http://www.ebi.ac.uk/services/index.html, using the series CAAC01000001-CAAC01000578, however, the Phusion mouse assembly described here was not accessioned. The mouse data was generated by the Mouse Genome Sequencing Consortium. The C. briggsae sequence was generated at The Wellcome Trust Sanger Institute and the Genome Sequencing Center, Washington University School of Medicine.

Animals↗

Integrative Long-Read Multi-Omics of a Patient With GPI Deficiency: A Molecular Case Study of a Candidate Dual-Effect GPI Variant.

The molecular determinants of phenotypic severity in red cell enzymopathies are often obscured by the disconnect between coding sequence variants and their regulatory landscapes. Here we present a single-patient molecular case study that uses an integrative multi-omic approach-combining short-read WGS, PacBio HiFi long-read sequencing, native CpG methylation profiling, and Iso-Seq full-length transcriptomics-to characterize a severe, transfusion-dependent hemolytic anaemia. We identified a compound heterozygous state in the glucose-6-phosphate isomerase (GPI) gene, with no wild-type allele present. One allele (Haplotype 1) carried a missense variant (p.His191Arg); the other (Haplotype 2) carried a distinct missense variant, c.1414C>T (p.Arg472Cys), previously reported as biochemically unstable. Long-read phasing placed the two variants in trans. Allele-resolved transcript counts showed a directionally consistent but statistically non-significant trend toward higher expression of Haplotype 2 across two Iso-Seq replicates. Notably, the c.1414C>T transition abolishes a local CpG dinucleotide; in a small number of haplotype-2 reads spanning this position, the corresponding cytosine on the wild-type/Haplotype-1 background was methylated. We did not measure GPI protein abundance, enzymatic activity, or stability in this patient, and we do not establish that methylation at this site regulates GPI transcription. On the basis of these correlative observations in a single patient, we propose-as a hypothesis for future testing-that a coding variant might simultaneously perturb protein stability and disrupt a local epigenetic mark, and we outline the experiments required to test whether such a dual effect contributes to disease. This case illustrates the value of integrative long-read multi-omics for generating mechanistic hypotheses about variants of uncertain significance, while underscoring that causal claims require dedicated functional validation.

Humans↗

Differential lineage-specific amplification of transposable elements is responsible for genome size variation in Gossypium.

The DNA content of eukaryotic nuclei (C-value) varies approximately 200,000-fold, but there is only a approximately 20-fold variation in the number of protein-coding genes. Hence, most C-value variation is ascribed to the repetitive fraction, although little is known about the evolutionary dynamics of the specific components that lead to genome size variation. To understand the modes and mechanisms that underlie variation in genome composition, we generated sequence data from whole genome shotgun (WGS) libraries for three representative diploid (n = 13) members of Gossypium that vary in genome size from 880 to 2460 Mb (1C) and from a phylogenetic outgroup, Gossypioides kirkii, with an estimated genome size of 588 Mb. Copy number estimates including all dispersed repetitive sequences indicate that 40%-65% of each genome is composed of transposable elements. Inspection of individual sequence types revealed differential, lineage-specific expansion of various families of transposable elements among the different plant lineages. Copia-like retrotransposable element sequences have differentially accumulated in the Gossypium species with the smallest genome, G. raimondii, while gypsy-like sequences have proliferated in the lineages with larger genomes. Phylogenetic analyses demonstrated a pattern of lineage-specific amplification of particular subfamilies of retrotransposons within each species studied. One particular group of gypsy-like retrotransposon sequences, Gorge3 (Gossypium retrotransposable gypsy-like element), appears to have undergone a massive proliferation in two plant lineages, accounting for a major fraction of genome-size change. Like maize, Gossypium has undergone a threefold increase in genome size due to the accumulation of LTR retrotransposons over the 5-10 Myr since its origin.

Base Sequence↗

Two isolectins from leaves of winged bean, Psophocarpus tetragonolobus (L.) DC.

Two isolectins were isolated from leaves of winged bean and characterized. They differed from each other in terms of their immunological properties, hemagglutinating activities, sugar inhibition patterns, and amino acid compositions. Both lectins were acidic and one of them (L-I), which was inactive toward trypsinized human type O erythrocytes, was similar to one of green shell lectins (WGS-1); which resembled basic seed lectin in its immunological properties. The amino-terminal sequence of L-I was homologous to that of WGS-1. The amino acid composition of L-I was similar to that of basic seed lectin, but the extent of the homology between amino-terminal sequences was low when L-I and basic seed lectin were compared. Examination by ELISA revealed that L-I and WGS-1 were distinct from the basic lectins of seeds and tuberous roots. L-I had a disulfide bridge between two subunits and it exhibited high hemagglutinating activity toward human type A erythrocytes, as compared to its activity toward other erythrocytes. By contrast, the properties of a second acidic lectin from winged bean leaves (L-II) were very similar to those of acidic lectins from seeds and tuberous roots, and the similarities extended as far as the immunological properties.

Amino Acid Sequence↗

Armenian Hamsters (Nothocricetulus migratorius): A New Host Susceptible to Corynebacterium bovis Infection and Disease.

Corynebacterium bovis causes skin disease in immunocompromised mice and possibly rats. In 2022, scaly skin and mortality were observed in 7- to 11-d-old neonates (n = 8) from a primiparous Armenian (Nothocricetulus migratorius) hamster breeding pair in a newly established colony. C. bovis was detected by culture and PCR, and affected animals had moderate to severe acanthotic, hyperkeratotic lesions with intralesional C. bovis confirmed by in situ hybridization. Intrafollicular Demodex cricetuli mites, an ectoparasite found in all laboratory-maintained Armenian hamsters, were also identified in affected animals. To elucidate the role of D. cricetuli on C. bovis-associated disease and maintain adult hamsters without the need for sustained mite treatment, a D. cricetuli-free colony was generated by treating breeding pairs and their 1- to 3-d-old neonates with topical fluralaner (35 mg/kg), and a prospective study was undertaken to compare C. bovis-associated pup mortality in D. cricetuli-free and D. cricetuli-infested hamsters. During the ensuing 22 mo, 4 of 96 (4.2%) litters born exhibited C. bovis-associated disease and/or mortality. The litters were born to 4 different nulliparous breeding pairs (n = 47, 9%). Of the 4 affected litters, 2 were D. cricetuli-infested while 2 were D. cricetuli-free. C. bovis was routinely cultured with a variable bacterial burden that had no association with mortality or skin lesion severity from all hamsters, independent of their D. cricetuli status. The severity of histologic pathology appeared to correlate with clinical presentation and mortality in neonates. Whole genome sequencing was performed on 4 hamster C. bovis isolates, which revealed a close genetic association among the isolates as well as with previously characterized mouse and rat C. bovis isolates.

DSS, deep skin scrape↗

The Drosophila melanogaster genome.

Drosophila's importance as a model organism made it an obvious choice to be among the first genomes sequenced, and the Release 1 sequence of the euchromatic portion of the genome was published in March 2000. This accomplishment demonstrated that a whole genome shotgun (WGS) strategy could produce a reliable metazoan genome sequence. Despite the attention to sequencing methods, the nucleotide sequence is just the starting point for genome-wide analyses; at a minimum, the genome sequence must be interpreted using expressed sequence tag (EST) and complementary DNA (cDNA) evidence and computational tools to identify genes and predict the structures of their RNA and protein products. The functions of these products and the manner in which their expression and activities are controlled must then be assessed-a much more challenging task with no clear endpoint that requires a wide variety of experimental and computational methods. We first review the current state of the Drosophila melanogaster genome sequence and its structural annotation and then briefly summarize some promising approaches that are being taken to achieve an initial functional annotation.

Animals↗

Characteristics of Tuberculosis Tests Performed during Postimport Quarantine of Nonhuman Primates, United States, 2021 to 2024.

Screening nonhuman primates (NHPs) for tuberculosis (TB) is important to protect the health of NHP colonies and people who interact with them. Screening is especially important for imported NHPs from countries where TB is prevalent and biosecurity practices may be lax. There are a variety of testing methods available for TB screening and diagnosis in NHPs; all have limitations, and their performance in different settings is incompletely characterized. The US Centers for Disease Control and Prevention (CDC) collects TB testing results as part of its regulatory oversight of NHP importation. We collated the results of tuberculin skin tests (TSTs), interferon-γ release assays (IGRAs), multiplexed fluorometric immunoassay (MFIA), Mycobacterium tuberculosis complex PCR, staining for acid-fast bacilli (AFB), and culture of bacteria from tissues for imported NHPs in CDC-mandated quarantine during fiscal years 2021 to 2024. We used these data to assess test performance and intertest agreement for the different tests used. Among 107 imported NHPs tested, TST and IGRA were the most common antemortem tests performed, but they agreed poorly with each other and with culture. AFB staining and PCR exhibited moderate agreement and high positive predictive values using culture as the gold standard. The most commonly affected tissues were lungs and tracheobronchial lymph nodes, regardless of the Mycobacterium sp. identified. Further research is needed to identify and validate additional methods for TB testing in NHPs, particularly for antemortem screening. Tissue acid-fast staining and PCR exhibited high positive predictive values and could be useful to inform policies and clinical decisions about colony management and occupational health while awaiting culture results.

AFB, acid-fast bacilli↗

Mouse SCO-spondin, a gene of the thrombospondin type 1 repeat (TSR) superfamily expressed in the brain.

SCO-spondin is specifically expressed in the subcommissural organ (SCO), a secretory ependymal differentiation lining the roof of the third ventricular cavity of the brain. When released into the cerebro-spinal fluid (CSF), SCO-spondin aggregates and forms Reissner's fiber (RF), a structure present in the central canal of the spinal cord. SCO-spondin belongs to the superfamily of proteins exhibiting conserved motifs called TSRs for 'thrombospondin type 1 repeats' and involved in axonal pathfinding during development. The mouse SCO-spondin coding sequence was searched by alignement of the coding bovine SCO-spondin sequence with the mouse whole genome shotgun (WGS) supercontig (NW 000250). Compared to the bovine, mouse SCO-spondin shows 66.8% identity of amino acids. This extracellular matrix glycoprotein has a modular arrangement of several conserved domains including 25 TSRs, 10 low-density lipoprotein receptor (LDLr) type A repeats and cystein-rich regions in the -NH2 and -COOH ends. The spatio-temporal expression of SCO-spondin was analyzed using specific antisera and an homospecific SCO-spondin riboprobe. In the adult, the patterns obtained by in situ hybridization (ISH) and immunohistochemistry correlated well in the SCO, while Reissner's fiber and the ampulla caudalis were immunoreactive only. In the fetus, both the immuno and ISH reactions appeared between 14 and 15 days post coïtum (dpc) in the SCO anlage. In addition, the mouse SCO-spondin gene was located at chromosome 6, between marker D6Mit352 and D6Mit119, in a conserved syntenic region.

Amino Acid Sequence↗

Chicken genome sequence: a centennial gift to poultry genetics.

A draft sequence of the chicken genome will be available by early 2004. This event conveniently marks the start of the second century of poultry genetics, coming 100 years after the use of the chicken to demonstrate Mendelian inheritance in animals by William Bateson. How will the second, post-genomic century of poultry genetics differ from the first? A whole genome shotgun (WGS) approach is being used to obtain the chicken sequence, with the goal of generating approximately six-fold coverage of the genome. Bacterial artificial chromosome (BAC) and fosmid clone end sequences, along with a BAC contig map integrated with genetic linkage and radiation hybrid maps, will form the platform for assembly of the WGS data. Rapid progress in global analysis of chicken gene expression patterns is also being made. Comparative genomics will link these new discoveries to the knowledge base for all other animal species. It's hoped that the genome sequence will also provide common ground on which to unite studies of the chicken as a model species with those aimed at agriculturally-relevant applications. The current status of chicken genomics will be assessed with projections for its near and long term future.

Animals↗

U6 snRNA variants isolated from the posterior silk gland of the silk moth Bombyx mori.

Five U6 small nuclear RNA (snRNA) isoforms were detected and characterized from the posterior silk gland (PSG) of the silk moth Bombyx mori (Nistari strain). Using the currently accepted U6 secondary structure model as a basis for comparison, the variants were analyzed for nucleotide differences across the sequence with a focus on known functional domains. Differences were observed primarily in single-stranded areas of which sixty percent were found in the highly conserved U4-U6 binding sites. In the Nistari strain, the U6A variant was found to be approximately four times more abundant as part of high molecular weight spliceosomal complexes when compared with U6A in the total unfractionated PSG cell lysate. Additionally, the European 703 B. mori strain total cell lysate U6 snRNA was analyzed and only the dominant U6A isoform initially identified in Nistari was found. Due to U6's essential role in pre-mRNA processing, variants may modulate assemblage of the catalytic core and in doing so potentially affect the rate of splicing. Phylogenetic analysis of the U6 snRNA sequences indicate an ancient divergence of U6 from the self-splicing group II intron module and a high degree of evolutionary conservation across species possibly due to functional constraints on the gene. Using in silico analysis, 35 full-length U6 variants were observed in the recently released Whole Genome Shotgun (WGS) database of the p50T strain. The consensus sequence of these U6 genes from p50T is identical to U6A identified in the Nistari strain. Furthermore p50T variant 1, which is represented in 14 genes, is equivalent to Nistari U6A.

Animals↗

A comparative analysis of HGSC and Celera human genome assemblies and gene sets.

MOTIVATION: Since the simultaneous publication of the human genome assembly by the International Human Genome Sequencing Consortium (HGSC) and Celera Genomics, several comparisons have been made of various aspects of these two assemblies. In this work, we set out to provide a more comprehensive comparative analysis of the two assemblies and their associated gene sets. RESULTS: The local sequence content for both draft genome assemblies has been similar since the early releases, however it took a year for the quality of the Celera assembly to approach that of HGSC, suggesting an advantage of HGSC's hierarchical shotgun (HS) sequencing strategy over Celera's whole genome shotgun (WGS) approach. While similar numbers of ab initio predicted genes can be derived from both assemblies, Celera's Otto approach consistently generated larger, more varied gene sets than the Ensembl gene build system. The presence of a non-overlapping gene set has persisted with successive data releases from both groups. Since most of the unique genes from either genome assembly could be mapped back to the other assembly, we conclude that the gene set discrepancies do not reflect differences in local sequence content but rather in the assemblies and especially the different gene-prediction methodologies.

Databases, Protein↗

The silk moth Bombyx mori U1 and U2 snRNA variants are differentially expressed.

Five U1 and eight U2 isoforms of the silk moth Bombyx mori exhibiting internal nucleotide differences have been previously identified and characterized in various tissues and developmental stages. In this investigation, it is demonstrated that the levels of some snRNA variants differ in egg and silk gland tissue and change during development. Qualitative and quantitative differences in the U1 and U2 variant populations were observed at three developmental points (early, middle and late) of the silk gland (SG) during the fifth instar larval stage of the silk moth. Statistical analyses of the various isoform populations across the fifth instar larval and egg stages show significant differences for some of the U1 and U2 variants. The representation of variant sequences in expressed U1 and U2 sequences (RT-PCR libraries) and in a whole-genome shotgun (WGS) assembly database was confirmed. In addition, conserved elements in the promoter 5'-flanking region of the U1 and U2 variants were identified in the WGS.

Animals↗

Dynamic building of a BAC clone tiling path for the Rat Genome Sequencing Project.

CLONEPICKER is a software pipeline that integrates sequence data with BAC clone fingerprints to dynamically select a minimal overlapping clone set covering the whole genome. In the Rat Genome Sequencing Project (RGSP), a hybrid strategy of "clone by clone" and "whole genome shotgun" approaches was used to maximize the merits of both approaches. Like the "clone by clone" method, one key challenge for this strategy was to select a low-redundancy clone set that covered the whole genome while the sequencing is in progress. The CLONEPICKER pipeline met this challenge using restriction enzyme fingerprint data, BAC end sequence data, and sequences generated from individual BAC clones as well as WGS reads. In the RGSP, an average of 7.5 clones was identified from each side of a seed clone, and the minimal overlapping clones were reliably selected. Combined with the assembled BAC fingerprint map, a set of BAC clones that covered >97% of the genome was identified and used in the RGSP.

Animals↗

DDBJ working on evaluation and classification of bacterial genes in INSDC.

DNA Data Bank of Japan (DDBJ) (http://www.ddbj.nig.ac.jp) newly collected and released 12,927,184 entries or 13,787,688,598 bases in the period from July 2005 to June 2006. The released data contain honeybee expressed sequence tags (ESTs), re-examined and re-annotated complete genome data of Escherichia coli K-12 W3110, medaka WGS and human MGA. We also systematically evaluated and classified the genes in the complete bacterial genomes submitted to the International Nucleotide Sequence Database Collaboration (INSDC, http://insdc.org) that is composed of DDBJ, EMBL Bank and GenBank. The examination and classification selected 557,000 genes as reliable ones among all the bacterial genes predicted by us.

Animals↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl/), maintained at the European Bioinformatics Institute (EBI), incorporates, organizes and distributes nucleotide sequences from public sources. The database is a part of an international collaboration with DDBJ (Japan) and GenBank (USA). Data are exchanged between the collaborating databases on a daily basis to achieve optimal synchrony. The web-based tool, Webin, is the preferred system for individual submission of nucleotide sequences, including Third Party Annotation (TPA) and alignment data. Automatic submission procedures are used for submission of data from large-scale genome sequencing centres and from the European Patent Office. Database releases are produced quarterly. The latest data collection can be accessed via FTP, email and WWW interfaces. The EBI's Sequence Retrieval System (SRS) integrates and links the main nucleotide and protein databases as well as many other specialist molecular biology databases. For sequence similarity searching, a variety of tools (e.g. FASTA and BLAST) are available that allow external users to compare their own sequences against the data in the EMBL Nucleotide Sequence Database, the complete genomic component subsection of the database, the WGS data sets and other databases. All available resources can be accessed via the EBI home page at http://www.ebi.ac.uk.

Animals↗