PubMed Health⌕ Search

Biomedical subjects

C Fields

Publications and source records attributed to C Fields.

At least 19 recordsLinked to original sources

The Genome Sequence DataBase version 1.0 (GSDB): from low pass sequences to complete genomes.

The Genome Sequence DataBase (GSDB) has completed its conversion to an improved relational database. The new database, GSDB 1.0, is fully operational and publicly available. Data contributions, including both original sequence submissions and community annotation, are being accomplished through the use of a graphical client-server interface tool, the GSDB Annotator, and via GIO (GSDB Input/Output) files. Data retrieval services are being provided through a new Web Query Tool and direct SQL. All methods of data contribution and data retrieval fully support the new data types that have been incorporated into GSDB, including discontiguous sequences, multiple sequence alignments, and community annotation.

Animals↗

Using case management systems to integrate clinical and financial data.

Integrated delivery systems can use automated case management information systems to better manage relevant clinical and financial data across the continuum of care. Effective case management systems can help caregivers track clinical and financial information; match appropriate resources to patient needs; and analyze populations to identify risk, enhance adherence to clinical guidelines, and understand provider treatment profiles.

Case Management↗

The Genome Sequence DataBase (GSDB): meeting the challenge of genomic sequencing.

The genome sequence database (GSDB) is a complete, publicly available relational database of DNA sequences and annotation maintained by the National Center for Genome Resources (NCGR) under a Cooperative Agreement with the US Department of Energy (DOE). GSDB provides direct, client- server access to the database for data contributions, community annotation and SQL queries. The GSDB Annotator, a multi-platform graphic user interface, is freely available. Automatically updated relational replicates of GSDB are also freely available.

Amino Acid Sequence↗

Informatics for ubiquitous sequencing.

Methods and technologies currently being developed promise an increase of between one and two orders of magnitude in the practical throughputs of DNA sequencing for gene discovery, expression analysis and variant analysis. Integrated laboratories will use all of these methods as components of a molecular strategy for the functional characterization of genes and their products. This review summarizes the types of data produced by these strategies, and the analysis and management challenges that they raise.

Biotechnology↗

Intraspecific variation in small-subunit rRNA sequences in GenBank: why single sequences may not adequately represent prokaryotic taxa.

Small-subunit rRNA (SSU rRNA) sequencing is a powerful tool to detect, identify, and classify prokaryotic organisms, and there is currently an explosion of SSU rRNA sequencing in the microbiology community. We report unexpectedly high levels of intraspecific variation (within and between strains) of prokaryote SSU rRNA sequences deposited in GenBank. A total of 82% of the prokaryote species with two published SSU rRNA sequences had more variable positions than a 0.1% random sequencing error would predict, and 48% of these sequence pairs had more variable positions than predicted by a 1.0% random sequencing error. Other sources of sequence variability must account for some of this intraspecific variation. Given these results, phylogenetic studies and biodiversity estimates obtained by using prokaryotic SSU rRNA sequences cannot proceed under the assumption that rRNA sequences of single operons from single isolates adequately represent their taxa. Sequencing SSU rRNA molecules from multiple operons and multiple isolates is highly recommended to obtain meaningful phylogenetic hypotheses, as is careful attention to accurate strain identification.

Bacteria↗

Expressed sequence tags identify a human isolog of the suil translation initiation factor.

The complete cDNA sequence of a human isolog of the yeast suil translation initiation factor gene was obtained by assembling over 40 expressed sequence tags (ESTs) for this gene obtained from a variety of tissue-specific cDNA libraries. The human suilisol gene product is a 113 amino-acid polypeptide similar to proteins known from yeast, rice, mosquito, and Methanococcus. The identification of suilisol illustrates the utility of assemblies of independent ESTs for deriving full-length cDNA sequences for new human genes.

Amino Acid Sequence↗

Analysis of gene expression by tissue and developmental stage.

High-throughput sequencing of cDNAs from multiple tissue- and stage-specific libraries is an efficient method for characterizing gene expression by tissue and developmental stage. When combined with functional information derived from the systematic study of transcription factors, signal transducers, and other regulatory molecules in model systems, data from expressed sequence tag projects provide an increasingly detailed picture of gene expression and its regulation. Understanding this picture will require the development of highly sophisticated databases to organize and correlate these data.

Animals↗

A quality control algorithm for DNA sequencing projects.

Heterologous DNA sequences from rearrangements with the genomes of host cells, genomic fragments from hybrid cells, or impure tissue sources can threaten the purity of libraries that are derived from RNA or DNA. Hybridization methods can only detect contaminants from known or suspected heterologous sources, and whole library screening is technically very difficult. Detection of contaminating heterologous clones by sequence alignment is only possible when related sequences are present in a known database. We have developed a statistical test to identify heterologous sequences that is based on the differences in hexamer composition of DNA from different organisms. This test does not require that sequences similar to potential heterologous contaminants are present in the database, and can in principle detect contamination by previously unknown organisms. We have applied this test to the major public expressed sequence tag (EST) data sets to evaluate its utility as a quality control measure and a peer evaluation tool. There is detectable heterogeneity in most human and C.elegans EST data sets but it is not apparently associated with cross-species contamination. However, there is direct evidence for both yeast and bacterial sequence contamination in some public database sequences annotated as human. Results obtained with the hexamer test have been confirmed with similarity searches using sequences from the relevant data sets.

Algorithms↗

3,400 new expressed sequence tags identify diversity of transcripts in human brain.

We present the results of the partial sequencing of over 3,400 expressed sequence tags (ESTs) from human brain cDNA clones, which increases the number of distinct genes expressed in the brain, that are represented by ESTs, to about 6,000. By choosing clones in an unbiased manner, it is possible to construct a profile of the transcriptional activity of the brain at different stages. Proteins that comprise the cytoskeleton are the most abundant; however, a large variety of regulatory proteins are also seen. About half of the ESTs predicted to contain a protein-coding region have no matches in the public peptide databases and may represent new gene families.

Amino Acid Sequence↗

Rapid cDNA sequencing (expressed sequence tags) from a directionally cloned human infant brain cDNA library.

A human infant brain cDNA library, made specifically for production of expressed sequence tags (ESTs) was evaluated by partial sequencing of over 1,600 clones. Advantages of this library, constructed for EST sequencing, include the use of directional cloning, size selection, very low numbers of mitochondrial and ribosomal transcripts, short polyA tails, few non-recombinants and a broad representation of transcripts. 37% of the clones were identified, based on matches to over 320 different genes in the public databases. Of these, two proteins similar to the Alzheimer's disease amyloid precursor protein were identified.

Amino Acid Sequence↗

The use of deficiencies to determine essential gene content in the let-56-unc-22 region of Caenorhabditis elegans.

We have investigated the possibility of using the polymerase chain reaction to detect deletions of coding elements in the unc-22-let-56 interval on chromosome IV in the nematode Caenorhabditis elegans. Our analysis of approximately 13 kb of genomic sequence immediately to the left of the unc-22 gene resulted in the identification of four possible genes. Partial cDNAs have been identified for three of them. To determine whether any of these coding elements are essential for development, we required a method for the induction and selection of mutations in these elements. Our approach was to identify a set of formaldehyde and gamma radiation induced unc-22 mutations that mapped to the unc-22-let-56 region, and then employ polymerase chain reaction methodology to identify deficiencies that affected one or more of the four identified coding elements. Two small deficiencies were identified in this manner. Characterization of these deficiencies shows that there are no coding elements between unc-22 and let-56 (the nearest mutationally identified gene to the left of unc-22), which are required in development under laboratory conditions. We conclude that the polymerase chain reaction is a practical tool for the detection of deletions of coding elements identified in this region, and that characterization of such deficiencies provides a method for assessing whether or not these elements are required for development.

Amino Acid Sequence↗

Splicing signals in Drosophila: intron size, information content, and consensus sequences.

A database of 209 Drosophila introns was extracted from Genbank (release number 64.0) and examined by a number of methods in order to characterize features that might serve as signals for messenger RNA splicing. A tight distribution of sizes was observed: while the smallest introns in the database are 51 nucleotides, more than half are less than 80 nucleotides in length, and most of these have lengths in the range of 59-67 nucleotides. Drosophila splice sites found in large and small introns differ in only minor ways from each other and from those found in vertebrate introns. However, larger introns have greater pyrimidine-richness in the region between 11 and 21 nucleotides upstream of 3' splice sites. The Drosophila branchpoint consensus matrix resembles C T A A T (in which branch formation occurs at the underlined A), and differs from the corresponding mammalian signal in the absence of G at the position immediately preceding the branchpoint. The distribution of occurrences of this sequence suggests a minimum distance between 5' splice sites and branchpoints of about 38 nucleotides, and a minimum distance between 3' splice sites and branchpoints of 15 nucleotides. The methods we have used detect no information in exon sequences other than in the few nucleotides immediately adjacent to the splice sites. However, Drosophila resembles many other species in that there is a discontinuity in A + T content between exons and introns, which are A + T rich.

Animals↗

Sequence identification of 2,375 human brain genes.

We recently described a new approach for the rapid characterization of expressed genes by partial DNA sequencing to generate 'expressed sequence tags'. From a set of 600 human brain complementary DNA clones, 348 were informative nuclear-encoded messenger RNAs. We have now partially sequenced 2,672 new, independent cDNA clones isolated from four human brain cDNA libraries to generate 2,375 expressed sequence tags to nuclear-encoded genes. These sequences, together with 348 brain expressed sequence tags from our previous study, comprise more than 2,500 new human genes and 870,769 base pairs of DNA sequence. These data represent an approximate doubling of the number of human genes identified by DNA sequencing and may represent as many as 5% of the genes in the human genome.

Brain Chemistry↗

Information contents and dinucleotide compositions of plant intron sequences vary with evolutionary origin.

The DNA sequence composition of 526 dicot and 345 monocot intron sequences have been characterized using computational methods. Splice site information content and bulk intron and exon dinucleotide composition were determined. Positions 4 and 5 of 5' splice sites contain different statistically significant levels of information in the two groups. Basal levels of information in introns are higher in dicots than in monocots. Two dinucleotide groups, WW (AA, AU, UA, UU) and SS (CC, CG, GC, GG) have significantly different frequencies in exons and introns of the two plant groups. These results suggest that the mechanisms of splice-site recognition and binding may differ between dicot and monocot plants.

Base Composition↗

Caenorhabditis elegans expressed sequence tags identify gene families and potential disease gene homologues.

A database containing mapped partial cDNA sequences from Caenorhabditis elegans will provide a ready starting point for identifying nematode homologues of important human genes and determining their functions in C. elegans. A total of 720 expressed sequence tags (ESTs) have been generated from 585 clones randomly selected from a mixed-stage C. elegans cDNA library. Comparison of these ESTs with sequence databases identified 422 new C. elegans genes, of which 317 are not similar to any sequences in the database. Twenty-six new genes have been mapped by YAC clone hybridization. Members of several gene families, including cuticle collagens, GTP-binding proteins, and RNA helicases were discovered. Many of the new genes are similar to known or potential human disease genes, including CFTR and the LDL receptor.

Amino Acid Sequence↗