PubMed Health⌕ Search

Biomedical subjects

K H Fasman

Publications and source records attributed to K H Fasman.

9 recordsLinked to original sources

A decision tree system for finding genes in DNA.

MORGAN is an integrated system for finding genes in vertebrate DNA sequences. MORGAN uses a variety of techniques to accomplish this task, the most distinctive of which is a decision tree classifier. The decision tree system is combined with new methods for identifying start codons, donor sites, and acceptor sites, and these are brought together in a frame-sensitive dynamic programming algorithm that finds the optimal segmentation of a DNA sequence into coding and noncoding regions (exons and introns). The optimal segmentation is dependent on a separate scoring function that takes a subsequence and assigns to it a score reflecting the probability that the sequence is an exon. The scoring functions in MORGAN are sets of decision trees that are combined to give a probability estimate. Experimental results on a database of 570 vertebrate DNA sequences show that MORGAN has excellent performance by many different measures. On a separate test set, it achieves an overall accuracy of 95 %, with a correlation coefficient of 0.78, and a sensitivity and specificity for coding bases of 83 % and 79%. In addition, MORGAN identifies 58% of coding exons exactly; i.e., both the beginning and end of the coding regions are predicted correctly. This paper describes the MORGAN system, including its decision tree routines and the algorithms for site recognition, and its performance on a benchmark database of vertebrate DNA.

Algorithms↗

The GDB Human Genome Database Anno 1997.

The value of the Genome Database (GDB) for the human genome research community has been greatly increased since the release of version 6. 0 last year. Thanks to the introduction of significant technical improvements, GDB has seen dramatic growth in the type and volume of information stored in the database. This article summarizes the types of data that are now available in the Genome Database, demonstrates how the database is interconnected with other biomedical resources on the World Wide Web, discusses how researchers can contribute new or updated information to the database, and describes our current efforts as well as planned improvements for the future.

Base Sequence↗

Finding genes in DNA with a Hidden Markov Model.

This study describes a new Hidden Markov Model (HMM) system for segmenting uncharacterized genomic DNA sequences into exons, introns, and intergenic regions. Separate HMM modules were designed and trained for specific regions of DNA: exons, introns, intergenic regions, and splice sites. The models were then tied together to form a biologically feasible topology. The integrated HMM was trained further on a set of eukaryotic DNA sequences and tested by using it to segment a separate set of sequences. The resulting HMM system which is called VEIL (Viterbi Exon-Intron Locator), obtains an overall accuracy on test data of 92% of total bases correctly labelled, with a correlation coefficient of 0.73. Using the more stringent test of exact exon prediction, VEIL correctly located both ends of 53% of the coding exons, and 49% of the exons it predicts are exactly correct. These results compare favorably to the best previous results for gene structure prediction and demonstrate the benefits of using HMMs for this problem.

Algorithms↗

Improvements to the GDB Human Genome Data Base.

Version 6.0 of the Human Genome Data Base introduces a number of significant improvements over previous releases of GDB. The most important of these are revised data representations for genes and genomic maps and a new curatorial model for the database. GDB 6.0 is the first major genomic database to provide read/write access directly to the scientific community, including capabilities for third-party annotation. The revised database can represent all major categories of genetic and physical maps, along with the underlying order and distance information used to construct them. The improved representation permits more sophisticated map queries to be posed and supports the graphical display of maps. In addition the new GDB has a richer model for gene information, better suited for supporting cross-references to databases describing gene function, structure, products, expression and associated phenotypes.

Animals↗

Restructuring the genome data base: a model for a federation of biological databases.

The creation of a federation of public biological databases has been proposed. Formerly independent systems will need to be modified to interoperate better within this federation. This will enable the federated system to provide biologists with an integrated view of biological data. The GDB Human Genome Data Base is being restructured to participate in the proposed federation. GDB itself will be organized into a collection of related data sets in support of human gene mapping. The techniques that will be used to link these data sets will be applicable to the federation as a whole. Links will be based on stable accession numbers that have no inherent information content and are guaranteed always to be recognized. Improvements will be made in the links between GDB and the nucleotide sequence databases to test this approach further.

Amino Acid Sequence↗

The GDB Human Genome Data Base anno 1994.

In 1991 the Genome Data Base at Johns Hopkins University School of Medicine was selected as the central repository for mapping data from the Human Genome Project, and was funded by NIH and DOE under a three year award. GDB has now finished 28 months of Federally funded operation. During this period a great deal of progress and many internal changes have taken place. In addition, many changes have also occurred in the external environment, and GDB has adapted its strategies to play an appropriate role in those changes as well. Recognizing the central role of mapping information in the genome project, it is important that GDB respond aggressively to the increasing demands of genomic researchers, as well as formulate a program of response to a number of long standing, but still unmet, needs of that community. It is even more important that GDB provide leadership in the genome informatics enterprise. Three themes described here are dominant in our future plans and represent the essence of the major changes made in the past year. They include: enhanced data acquisition, better map representation, and full integration into the collection of genomic databases.

Computer Communication Networks↗

The GDB human genome data base anno 1993.

Version 5.0 of the Genome Data Base (GDB) was released in March 1993. This document describes some of the significant changes to the types of data which are stored within the GDB. In addition to handling a wider scope of data, the GDB 5.0 application software now supports the X-Windows protocol. Although the GDB software still remains the most widely utilized method for accessing the data, alternate methods of access are now available, including direct SQL (Structured Query Language) queries, FTP (Internet File Transfer Protocol), WAIS (Wide Area Information Server), and other tools produced by third-party developers.

Chromosome Mapping↗