PubMed · 1127034
Data compression of large document data bases.
Abstract
Consideration is given to a document data base that is structured for information retrieval purposes by means of an inverted index and term dictionary. Vocabulary characteristics of various fields are described, and it is shown how the data base may be stored in a compressed form by use of restricted variable length codes that produce a compression not greatly in excess of the optimum that could be achieved through use of Huffman codes. The coding is word oriented. An alternative scheme of word fragment coding is described. It has the advantage that it allows the use of a small dictionary, but is less efficient with respect to compression of the data base.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
H S Heaps. 1975. Data compression of large document data bases.. https://doi.org/10.1021/ci60001a011
Cite the original work for its findings. Save a collection to share your selection of sources.