We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.IT

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Information Theory

Title: Optimal alphabet for single text compression

Abstract: A text can be viewed via different representations, i.e. as a sequence of letters, n-grams of letters, syllables, words, and phrases. Here we study the optimal noiseless compression of texts using the Huffman code, where the alphabet of encoding coincides with one of those representations. We show that it is necessary to account for the codebook when compressing a single text. Hence, the total compression comprises of the optimally compressed text -- characterized by the entropy of the alphabet elements -- and the codebook which is text-specific and therefore has to be included for noiseless (de)compression. For texts of Project Gutenberg the best compression is provided by syllables, i.e. the minimal meaning-expressing element of the language. If only sufficiently short texts are retained, the optimal alphabet is that of letters or 2-grams of letters depending on the retained length.
Comments: 11 pages, 12 figures, 1 table
Subjects: Information Theory (cs.IT); Computation and Language (cs.CL); Data Analysis, Statistics and Probability (physics.data-an)
Cite as: arXiv:2201.05234 [cs.IT]
  (or arXiv:2201.05234v1 [cs.IT] for this version)

Submission history

From: Armen Allahverdyan [view email]
[v1] Thu, 13 Jan 2022 22:16:51 GMT (1293kb,D)

Link back to: arXiv, form interface, contact.