We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:


Current browse context:


Change to browse by:

References & Citations


(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Quantitative Biology > Genomics

Title: Genome Compression Against a Reference

Abstract: Being able to store and transmit human genome sequences is an important part in genomic research and industrial applications. The complete human genome has 3.1 billion base pairs (haploid), and storing the entire genome naively takes about 3 GB, which is infeasible for large scale usage.
However, human genomes are highly redundant. Any given individual's genome would differ from another individual's genome by less than 1%. There are tools like DNAZip, which express a given genome sequence by only noting down the differences between the given sequence and a reference genome sequence. This allows losslessly compressing the given genome to ~ 4 MB in size.
In this work, we demonstrate additional improvements on top of the DNAZip library, where we show an additional ~ 11% compression on top of DNAZip's already impressive results. This would allow further savings in disk space and network costs for transmitting human genome sequences.
Subjects: Genomics (q-bio.GN)
Cite as: arXiv:2010.02286 [q-bio.GN]
  (or arXiv:2010.02286v1 [q-bio.GN] for this version)

Submission history

From: Gaurav Menghani [view email]
[v1] Mon, 5 Oct 2020 19:00:23 GMT (12kb,D)

Link back to: arXiv, form interface, contact.