We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CR

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Cryptography and Security

Title: MALIGN: Explainable Static Raw-byte Based Malware Family Classification using Sequence Alignment

Abstract: For a long time, malware classification and analysis have been an arms-race between antivirus systems and malware authors. Though static analysis is vulnerable to evasion techniques, it is still popular as the first line of defense in antivirus systems. But most of the static analyzers failed to gain the trust of practitioners due to their black-box nature. We propose MAlign, a novel static malware family classification approach inspired by genome sequence alignment that can not only classify malware families but can also provide explanations for its decision. MAlign encodes raw bytes using nucleotides and adopts genome sequence alignment approaches to create a signature of a malware family based on the conserved code segments in that family, without any human labor or expertise. We evaluate MAlign on two malware datasets, and it outperforms other state-of-the-art machine learning based malware classifiers (by 4.49% - 0.07%), especially on small datasets (by 19.48% - 1.2%). Furthermore, we explain the generated signatures by MAlign on different malware families illustrating the kinds of insights it can provide to analysts, and show its efficacy as an analysis tool. Additionally, we evaluate its theoretical and empirical robustness against some common attacks. In this paper, we approach static malware analysis from a unique perspective, aiming to strike a delicate balance among performance, interpretability, and robustness.
Subjects: Cryptography and Security (cs.CR)
DOI: 10.1016/j.cose.2024.103714
Cite as: arXiv:2111.14185 [cs.CR]
  (or arXiv:2111.14185v3 [cs.CR] for this version)

Submission history

From: Shoumik Saha [view email]
[v1] Sun, 28 Nov 2021 15:57:28 GMT (842kb,D)
[v2] Sun, 20 Aug 2023 13:25:24 GMT (2176kb,D)
[v3] Fri, 12 Jan 2024 21:14:41 GMT (1788kb,D)

Link back to: arXiv, form interface, contact.