Learning Fast and Slow: PROPEDEUTICA for Real-time Malware Detection

Sun, Ruimin; Yuan, Xiaoyong; He, Pan; Zhu, Qile; Chen, Aokun; Gregio, Andre; Oliveira, Daniela; Li, Xiaolin

Full-text links:

Download:

Current browse context:

cs.CR

< prev | next >

new | recent | 1712

Computer Science > Cryptography and Security

Title: Learning Fast and Slow: PROPEDEUTICA for Real-time Malware Detection

Authors: Ruimin Sun, Xiaoyong Yuan, Pan He, Qile Zhu, Aokun Chen, Andre Gregio, Daniela Oliveira, Xiaolin Li

(Submitted on 4 Dec 2017 (this version), latest version 17 Oct 2021 (v2))

Abstract: In this paper, we introduce and evaluate PROPEDEUTICA, a novel methodology and framework for efficient and effective real-time malware detection, leveraging the best of conventional machine learning (ML) and deep learning (DL) algorithms. In PROPEDEUTICA, all software processes in the system start execution subjected to a conventional ML detector for fast classification. If a piece of software receives a borderline classification, it is subjected to further analysis via more performance expensive and more accurate DL methods, via our newly proposed DL algorithm DEEPMALWARE. Further, we introduce delays to the execution of software subjected to deep learning analysis as a way to "buy time" for DL analysis and to rate-limit the impact of possible malware in the system. We evaluated PROPEDEUTICA with a set of 9,115 malware samples and 877 commonly used benign software samples from various categories for the Windows OS. Our results show that the false positive rate for conventional ML methods can reach 20%, and for modern DL methods it is usually below 6%. However, the classification time for DL can be 100X longer than conventional ML methods. PROPEDEUTICA improved the detection F1-score from 77.54% (conventional ML method) to 90.25%, and reduced the detection time by 54.86%. Further, the percentage of software subjected to DL analysis was approximately 40% on average. Further, the application of delays in software subjected to ML reduced the detection time by approximately 10%. Finally, we found and discussed a discrepancy between the detection accuracy offline (analysis after all traces are collected) and on-the-fly (analysis in tandem with trace collection). Our insights show that conventional ML and modern DL-based malware detectors in isolation cannot meet the needs of efficient and effective malware detection: high accuracy, low false positive rate, and short classification time.

Comments:	17 pages, 7 figures
Subjects:	Cryptography and Security (cs.CR); Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1712.01145 [cs.CR]
	(or arXiv:1712.01145v1 [cs.CR] for this version)

Submission history

From: Xiaoyong Yuan [view email]
[v1] Mon, 4 Dec 2017 15:30:03 GMT (2267kb,D)
[v2] Sun, 17 Oct 2021 16:33:29 GMT (8472kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1712.01145v1

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Cryptography and Security

Title: Learning Fast and Slow: PROPEDEUTICA for Real-time Malware Detection

Submission history