Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification

Jorgensen, Steven; Holodnak, John; Dempsey, Jensen; de Souza, Karla; Raghunath, Ananditha; Rivet, Vernon; DeMoes, Noah; Alejos, Andrés; Wollaber, Allan

doi:10.1109/TAI.2023.3244168

Full-text links:

Download:

Current browse context:

cs.CR

< prev | next >

new | recent | 2205

Computer Science > Cryptography and Security

Title: Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification

Authors: Steven Jorgensen, John Holodnak, Jensen Dempsey, Karla de Souza, Ananditha Raghunath, Vernon Rivet, Noah DeMoes, Andrés Alejos, Allan Wollaber (MIT Lincoln Laboratory)

(Submitted on 11 May 2022 (v1), last revised 6 Oct 2023 (this version, v3))

Abstract: With the increasing prevalence of encrypted network traffic, cyber security analysts have been turning to machine learning (ML) techniques to elucidate the traffic on their networks. However, ML models can become stale as new traffic emerges that is outside of the distribution of the training set. In order to reliably adapt in this dynamic environment, ML models must additionally provide contextualized uncertainty quantification to their predictions, which has received little attention in the cyber security domain. Uncertainty quantification is necessary both to signal when the model is uncertain about which class to choose in its label assignment and when the traffic is not likely to belong to any pre-trained classes.
We present a new, public dataset of network traffic that includes labeled, Virtual Private Network (VPN)-encrypted network traffic generated by 10 applications and corresponding to 5 application categories. We also present an ML framework that is designed to rapidly train with modest data requirements and provide both calibrated, predictive probabilities as well as an interpretable "out-of-distribution" (OOD) score to flag novel traffic samples. We describe calibrating OOD scores using p-values of the relative Mahalanobis distance.
We demonstrate that our framework achieves an F1 score of 0.98 on our dataset and that it can extend to an enterprise network by testing the model: (1) on data from similar applications, (2) on dissimilar application traffic from an existing category, and (3) on application traffic from a new category. The model correctly flags uncertain traffic and, upon retraining, accurately incorporates the new data.

Comments:	Paper is 15 pages and has 10 figures. Published in IEEE Transactions on Artificial Intelligence (this https URL). For associated dataset, see this https URL
Subjects:	Cryptography and Security (cs.CR); Machine Learning (cs.LG)
MSC classes:	68T07 (Primary), 68M25 (Secondary)
ACM classes:	I.2.6; K.6.5
DOI:	10.1109/TAI.2023.3244168
Cite as:	arXiv:2205.05628 [cs.CR]
	(or arXiv:2205.05628v3 [cs.CR] for this version)

Submission history

From: Steven Jorgensen [view email]
[v1] Wed, 11 May 2022 16:54:37 GMT (807kb,D)
[v2] Wed, 8 Mar 2023 18:45:17 GMT (1841kb,D)
[v3] Fri, 6 Oct 2023 20:11:03 GMT (919kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2205.05628

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Cryptography and Security

Title: Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification

Submission history