MalWhiteout: Reducing Label Errors in Android Malware Detection

Wang, L; Wang, H; Luo, X; Sui, Y

MalWhiteout: Reducing Label Errors in Android Malware Detection

Wang, L Wang, H Luo, X Sui, Y

Permalink

Publisher:: ACM
Publication Type:: Conference Proceeding
Citation:: ACM International Conference Proceeding Series, 2022, pp. 1-13
Issue Date:: 2022-09-19

Closed Access

	Filename	Description	Size
	3551349.3560418_published.pdf	Published version	1.04 MB	Adobe PDF	View/Open

Copyright Clearance Process

Recently Added
In Progress
Closed Access

This item is closed access and not available.

Full metadata record

Field	Value	Language
dc.contributor.author	Wang, L
dc.contributor.author	Wang, H
dc.contributor.author	Luo, X
dc.contributor.author	Sui, Y https://orcid.org/0000-0002-9510-6574
dc.date.accessioned	2023-03-24T06:22:27Z
dc.date.available	2023-03-24T06:22:27Z
dc.date.issued	2022-09-19
dc.identifier.citation	ACM International Conference Proceeding Series, 2022, pp. 1-13
dc.identifier.isbn	9781450396240
dc.identifier.uri	http://hdl.handle.net/10453/168344
dc.description.abstract	Machine learning based Android malware detection has attracted a great deal of research work in recent years. A reliable malware dataset is critical to evaluate the effectiveness of malware detection approaches. Unfortunately, existing malware datasets used in our community are mainly labelled by leveraging existing anti-virus services (i.e., VirusTotal), which are prone to mislabelling. This, however, would lead to the inaccurate evaluation of the malware detection techniques. Removing label noises from Android malware datasets can be quite challenging, especially at a large data scale. To address this problem, we propose an effective approach called MalWhiteout to reduce label errors in Android malware datasets. Specifically, we creatively introduce Confident Learning (CL), an advanced noise estimation approach, to the domain of Android malware detection. To combat false positives introduced by CL, we incorporate the idea of ensemble learning and inter-app relation to achieve a more robust capability in noise detection. We evaluate MalWhiteout on a curated large-scale and reliable benchmark dataset. Experimental results show that MalWhiteout is capable of detecting label noises with over 94% accuracy even at a high noise ratio (i.e., 30%) of the dataset. MalWhiteout outperforms the state-of-the-art approach in terms of both effectiveness (8% to 218% improvement) and efficiency (70 to 249 times faster) across different settings. By reducing label noises, we show that the performance of existing malware detection approaches can be improved.
dc.language	en
dc.publisher	ACM
dc.relation.ispartof	ACM International Conference Proceeding Series
dc.relation.ispartof	37th IEEE/ACM International Conference on Automated Software Engineering
dc.relation.isbasedon	10.1145/3551349.3560418
dc.rights	info:eu-repo/semantics/closedAccess
dc.title	MalWhiteout: Reducing Label Errors in Android Malware Detection
dc.type	Conference Proceeding
pubs.organisational-group	/University of Technology Sydney
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology
pubs.organisational-group	/University of Technology Sydney/Strength - AAII - Australian Artificial Intelligence Institute
utslib.copyright.status	closed_access	*
dc.date.updated	2023-03-24T06:22:26Z
pubs.publication-status	Published

Abstract:

Machine learning based Android malware detection has attracted a great deal of research work in recent years. A reliable malware dataset is critical to evaluate the effectiveness of malware detection approaches. Unfortunately, existing malware datasets used in our community are mainly labelled by leveraging existing anti-virus services (i.e., VirusTotal), which are prone to mislabelling. This, however, would lead to the inaccurate evaluation of the malware detection techniques. Removing label noises from Android malware datasets can be quite challenging, especially at a large data scale. To address this problem, we propose an effective approach called MalWhiteout to reduce label errors in Android malware datasets. Specifically, we creatively introduce Confident Learning (CL), an advanced noise estimation approach, to the domain of Android malware detection. To combat false positives introduced by CL, we incorporate the idea of ensemble learning and inter-app relation to achieve a more robust capability in noise detection. We evaluate MalWhiteout on a curated large-scale and reliable benchmark dataset. Experimental results show that MalWhiteout is capable of detecting label noises with over 94% accuracy even at a high noise ratio (i.e., 30%) of the dataset. MalWhiteout outperforms the state-of-the-art approach in terms of both effectiveness (8% to 218% improvement) and efficiency (70 to 249 times faster) across different settings. By reducing label noises, we show that the performance of existing malware detection approaches can be improved.

Please use this identifier to cite or link to this item:

http://hdl.handle.net/10453/168344