Clustering-based missing value imputation for data preprocessing

Zhang, C; Qin, Y; Zhu, X; Zhang, J; Zhang, S

Clustering-based missing value imputation for data preprocessing

Zhang, C

Qin, Y Zhu, X Zhang, J Zhang, S

Permalink

Publication Type:: Conference Proceeding
Citation:: 2006 IEEE International Conference on Industrial Informatics, INDIN'06, 2006, pp. 1081 - 1086
Issue Date:: 2006-01-01

Closed Access

	Filename	Description	Size
	2006005160.pdf		6.69 MB	Adobe PDF	View/Open

Copyright Clearance Process

Recently Added
In Progress
Closed Access

This item is closed access and not available.

Full metadata record

Field	Value	Language
dc.contributor.author	Zhang, C https://orcid.org/0000-0001-5715-7154	en_US
dc.contributor.author	Qin, Y	en_US
dc.contributor.author	Zhu, X	en_US
dc.contributor.author	Zhang, J	en_US
dc.contributor.author	Zhang, S	en_US
dc.date.issued	2006-01-01	en_US
dc.identifier.citation	2006 IEEE International Conference on Industrial Informatics, INDIN'06, 2006, pp. 1081 - 1086	en_US
dc.identifier.isbn	0780397010	en_US
dc.identifier.isbn	9780780397019	en_US
dc.identifier.uri	http://hdl.handle.net/10453/2673
dc.description.abstract	Missing value imputation is an actual yet challenging issue confronted by machine learning and data mining. Existing missing value imputation is a procedure that replaces the missing values in a dataset by some plausible values. The plausible values are generally generated from the dataset using a deterministic, or random method. In this paper we propose a new and efficient missing value imputation based on data clustering, called CRI (Clustering-based Random Imputation). In our approach, we fill up the missing values of an instance with those plausible values that are generated from the data similar to this instance using a kernel-based random method. Specifically, we first divide the dataset (exclude instances with missing values) into clusters. And then each of those instances with missing-values is assigned to a cluster most similar to it. Finally, missing values of an instance A are thus patched up with those plausible values that are generated using a kernel-based method to those instances from A's cluster. Our experiments (some of them are with the decision tree induction system C5.0) have proved the effectiveness of our proposed method in missing value imputation task. © 2006 IEEE.	en_US
dc.relation	http://purl.org/au-research/grants/arc/DP0449535
dc.relation	http://purl.org/au-research/grants/arc/DP0667060
dc.relation	http://purl.org/au-research/grants/arc/DP0559536
dc.relation.ispartof	2006 IEEE International Conference on Industrial Informatics, INDIN'06	en_US
dc.relation.isbasedon	10.1109/INDIN.2006.275767	en_US
dc.title	Clustering-based missing value imputation for data preprocessing	en_US
dc.type	Conference Proceeding
utslib.for	080704 Information Retrieval and Web Search	en_US
utslib.for	080604 Database Management	en_US
utslib.for	080109 Pattern Recognition and Data Mining	en_US
dc.location.activity	Singapore	en_US
pubs.embargo.period	Not known	en_US
pubs.organisational-group	/University of Technology Sydney
pubs.organisational-group	/University of Technology Sydney/DVC (International)
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology
pubs.organisational-group	/University of Technology Sydney/Strength - ACRI - Australia China Relations Institute
pubs.organisational-group	/University of Technology Sydney/Strength - CAI - Centre for Artificial Intelligence
utslib.copyright.status	closed_access
pubs.publication-status	Published	en_US

Abstract:

Missing value imputation is an actual yet challenging issue confronted by machine learning and data mining. Existing missing value imputation is a procedure that replaces the missing values in a dataset by some plausible values. The plausible values are generally generated from the dataset using a deterministic, or random method. In this paper we propose a new and efficient missing value imputation based on data clustering, called CRI (Clustering-based Random Imputation). In our approach, we fill up the missing values of an instance with those plausible values that are generated from the data similar to this instance using a kernel-based random method. Specifically, we first divide the dataset (exclude instances with missing values) into clusters. And then each of those instances with missing-values is assigned to a cluster most similar to it. Finally, missing values of an instance A are thus patched up with those plausible values that are generated using a kernel-based method to those instances from A's cluster. Our experiments (some of them are with the decision tree induction system C5.0) have proved the effectiveness of our proposed method in missing value imputation task. © 2006 IEEE.

Please use this identifier to cite or link to this item:

http://hdl.handle.net/10453/2673