Efficient rank based KNN query processing over uncertain data

Zhang, Y; Lin, X; Zhu, G; Zhang, W; Lin, Q

Efficient rank based KNN query processing over uncertain data

Zhang, Y

Lin, X Zhu, G Zhang, W Lin, Q

Permalink

Publication Type:: Conference Proceeding
Citation:: Proceedings - International Conference on Data Engineering, 2010, pp. 28 - 39
Issue Date:: 2010-06-01

Closed Access

	Filename	Description	Size
	2013005467OK.pdf		337.16 kB	Adobe PDF	View/Open

Copyright Clearance Process

Recently Added
In Progress
Closed Access

This item is closed access and not available.

Full metadata record

Field	Value	Language
dc.contributor.author	Zhang, Y https://orcid.org/0000-0002-2674-1638	en_US
dc.contributor.author	Lin, X	en_US
dc.contributor.author	Zhu, G	en_US
dc.contributor.author	Zhang, W	en_US
dc.contributor.author	Lin, Q	en_US
dc.date.issued	2010-06-01	en_US
dc.identifier.citation	Proceedings - International Conference on Data Engineering, 2010, pp. 28 - 39	en_US
dc.identifier.isbn	9781424454440	en_US
dc.identifier.issn	1084-4627	en_US
dc.identifier.uri	http://hdl.handle.net/10453/28962
dc.description.abstract	Uncertain data are inherent in many applications such as environmental surveillance and quantitative economics research. As an important problem in many applications, KNN query has been extensively investigated in the literature. In this paper, we study the problem of processing rank based KNN query against uncertain data. Besides applying the expected rank semantic to compute KNN, we also introduce the median rank which is less sensitive to the outliers. We show both ranking methods satisfy nice top-k properties such as exact-k, containment, unique ranking, value invariance, stability and fairfulness. For given query q, IO and CPU efficient algorithms are proposed in the paper to compute KNN based on expected (median) ranks of the uncertain objects. To tackle the correlations of the uncertain objects and high IO cost caused by large number of instances of the uncertain objects, randomized algorithms are proposed to approximately compute KNN with theoretical guarantees. Comprehensive experiments are conducted on both real and synthetic data to demonstrate the efficiency of our techniques. © 2010 IEEE.	en_US
dc.relation.ispartof	Proceedings - International Conference on Data Engineering	en_US
dc.relation.isbasedon	10.1109/ICDE.2010.5447874	en_US
dc.title	Efficient rank based KNN query processing over uncertain data	en_US
dc.type	Conference Proceeding
utslib.for	0806 Information Systems	en_US
dc.location.activity	Long Beach, USA	en_US
pubs.embargo.period	Not known	en_US
pubs.organisational-group	/University of Technology Sydney
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology/School of Computer Science
pubs.organisational-group	/University of Technology Sydney/Strength - CAI - Centre for Artificial Intelligence
utslib.copyright.status	closed_access
pubs.publication-status	Published	en_US

Abstract:

Uncertain data are inherent in many applications such as environmental surveillance and quantitative economics research. As an important problem in many applications, KNN query has been extensively investigated in the literature. In this paper, we study the problem of processing rank based KNN query against uncertain data. Besides applying the expected rank semantic to compute KNN, we also introduce the median rank which is less sensitive to the outliers. We show both ranking methods satisfy nice top-k properties such as exact-k, containment, unique ranking, value invariance, stability and fairfulness. For given query q, IO and CPU efficient algorithms are proposed in the paper to compute KNN based on expected (median) ranks of the uncertain objects. To tackle the correlations of the uncertain objects and high IO cost caused by large number of instances of the uncertain objects, randomized algorithms are proposed to approximately compute KNN with theoretical guarantees. Comprehensive experiments are conducted on both real and synthetic data to demonstrate the efficiency of our techniques. © 2010 IEEE.

Please use this identifier to cite or link to this item:

http://hdl.handle.net/10453/28962