Top-k similarity join over multi-valued objects

Zhang, W; Xu, J; Liang, X; Zhang, Y; Lin, X

Top-k similarity join over multi-valued objects

Zhang, W Xu, J Liang, X Zhang, Y

Lin, X

Permalink

Publication Type:: Conference Proceeding
Citation:: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2012, 7238 LNCS (PART 1), pp. 509 - 525
Issue Date:: 2012-05-11

Closed Access

	Filename	Description	Size
	2013005460OK.pdf		387.69 kB	Adobe PDF	View/Open

Copyright Clearance Process

Recently Added
In Progress
Closed Access

This item is closed access and not available.

Full metadata record

Field	Value	Language
dc.contributor.author	Zhang, W	en_US
dc.contributor.author	Xu, J	en_US
dc.contributor.author	Liang, X	en_US
dc.contributor.author	Zhang, Y https://orcid.org/0000-0002-2674-1638	en_US
dc.contributor.author	Lin, X	en_US
dc.date.issued	2012-05-11	en_US
dc.identifier.citation	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2012, 7238 LNCS (PART 1), pp. 509 - 525	en_US
dc.identifier.isbn	9783642290374	en_US
dc.identifier.issn	0302-9743	en_US
dc.identifier.uri	http://hdl.handle.net/10453/28954
dc.description.abstract	The top-k similarity joins have been extensively studied and used in a wide spectrum of applications such as information retrieval, decision making, spatial data analysis and data mining. Given two sets of objects and U and V, a top-k similarity join returns k pairs of most similar objects from U x V. In the conventional model of top-k similarity join processing, an object is usually regarded as a point in a multi-dimensional space and the similarity between two objects is usually measured by distance metrics such as Euclidean distance. However, in many applications an object may be described by multiple values (instances) and the conventional model is not applicable since it does not address the distributions of object instances. In this paper, we study top-k similarity join queries over multi-valued objects. We apply quantile based distance to explore the relative instance distribution among the multiple instances of objects. Efficient and effective techniques to process top-k similarity joins over multi-valued objects are developed following a filtering-refinement framework. Novel distance, statistic and weight based pruning techniques are proposed. Comprehensive experiments on both real and synthetic datasets demonstrate the efficiency and effectiveness of our techniques. © 2012 Springer-Verlag.	en_US
dc.relation.ispartof	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)	en_US
dc.relation.isbasedon	10.1007/978-3-642-29038-1_37	en_US
dc.subject.classification	Artificial Intelligence & Image Processing	en_US
dc.title	Top-k similarity join over multi-valued objects	en_US
dc.type	Conference Proceeding
utslib.citation.volume	PART 1	en_US
utslib.citation.volume	7238 LNCS	en_US
utslib.for	0806 Information Systems	en_US
dc.location.activity	Busan, Korea	en_US
pubs.embargo.period	Not known	en_US
pubs.organisational-group	/University of Technology Sydney
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology/School of Computer Science
pubs.organisational-group	/University of Technology Sydney/Strength - CAI - Centre for Artificial Intelligence
utslib.copyright.status	closed_access
pubs.issue	PART 1	en_US
pubs.publication-status	Published	en_US
pubs.volume	7238 LNCS	en_US

Abstract:

The top-k similarity joins have been extensively studied and used in a wide spectrum of applications such as information retrieval, decision making, spatial data analysis and data mining. Given two sets of objects and U and V, a top-k similarity join returns k pairs of most similar objects from U x V. In the conventional model of top-k similarity join processing, an object is usually regarded as a point in a multi-dimensional space and the similarity between two objects is usually measured by distance metrics such as Euclidean distance. However, in many applications an object may be described by multiple values (instances) and the conventional model is not applicable since it does not address the distributions of object instances. In this paper, we study top-k similarity join queries over multi-valued objects. We apply quantile based distance to explore the relative instance distribution among the multiple instances of objects. Efficient and effective techniques to process top-k similarity joins over multi-valued objects are developed following a filtering-refinement framework. Novel distance, statistic and weight based pruning techniques are proposed. Comprehensive experiments on both real and synthetic datasets demonstrate the efficiency and effectiveness of our techniques. © 2012 Springer-Verlag.

Please use this identifier to cite or link to this item:

http://hdl.handle.net/10453/28954