Cooperative strategy for Web data mining and cleaning

Publisher:
Taylor & Francis Inc
Publication Type:
Journal Article
Citation:
Applied Artificial Intelligence, 2003, 17 (5-6), pp. 443 - 460
Issue Date:
2003-01
Full metadata record
Files in This Item:
Filename Description Size
Thumbnail2003000345.pdf1.13 MB
Adobe PDF
While the Internet and World Wide Web have put a huge volume of low-quality information at the easy access of an information gathering system, filtering out irrelevant information has become a big challenge. In this paper, a Web data mining and cleaning strategy for information gathering is proposed. A data-mining model is presented for the data that come from multiple agents. Using the model, a data-cleaning algorithm is then presented to eliminate irrelevant data. To evaluate the data-cleaning strategy, an interpretation is given for the mining model according to evidence theory. An experiment is also conducted to evaluate the strategy using Web data. The experimental results have shown that the proposed strategy is efficient and promising.
Please use this identifier to cite or link to this item: