Cooperative strategy for web data mining and cleaning
- Publication Type:
- Journal Article
- Applied Artificial Intelligence, 2003, 17 (5-6), pp. 443 - 460
- Issue Date:
While the Internet and World Wide Web have put a huge volume of low-quality information at the easy access of an information gathering system, filtering out irrelevant information has become a big challenge. In this paper, a Web data mining and cleaning strategy for information gathering is proposed. A data-mining model is presented for the data that come from multiple agents. Using the model, a data-cleaning algorithm is then presented to eliminate irrelevant data. To evaluate the data-cleaning strategy, an interpretation is given for the mining model according to evidence theory. An experiment is also conducted to evaluate the strategy using Web data. The experimental results have shown that the proposed strategy is efficient and promising. © 2003 Taylor and Francis Group, LLC.
Please use this identifier to cite or link to this item: