Detecting drifting concepts on the internet

Chien I. Lee, Cheng Jung Tsai, Chien Hui Hsieh

研究成果: Article

2 引文 斯高帕斯(Scopus)

摘要

With the explosive growth of information sources available on the World Wide Web, it has become increasingly necessary to utilize automated tools to discovery interesting and potentially useful patterns from data on the Internet. Since the data on the Internet such as communication packages, email, and e-commerce transactions come consecutively, an efficient and accurate incremental learning approach is required. Moreover, since the labels of these data may change over time, the problem of concept drift must be considered while incrementally learning from the data on the Internet. In this paper, we give a detailed discussion of the concept-drifting problem on the Internet. We also address a new problem called two-way drift. An approach adapted to the occurrence of concept drift is then proposed as the solution to incrementally learn from the data on the Internet. Our approach works as a preprocessor to detect the occurrence of concept drift and can be incorporated into any existing classification techniques. Our approach can also reveal which attribute values cause concept drift and therefore enables systems or decision makers to adopt proper decision in advance.

原文English
頁(從 - 到)229-236
頁數8
期刊Journal of Internet Technology
9
發行號3
出版狀態Published - 2008 七月 1

    指紋

All Science Journal Classification (ASJC) codes

  • Software
  • Computer Networks and Communications

引用此