Detecting drifting concepts on the internet

Chien I. Lee, Cheng Jung Tsai, Chien Hui Hsieh

Research output: Contribution to journalArticle

2 Citations (Scopus)

Abstract

With the explosive growth of information sources available on the World Wide Web, it has become increasingly necessary to utilize automated tools to discovery interesting and potentially useful patterns from data on the Internet. Since the data on the Internet such as communication packages, email, and e-commerce transactions come consecutively, an efficient and accurate incremental learning approach is required. Moreover, since the labels of these data may change over time, the problem of concept drift must be considered while incrementally learning from the data on the Internet. In this paper, we give a detailed discussion of the concept-drifting problem on the Internet. We also address a new problem called two-way drift. An approach adapted to the occurrence of concept drift is then proposed as the solution to incrementally learn from the data on the Internet. Our approach works as a preprocessor to detect the occurrence of concept drift and can be incorporated into any existing classification techniques. Our approach can also reveal which attribute values cause concept drift and therefore enables systems or decision makers to adopt proper decision in advance.

Original languageEnglish
Pages (from-to)229-236
Number of pages8
JournalJournal of Internet Technology
Volume9
Issue number3
Publication statusPublished - 2008 Jul 1

All Science Journal Classification (ASJC) codes

  • Software
  • Computer Networks and Communications

Fingerprint Dive into the research topics of 'Detecting drifting concepts on the internet'. Together they form a unique fingerprint.

  • Cite this