从带有主题漂移的短文本流中学习

Learning From Short Text Streams With Topic Drifts

IEEE Transactions on Cybernetics · 2017
被引 37
ABS 3

中文导读

提出一种利用大规模语义网络扩展特征的方法,结合增量集成分类模型,有效处理短文本流中的主题漂移问题,提升分类效率与准确性。

Abstract

Short text streams such as search snippets and micro blogs have been popular on the Web with the emergence of social media. Unlike traditional normal text streams, these data present the characteristics of short length, weak signal, high volume, high velocity, topic drift, etc. Short text stream classification is hence a very challenging and significant task. However, this challenge has received little attention from the research community. Therefore, a new feature extension approach is proposed for short text stream classification with the help of a large-scale semantic network obtained from a Web corpus. It is built on an incremental ensemble classification model for efficiency. First, more semantic contexts based on the senses of terms in short texts are introduced to make up of the data sparsity using the open semantic network, in which all terms are disambiguated by their semantics to reduce the noise impact. Second, a concept cluster-based topic drifting detection method is proposed to effectively track hidden topic drifts. Finally, extensive studies demonstrate that as compared to several well-known concept drifting detection methods in data stream, our approach can detect topic drifts effectively, and it enables handling short text streams effectively while maintaining the efficiency as compared to several state-of-the-art short text classification approaches.

计算机科学数据挖掘自然语言处理文本分类流数据挖掘