利用全文特征预测科学概念的影响

Predicting the impact of scientific concepts using full‐text features

Journal of the Association for Information Science and Technology (JASIST) · 2016
被引 75
ABS 3

中文导读

研究了一个系统,利用近期研究论文的全文特征(如修辞句分析、信息抽取和时间序列分析)来预测新科学概念的未来影响力,发现全文特征比仅用元数据更有效,且两者结合效果最佳。

Abstract

New scientific concepts, interpreted broadly, are continuously introduced in the literature, but relatively few concepts have a long‐term impact on society. The identification of such concepts is a challenging prediction task that would help multiple parties—including researchers and the general public—focus their attention within the vast scientific literature. In this paper we present a system that predicts the future impact of a scientific concept, represented as a technical term, based on the information available from recently published research articles. We analyze the usefulness of rich features derived from the full text of the articles through a variety of approaches, including rhetorical sentence analysis, information extraction, and time‐series analysis. The results from two large‐scale experiments with 3.8 million full‐text articles and 48 million metadata records support the conclusion that full‐text features are significantly more useful for prediction than metadata‐only features and that the most accurate predictions result from combining the metadata and full‐text features. Surprisingly, these results hold even when the metadata features are available for a much larger number of documents than are available for the full‐text features.

科学计量学信息检索自然语言处理预测分析