使用文本挖掘追踪全球新兴疾病监测中的疫情趋势:ProMED-mail

Using Text Mining to Track Outbreak Trends in Global Surveillance of Emerging Diseases: ProMED-mail

Journal of the Royal Statistical Society. Series A: Statistics in Society · 2021
被引 10
ABS 3

中文导读

针对ProMED-mail报告分布不均的问题,提出结合TextRank关键词提取与共现网络分析的方法,以埃博拉和寨卡疫情验证,能有效提取疫情演化信息,辅助监测与响应。

Abstract

Abstract ProMED-mail (Program for Monitoring Emerging Disease) is an international disease outbreak monitoring and early warning system. Every year, users contribute thousands of reports that include reference to infectious diseases and toxins. However, due to the uneven distribution of the reports for each disease, traditional statistics-based text mining techniques, represented by term frequency-related algorithm, are not suitable. Thus, we conducted a study in three steps (i) report filtering, (ii) keyword extraction from reports and finally (iii) word co-occurrence network analysis to fill the gap between ProMED and its utilization. The keyword extraction was performed with the TextRank algorithm, keywords co-occurrence networks were then produced using the top keywords from each document and multiple network centrality measures were computed to analyse the co-occurrence networks. We used two major outbreaks in recent years, Ebola, 2014 and Zika 2015, as cases to illustrate and validate the process. We found that the extracted information structures are consistent with World Health Organisation description of the timeline and phases of the epidemics. Our research presents a pipeline that can extract and organize the information to characterize the evolution of epidemic outbreaks. It also highlights the potential for ProMED to be utilized in monitoring, evaluating and improving responses to outbreaks.

传染病监测文本挖掘疫情预警网络分析