复杂调查中自由回答文本数据的主题建模

Topic modelling for free-response text data from a complex survey

Journal of the Royal Statistical Society. Series A: Statistics in Society · 2025
被引 0
ABS 3

中文导读

针对复杂调查中忽略抽样设计会导致估计偏差的问题,提出一种结合调查权重的伪似然混合一元模型,并在ANES数据上验证其提取主题的有效性。

Abstract

Abstract Topic modeling, particularly the Mixture of Unigrams (MoU), is a standard statistical tool for identifying thematic structure and clustering documents, especially useful for analyzing open-ended survey responses. However, complex surveys often involve informative sampling, which, if ignored, leads to biased estimates. We address this by proposing a pseudolikelihood-based MoU model that incorporates survey weights to properly account for the sample design. We evaluate this approach using a simulation study and apply it to two American National Election Studies (ANES) datasets, demonstrating its effectiveness in extracting meaningful topics compared to the traditional MoU. Additionally, we introduce a Hierarchical Mixture of Unigrams (hMoU) for informative sampling, where topic proportions are modeled as a function of document-level fixed and random effects. We apply the hMoU to ANES data, comparing topic proportions across respondent factors such as sex, race, age, and state.

主题模型调查数据分析文本聚类统计方法