Methodological challenges in content‐based citation analysis: Expertise, reliability, and the primacy of citance identification
研究了引文分割和标注者专业知识如何影响分类一致性,发现引文长度是分歧的最强预测因子,强调引文识别的方法论改进对提高可重复性和研究评价有效性的关键作用。
Abstract Content‐based citation analysis seeks to capture the meaning and functions of citations but continues to face unresolved methodological challenges. This study analyzes a stratified sample of library and information science publications to examine how citance segmentation and annotator expertise influence the consistency of classification. Using two annotators with different professional backgrounds, the findings show that agreement is high when citances are defined identically, but reliability decreases sharply once text boundaries diverge. Citance length, rather than subject category or citation density, emerges as the strongest predictor of disagreement. These results identify segmentation as a methodological rather than a purely technical issue, shaping both human and automated tagging outcomes. By highlighting the interplay between expertise effects and boundary definitions, the study underscores the need for clearer operational frameworks in citation analysis. The contribution lies in demonstrating that methodological refinements in citance identification are essential for improving reproducibility, enhancing hybrid human–machine approaches, and strengthening the validity of citation‐based indicators in research evaluation.