Accounting for Measurement Bias: A New Framework for Reliable Country Ranking in Large-Scale Educational Assessments
针对PISA和TIMSS等国际大规模评估中因语言、文化差异导致的测量偏差,提出一种无需预设锚题或参照组的新方法,能高效校正排名,并应用于PISA 2022数据得到修正后的数学、科学和阅读排名。
International Large-scale Assessments (ILSAs), such as the Program for International Student Assessment (PISA) and the Trends in International Mathematics and Science Study (TIMSS), are cornerstone tools for global educational research and policy-making. By benchmarking educational quality and performance trends, these assessments enable countries to evaluate and share effective pedagogical structures. Specifically, ILSAs employ Item Response Theory (IRT) models to rank countries by students’ performance on cognitive items. However, measurement bias—arising from linguistic, cultural, and curricular differences—poses a significant threat to the statistical inference of IRT models and, consequently, the validity of the resulting rankings. Neglecting this bias can lead to systematic errors in parameter estimation, ultimately distorting national standings. To address this, we propose a novel method that avoids the restrictive assumptions typical of existing approaches, such as the prior identification of unbiased “anchor items" or designated reference groups. Our approach is computationally efficient and provides theoretical guarantees for the reliable recovery of group rankings. We apply this method to PISA 2022 data across the mathematics, science, and reading domains, yielding corrected performance rankings and insights into the survey’s measurement-bias structures.