The aboutness of words
定义了词汇与主题之间的关于性关系,并开发了关于性系数来估计这种关系的强度。通过分析OCLC WorldCat数据库中非小说类英文图书标题中的词汇使用模式,计算常见标题词汇的关于性系数,发现低系数词汇常见于停用词表,高系数词汇则具有明确主题关联。该系数可增强索引、规范控制和检索效果。
Word aboutness is defined as the relationship between words and subjects associated with them. An aboutness coefficient is developed to estimate the strength of the aboutness relationship. Words that are randomly distributed across subjects are assumed to lack aboutness and the degree to which their usage deviates from a random pattern indicates the strength of the aboutness. To estimate aboutness, title words and their associated subjects are extracted from the titles of non‐fiction English language books in the OCLC WorldCat database. The usage patterns of the title words are analyzed and used to compute aboutness coefficients for each of the common title words. Words with low aboutness coefficients ( An and In ) are commonly found in stop word lists, whereas words with high aboutness coefficients ( Carbonate , Autism ) are unambiguous and have a strong subject association. The aboutness coefficient potentially can enhance indexing, advance authority control, and improve retrieval.