<p>This paper describes a dataset of community-level diversity indicators for origin, gender and academic age across research communities in UK higher education institutions. The indicators were constructed from individual-level bibliometric metadata for 535,456 research-active individuals in OpenAlex, using onomastic inference for gender and country of origin and publication histories for academic age, with confidence-based filtering to manage inference uncertainty. The unit of analysis is the institution–Unit of Assessment pair defined by the UK Research Excellence Framework (REF) 2021, and the resulting REF 2021-Diversity dataset comprises 1,848 records across 155 institutions and 34 Units of Assessment. The indicators range from simple compositional statistics through standard entropy-based indices to multidimensional Leinster–Cobbold indices to accommodate various use cases and assumptions. The dataset can support benchmarking of research community composition, methodological comparison of diversity measures, and cautiously specified association analyses between community-level diversity indicators and REF-linked quality profiles.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Origin, gender and academic age diversity indicators for UK research communities

  • Abdullah Gök,
  • Greg Macmillan

摘要

This paper describes a dataset of community-level diversity indicators for origin, gender and academic age across research communities in UK higher education institutions. The indicators were constructed from individual-level bibliometric metadata for 535,456 research-active individuals in OpenAlex, using onomastic inference for gender and country of origin and publication histories for academic age, with confidence-based filtering to manage inference uncertainty. The unit of analysis is the institution–Unit of Assessment pair defined by the UK Research Excellence Framework (REF) 2021, and the resulting REF 2021-Diversity dataset comprises 1,848 records across 155 institutions and 34 Units of Assessment. The indicators range from simple compositional statistics through standard entropy-based indices to multidimensional Leinster–Cobbold indices to accommodate various use cases and assumptions. The dataset can support benchmarking of research community composition, methodological comparison of diversity measures, and cautiously specified association analyses between community-level diversity indicators and REF-linked quality profiles.