Building Sign Language Datasets
摘要
This chapter focuses on the methodologies and frameworks essential for building comprehensive sign language datasets. It outlines the critical steps in data collection, including participant recruitment, video recording, and annotation processes, ensuring high-quality and representative data. The chapter discusses the different types of sign language datasets, such as lexical databases, conversational corpora, and annotated video corpora, highlighting their importance for various research and technological applications. Additionally, it addresses the challenges and best practices in dataset creation, emphasizing the need for ethical considerations and community involvement. By providing a detailed guide on constructing robust sign language datasets, the chapter aims to support researchers and developers in advancing sign language technologies and promoting inclusive linguistic research.