Encoding and Annotation Schemes
摘要
The internet is multilingual and languages use different scripts. This chapter describes Unicode, a standard to encode about all the existing characters, and the Unicode regular expressions. Once encoded, most texts not only contain sequences of characters, but also embed a structure in the form of markups. This chapter outlines them as well as elementary techniques to collect corpora from the internet and parse their markup.