Identifying Discourse Markers in French Spoken Corpora: Using Machine Learning and Rule-Based Approaches
摘要
The objective of this work is to study the identification of French discourse markers (DM), in particular the polyfunctional occurrences such as ‘attetion’, bon, quoi, la preuve. A number of words identified as DM, and traditionally considered as adverbs or interjections, are also, for instance, adjectives or nouns. For example bon can be a DM or an adjective, ‘attetion’ can be a DM or a noun, etc. Hand annotation is in general robust but time consuming. The main difficulty with automatic identification is to take the context of the DM candidate correctly into account. To do that, a mechanisms based on rule-based and machine learning approaches was built, in order to reach an acceptable level of performance and reduce the expert effort. This study will provide a comprehensive use case of a machine learning algorithm, which has proved a good efficiency in dealing with such linguistic phenomena. In addition, an evaluation was done for the Unitex platform in order to determine the efficiency and drawbacks of this platform when dealing with such type of tasks.