<p>This paper aims at testing the ability of artificial text samples generated by transformers of replicating the writing style of various authors across different languages. We fine-tune GPT-2-based models with corpora from Jane Austen (English), Jules Verne (French) and Giovanni Verga (Italian). Then we analyse the samples in terms of (i) lexical distribution; (ii) long term correlations; and (iii) entropy. As a benchmark, we use text samples generated as Markov chains of different orders trained on the corpora of the same authors. Our results show that transformers represent a great improvement in terms of capturing long range correlations and entropy reduction, although the same cannot be said about lexical distribution.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Linguistic statistical universals: comparing computer- and human-generated texts

  • Marco Civico

摘要

This paper aims at testing the ability of artificial text samples generated by transformers of replicating the writing style of various authors across different languages. We fine-tune GPT-2-based models with corpora from Jane Austen (English), Jules Verne (French) and Giovanni Verga (Italian). Then we analyse the samples in terms of (i) lexical distribution; (ii) long term correlations; and (iii) entropy. As a benchmark, we use text samples generated as Markov chains of different orders trained on the corpora of the same authors. Our results show that transformers represent a great improvement in terms of capturing long range correlations and entropy reduction, although the same cannot be said about lexical distribution.