Linguistic statistical universals: comparing computer- and human-generated texts
摘要
This paper aims at testing the ability of artificial text samples generated by transformers of replicating the writing style of various authors across different languages. We fine-tune GPT-2-based models with corpora from Jane Austen (English), Jules Verne (French) and Giovanni Verga (Italian). Then we analyse the samples in terms of (i) lexical distribution; (ii) long term correlations; and (iii) entropy. As a benchmark, we use text samples generated as Markov chains of different orders trained on the corpora of the same authors. Our results show that transformers represent a great improvement in terms of capturing long range correlations and entropy reduction, although the same cannot be said about lexical distribution.