Language in the Technology Trap: The Impact of the Increasingly Large Proportion of Machine-Generated Texts on Language Use
摘要
Neural machine translation programs can operate only within the framework of variants they have been trained with. Their efficiency has been increased at the expense of transparency: in fact, the algorithms according to which the machine prefers one variant over the other - probability, frequency, coercivity - determine the output of Deep Learning. A neural machine translation program provides significantly better results if it runs first through an unsupervised training on row data, and then it undergoes a human-supervised training. In disadvantageous cases, this training with human-selected data entails ready-made assumptions and biases, and it might be influenced by the trainer’s subjective choices. This particulate corpus should represent the language norm. Based on that standard, the program will calculate probabilities for sentence constructions. A lower probability might result for a correct sentence (or any grammatically correct one) than for an incorrect one: for example, if several of its word forms did not appear in the training corpus. In that case, language technology-based applications and increasing reliance on real-time language technologies can have a major impact on the language structure. The structures using non-inflected forms are preferred in the machine translation process. Contemporary development of linguistic technologies - translation technology in particular - changes the proportions of natural and artificial text production by influencing the emergence and disappearance of diverse language patterns in a non-transparent way. Indeed, an exacerbation of frequently observed patterns and a loss of less frequent ones not only strengthen biases present in used datasets, but they could also lead to an artificially depleted language caused by algorithmic biases.