Leveraging Domain Expertise: The Impact of Media Experts in Developing Large Language Models
摘要
Large language models (LLMs) have shown remarkable capabilities in generating high-quality text based on large amounts of data. Nevertheless, in practical applications, the general LLMs do not sufficiently address the unique demands of the media domain, necessitating the development of media-domain-specific LLMs. Based on this, we conduct an experiment to analyze whether the news expert’s knowledge can help optimize LLMs. In the process, we fine-tune a LLM named MediaGPT, using a prompt framework that fits diversified journalism task and a journalistic professional data set.Through human experts and strong model evaluation, this paper demonstrates that MediaGPT outperforms mainstream models on various Chinese media domain tasks and verifies the importance of domain data and domain-defined prompt types for building an effective domain-specific LLM. Meanwhile, this experiment has also proven that journalists play an irreplaceable role in the development of the emerging LLM technology. In the future, media-domain-specific LLMs, which will play increasingly important role in news production, will require a collaborative effort between news experts and technologists to advance. Only through such cooperation can we create advanced tools that better meet the real needs of the media industry.