<p>The automatic analysis of political discourse requires systematic resources, particularly for non-Anglophone electoral debates, which differ from parliamentary sessions. This paper describes the DebatES dataset, containing manually reviewed transcriptions of nearly all Spanish nationally televised general election debates since 1993. Transcripts are segmented into turns and thematic blocks and include participant metadata. The data is enriched with linguistic-stylistic metrics derived from NLP analysis and extensive annotations combining Large Language Models and manual validation. Annotations cover turn topics, emotions, relevant entity mentions, electoral proposals, and factual claims. To facilitate access and reuse, the dataset is distributed in standard formats (XML and CSV), accompanied by interactive reports for visual exploration. This dataset provides a valuable resource for researchers in linguistics, political science, and language technologies, opening new avenues for studying the evolution of discourse, persuasive strategies, and ideological polarisation within the Spanish political context.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Annotated Spanish general election debate transcriptions 1993-2023

  • Fermín L. Cruz,
  • Fernando Enríquez,
  • F. Javier Ortega,
  • José A. Troyano

摘要

The automatic analysis of political discourse requires systematic resources, particularly for non-Anglophone electoral debates, which differ from parliamentary sessions. This paper describes the DebatES dataset, containing manually reviewed transcriptions of nearly all Spanish nationally televised general election debates since 1993. Transcripts are segmented into turns and thematic blocks and include participant metadata. The data is enriched with linguistic-stylistic metrics derived from NLP analysis and extensive annotations combining Large Language Models and manual validation. Annotations cover turn topics, emotions, relevant entity mentions, electoral proposals, and factual claims. To facilitate access and reuse, the dataset is distributed in standard formats (XML and CSV), accompanied by interactive reports for visual exploration. This dataset provides a valuable resource for researchers in linguistics, political science, and language technologies, opening new avenues for studying the evolution of discourse, persuasive strategies, and ideological polarisation within the Spanish political context.