<p>Humor in human communication is intrinsically multimodal, emerging from complex interactions among language, voice, facial expression, and affect. Yet computational studies have focused almost exclusively on English, leaving Spanish—and Latin-American varieties in particular—largely unexplored. We introduce <span>LS-FUNNY</span>, the first public corpus for multimodal humor detection in Latin-American Spanish, built from 272 TEDx talks and comprising 2040 balanced instances annotated via laughter markers. A unified pipeline extracts five complementary modalities: visual cues with OpenFace, acoustic descriptors with <i>Librosa</i>, contextual sentence embeddings from XLM-RoBERTa, affective scores from the NRC VAD lexicon, and a binary flag denoting figurative language detected by GPT-4o. We benchmark three off-the-shelf classifiers—SVM, XGBoost, and a feed-forward MLP—over all feature combinations. The best result (accuracy and F1 = 0.65) is obtained by XGBoost when fusing <i>all</i> modalities, confirming the benefit of multimodal integration. Remarkably, a zero-shot GPT-4o baseline fed with text only reaches 0.66 accuracy and 0.64 F1, highlighting both the strength of large language models and the residual cues missed without audio–visual context. We release the dataset and processing scripts to foster reproducible research and outline concrete paths for enlarging the corpus through automatic transcription and laughter detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Multimodal Humor Detection in Latin-American Spanish with LS-FUNNY

  • Eduardo Herrera-Alba,
  • Rubén Manrique

摘要

Humor in human communication is intrinsically multimodal, emerging from complex interactions among language, voice, facial expression, and affect. Yet computational studies have focused almost exclusively on English, leaving Spanish—and Latin-American varieties in particular—largely unexplored. We introduce LS-FUNNY, the first public corpus for multimodal humor detection in Latin-American Spanish, built from 272 TEDx talks and comprising 2040 balanced instances annotated via laughter markers. A unified pipeline extracts five complementary modalities: visual cues with OpenFace, acoustic descriptors with Librosa, contextual sentence embeddings from XLM-RoBERTa, affective scores from the NRC VAD lexicon, and a binary flag denoting figurative language detected by GPT-4o. We benchmark three off-the-shelf classifiers—SVM, XGBoost, and a feed-forward MLP—over all feature combinations. The best result (accuracy and F1 = 0.65) is obtained by XGBoost when fusing all modalities, confirming the benefit of multimodal integration. Remarkably, a zero-shot GPT-4o baseline fed with text only reaches 0.66 accuracy and 0.64 F1, highlighting both the strength of large language models and the residual cues missed without audio–visual context. We release the dataset and processing scripts to foster reproducible research and outline concrete paths for enlarging the corpus through automatic transcription and laughter detection.