The use of Transformers for text processing has attracted a large deal of attention in the last years. This is particularly true for sentence models, which present high capacity to comprehend and generate text contextually, improving the predictive performance in different Natural Language Processing tasks, when compared with previous approaches. Even so, there are still several challenges when applied to long documents, especially for some knowledge areas with very specific characteristics, such as legislative proposals. This study investigated different strategies for utilizing BERT-based models in long document retrieval written in Brazilian Portuguese. We used three corpora from the Brazilian Chamber of Deputies to build a dataset and assess the models, incorporating zero-shot and fine-tuning strategies. Five sentence models were evaluated: BERTimbau, LegalBert, LegalBert-pt, LegalBERTimbau, and LaBSE. We also assessed a summarized corpus of bills considering the input size limitation of the sentence models. Finaly, we propose a hybrid model, named HIRS, combining BM25 and BERTimbau with fine-tuning. According to the experimental results, the predictive performance obtained by HIRS was superior to the performance obtained by the other models, with a Recall of 84.78% for 20 documents.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HIRS: A Hybrid Information Retrieval System for Legislative Documents

  • José Antônio dos Santos,
  • Ellen Souza,
  • Carmelo J. A. Bastos Filho,
  • Hidelberg O. Albuquerque,
  • Douglas Vitório,
  • Danilo Carlos Gouveia de Lucena,
  • Nádia Silva,
  • André de Carvalho

摘要

The use of Transformers for text processing has attracted a large deal of attention in the last years. This is particularly true for sentence models, which present high capacity to comprehend and generate text contextually, improving the predictive performance in different Natural Language Processing tasks, when compared with previous approaches. Even so, there are still several challenges when applied to long documents, especially for some knowledge areas with very specific characteristics, such as legislative proposals. This study investigated different strategies for utilizing BERT-based models in long document retrieval written in Brazilian Portuguese. We used three corpora from the Brazilian Chamber of Deputies to build a dataset and assess the models, incorporating zero-shot and fine-tuning strategies. Five sentence models were evaluated: BERTimbau, LegalBert, LegalBert-pt, LegalBERTimbau, and LaBSE. We also assessed a summarized corpus of bills considering the input size limitation of the sentence models. Finaly, we propose a hybrid model, named HIRS, combining BM25 and BERTimbau with fine-tuning. According to the experimental results, the predictive performance obtained by HIRS was superior to the performance obtained by the other models, with a Recall of 84.78% for 20 documents.