An Empirical Study of Structural Rewards in Reinforcement Learning for Legal Summarization
摘要
Legal case summarization requires not only semantic correctness but also coherent argumentative structure. However, existing summarization models do not capture the structural regularities observed in human-written legal summaries. In this paper, we study structure-aware reinforcement learning for legal summarization with rewards that explicitly capture argumentative element proportions, length alignment, and positional organization. We first conduct an analysis of argumentative structure in human-written legal summaries. Based on the empirical analysis, we design a set of structure-aware rewards. Experiments on the Indian Supreme Court dataset show that these rewards consistently improve structural alignment. Through diagnostic analysis, we further reveal a structural–semantic trade-off: structural constraints can restrict semantically reasoning components, leading to mild degradation in semantic quality. Our findings highlight the importance of flexibility-aware structural guidance for reinforcement learning–based legal summarization.