<p>Patent landscaping is the process of identifying all patents related to a particular technological area, and is important for assessing various aspects of the intellectual property context. Traditionally, constructing patent landscapes is intensely laborious and expensive, and the rapid expansion of patenting activity in recent decades has driven an increasing need for efficient and effective automated patent landscaping approaches. In particular, it is critical that we be able to construct patent landscapes using a minimal number of labeled examples, as labeling patents for a narrow technology area requires highly specialized (and hence expensive) technical knowledge. We present an automated neural patent landscaping system that demonstrates significantly improved performance on difficult examples (0.69 <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10506_2025_9483_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="18" /> </InlineMediaObject> <EquationSource Format="TEX">\(F_1\)</EquationSource> </InlineEquation> on ‘hard’ examples, versus 0.6 for previously reported systems), and also significant improvements with much less training data (overall 0.75 <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10506_2025_9483_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="18" /> </InlineMediaObject> <EquationSource Format="TEX">\(F_1\)</EquationSource> </InlineEquation> on as few as 24 examples). Furthermore, in evaluating such automated landscaping systems, acquiring good data is challenge; we demonstrate a higher-quality training data generation procedure by merging (Abood and Feltenberger Artif Intell Law 26:103–125 <CitationRef CitationID="CR1">2018</CitationRef>) “seed/anti-seed” approach with active learning to collect difficult labeled examples near the decision boundary. Using this procedure we created a new dataset of labeled AI patents for training and testing. As in prior work we compare our approach with a number of baseline systems, and we release our code and data for others to build upon “(Code and data may be downloaded from <a href="https://doi.org/10.34703/gzx1-9v95/QDLKVW">https://doi.org/10.34703/gzx1-9v95/QDLKVW</a>Code and data are released under the Creative Commons NC-BY 4.0 license at <a href="https://creativecommons.org/licenses/by-nc/4.0/">https://creativecommons.org/licenses/by-nc/4.0/</a>)”.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated neural patent landscaping in the small data regime using citations and CPC codes

  • Tisa Islam Erana,
  • Mark A. Finlayson

摘要

Patent landscaping is the process of identifying all patents related to a particular technological area, and is important for assessing various aspects of the intellectual property context. Traditionally, constructing patent landscapes is intensely laborious and expensive, and the rapid expansion of patenting activity in recent decades has driven an increasing need for efficient and effective automated patent landscaping approaches. In particular, it is critical that we be able to construct patent landscapes using a minimal number of labeled examples, as labeling patents for a narrow technology area requires highly specialized (and hence expensive) technical knowledge. We present an automated neural patent landscaping system that demonstrates significantly improved performance on difficult examples (0.69 \(F_1\) on ‘hard’ examples, versus 0.6 for previously reported systems), and also significant improvements with much less training data (overall 0.75 \(F_1\) on as few as 24 examples). Furthermore, in evaluating such automated landscaping systems, acquiring good data is challenge; we demonstrate a higher-quality training data generation procedure by merging (Abood and Feltenberger Artif Intell Law 26:103–125 2018) “seed/anti-seed” approach with active learning to collect difficult labeled examples near the decision boundary. Using this procedure we created a new dataset of labeled AI patents for training and testing. As in prior work we compare our approach with a number of baseline systems, and we release our code and data for others to build upon “(Code and data may be downloaded from https://doi.org/10.34703/gzx1-9v95/QDLKVWCode and data are released under the Creative Commons NC-BY 4.0 license at https://creativecommons.org/licenses/by-nc/4.0/)”.