\(
\def\WIPO{World Intellectual Property Organisation}
\)
Rethinking patent retrieval with language models: Toward scalable and efficient search
2026
Formats
| Format | |
|---|---|
| BibTeX | |
| MARCXML | |
| TextMARC | |
| MARC | |
| DataCite | |
| DublinCore | |
| EndNote | |
| NLM | |
| RefWorks | |
| RIS | |
Citer
Titre
Rethinking patent retrieval with language models: Toward scalable and efficient search
Type d’élément
Journal article
Description
1 volume.
Résumé
Semantic search with embedding models offers an alternative to traditional keyword-based patent retrieval but often struggles with computational cost and efficiency in real-time scenarios compared to methods like BM25. Meanwhile, the rapid advancement of language models raises questions about the necessity of domain-specific models versus the viability of general-purpose ones. This work presents a comprehensive evaluation of embedding-based patent search using the CLEF-IP 2011 dataset. We assess 10 configurations employing language models as retrievers, re-rankers, or hybrids, across 9 models, both patent-specific and general-purpose, tested in 105 experimental setups. Our best configurations deliver a 14.81% absolute MAP improvement over state-of-the-art baselines and outperform patent-specific embeddings by at least 28.95% in MAP. We further show that embedding quantization enables large-scale patent search with up to 30×faster retrieval and 32×lower memory usage. These results provide practical guidance for integrating embedding models into patent prior art search while addressing performance and scalability constraints.
Source of Description
Crossref
Série
World Patent Information ; 84, March, 2026
Dans
World Patent Information
Ressources liées
Publié
Oxford [England] : Elsevier Ltd., 2026.
Langue
Anglais
Informations relatives au droit d’auteur
https://www.sciencedirect.com/science/article/abs/pii/S0172219023000108
Le document apparaît dans