Effects of Language and Topic Size in Patent IR: An Empirical Study

F. Piroi, M. Lupu, A. Hanbury:
"Effects of Language and Topic Size in Patent IR: An Empirical Study";
Vortrag: Information Access Evaluation. Multilinguality, Multimodality, and Visual Analytics (CLEF 2012), Italy; 17.09.2012 - 20.09.2012; in:"CLEF 2012 - Conference and Labs of the Evaluation Forum", Lecture Notes in Computer Science, 7488 (2012), ISBN: 978-3-642-33246-3; S. 54 - 66.

We revisit the effects that various characteristics of the topic documents have on the effectiveness of the systems for the task of finding prior art in the patent domain. In doing so, we provide the reader interested in approaching the domain a guide of the issues that need to be addressed in this context.
For the current study, we select two patent based test collections with a common document representation schema and look at topic characteristics specific to the objectives of the collections. We look at the effect of languages on retrieval and at the length of the topic documents. We present the correlations between these topic facets and their retrieval results, as well as their relevant documents.