Other «Previous Next»

UMass at TREC 2004: Novelty and HARD

Nasreen Abdul-Jaleel, University of Massachusetts - Amherst
James Allan, University of Massachusetts - Amherst
W. Bruce Croft, University of Massachusetts - Amherst
Fernando Diaz, University of Massachusetts - Amherst
Leah Larkey, University of Massachusetts - Amherst
Xiaoyan Li, University of Massachusetts - Amherst
Mark D. Smucker, University of Massachusetts - Amherst
Courtney Wade, University of Massachusetts - Amherst

Article comments

This paper was harvested from CiteSeer

Abstract

For the TREC 2004 Novelty track, UMass participated in all four tasks. Although finding relevant sentences was harder this year than last, we continue to show marked improvements over the baseline of calling all sentences relevant, with a variant of tfidf being the most successful approach. We achieve 5–9%improvements over the baseline in locating novel sentences, primarily by looking at the similarity of a sentence to earlier sentences and focusing on named entities. For the High Accuracy Retrieval from Documents (HARD) track, we investigated the use of clarification forms, fixed- and variable-length passage retrieval, and the use of metadata. Clarification form results indicate that passage level feedback can provide improvements comparable to user supplied related-text for document evaluation and outperforms related-text for passage evaluation. Document retrieval methods without a query expansion component show themost gains fromrelated-text. We also found that displaying the top passages for feedback outperformed displaying centroid passages. Named entity feedback resulted in mixed performance. Our primary findings for passage retrieval are that document retrieval methods performed better than passage retrieval methods on the passage evaluation metric of binary preference at 12,000 characters, and that clarification forms improved passage retrieval for every retrieval method explored. We found no benefit to using variable-length passages over fixed-length passages for this corpus. Our use of geography and genremetadata resulted in no significant changes in retrieval performance.

Suggested Citation

Nasreen Abdul-Jaleel, James Allan, W. Bruce Croft, Fernando Diaz, Leah Larkey, Xiaoyan Li, Mark D. Smucker, and Courtney Wade. "UMass at TREC 2004: Novelty and HARD" 2004
Available at: http://works.bepress.com/james_allan/9