References & Citations
Computer Science > Information Retrieval
Title: A Study on the Efficiency and Generalization of Light Hybrid Retrievers
(Submitted on 4 Oct 2022 (v1), last revised 23 May 2023 (this version, v2))
Abstract: Hybrid retrievers can take advantage of both sparse and dense retrievers. Previous hybrid retrievers leverage indexing-heavy dense retrievers. In this work, we study "Is it possible to reduce the indexing memory of hybrid retrievers without sacrificing performance"? Driven by this question, we leverage an indexing-efficient dense retriever (i.e. DrBoost) and introduce a LITE retriever that further reduces the memory of DrBoost. LITE is jointly trained on contrastive learning and knowledge distillation from DrBoost. Then, we integrate BM25, a sparse retriever, with either LITE or DrBoost to form light hybrid retrievers. Our Hybrid-LITE retriever saves 13X memory while maintaining 98.0% performance of the hybrid retriever of BM25 and DPR. In addition, we study the generalization capacity of our light hybrid retrievers on out-of-domain dataset and a set of adversarial attacks datasets. Experiments showcase that light hybrid retrievers achieve better generalization performance than individual sparse and dense retrievers. Nevertheless, our analysis shows that there is a large room to improve the robustness of retrievers, suggesting a new research direction.
Submission history
From: Man Luo [view email][v1] Tue, 4 Oct 2022 04:22:46 GMT (1716kb,D)
[v2] Tue, 23 May 2023 09:45:30 GMT (1792kb,D)
Link back to: arXiv, form interface, contact.