CacheChecker.run issues one filter_documents call per item (N+1 query pattern)
Bug
CacheChecker.run (haystack/components/caching/cache_checker.py:75) loops over the input items list and calls self.document_store.filter_documents(...) once per item:
for item in items:
filters = {"field": self.cache_field, "operator": "==", "value": item}
found = self.document_store.filter_documents(filters=filters)
if found:
found_documents.extend(found)
else:
misses.append(item)
For an input of N items, this is N round-trips to the document store. For InMemoryDocumentStore the constant is small; for QdrantDocumentStore, WeaviateDocumentStore, OpenSearchDocumentStore, etc., every call is a network request. With a 10 ms per-call latency, 200 items takes ~2 s; 5 000 items (a realistic batch) takes ~50 s.
The same pattern is duplicated in CacheChecker.run_async (cache_checker.py:99).
The Haystack filter spec already supports the in operator (haystack/document_stores/types/protocol.py:70, haystack/utils/filters.py:262), so the whole check can be done in a single call.
Reproduction
from haystack import Document
from haystack.components.caching.cache_checker import CacheChecker
from haystack.document_stores.in_memory import InMemoryDocumentStore
store = InMemoryDocumentStore()
store.write_documents(
[Document(content=f"d{i}", meta={"url": f"https://example.com/{i}"}) for i in range(200)]
)
calls = 0
orig = store.filter_documents
def counting(*a, **kw):
global calls; calls += 1
return orig(*a, **kw)
store.filter_documents = counting
CacheChecker(store, cache_field="url").run(
items=[f"https://example.com/{i}" for i in range(200)]
)
print("filter_documents calls:", calls) # -> 200
For any non-in-memory store, each call is a separate network round-trip.
Expected behavior
A single filter_documents call should be issued for the whole items list, using the in operator:
filters = {"field": self.cache_field, "operator": "in", "value": items}
found = self.document_store.filter_documents(filters=filters)
misses should be derived from the items whose value did not appear in the result set, preserving the current return contract.
Why this matters
- Cache-checking on URL- or ID-keyed lookups is a common pattern for Web RAG pipelines. Today a 1 000-URL crawl triggers 1 000 round-trips against the document store.
- The proposed change is a pure performance fix: semantics of the return value are preserved (verified locally on
InMemoryDocumentStore: identical hits set and misses list for inputs with duplicates, missing values, and nested cache_field paths).
Proposed fix
@component.output_types(hits=list[Document], misses=list)
def run(self, items: list[Any]) -> dict[str, Any]:
if not items:
return {"hits": [], "misses": []}
filters = {"field": self.cache_field, "operator": "in", "value": items}
found = self.document_store.filter_documents(filters=filters) or []
seen = {self._get_field_value(d) for d in found}
misses = [i for i in items if i not in seen]
return {"hits": list(found), "misses": misses}
(Plus a _get_field_value helper that mirrors the field-path handling in haystack/utils/filters.py:304, and the same change in run_async.)
Acceptance criteria
CacheChecker.run issues exactly one filter_documents call regardless of len(items).
- Returned
hits and misses are identical to the current implementation for: single item, many items, items with duplicates, items all missing, items partially missing, and a nested cache_field like "meta.foo.bar".
run_async mirrors the change and still issues exactly one filter_documents_async call.
- Existing
test_filters_syntax is updated to assert the new single-call filter shape.
- New unit tests cover the per-call count and the miss-derivation logic.
Backward compatibility
- Public API (
run / run_async signatures, return type, return value semantics) is unchanged.
- The new filter shape is observable only through
document_store.filter_documents calls and the test_filters_syntax assertion, both internal.
Risks
- Some document stores may have a practical limit on
in-list size. If a back-end rejects large lists, falling back to the existing per-item loop is a straightforward mitigation.
- The miss-derivation assumes
cache_field values are hashable. The current == filter requires the same, so no new constraint is introduced.
CacheChecker.runissues onefilter_documentscall per item (N+1 query pattern)Bug
CacheChecker.run(haystack/components/caching/cache_checker.py:75) loops over the inputitemslist and callsself.document_store.filter_documents(...)once per item:For an input of
Nitems, this isNround-trips to the document store. ForInMemoryDocumentStorethe constant is small; forQdrantDocumentStore,WeaviateDocumentStore,OpenSearchDocumentStore, etc., every call is a network request. With a 10 ms per-call latency, 200 items takes ~2 s; 5 000 items (a realistic batch) takes ~50 s.The same pattern is duplicated in
CacheChecker.run_async(cache_checker.py:99).The Haystack filter spec already supports the
inoperator (haystack/document_stores/types/protocol.py:70,haystack/utils/filters.py:262), so the whole check can be done in a single call.Reproduction
For any non-in-memory store, each call is a separate network round-trip.
Expected behavior
A single
filter_documentscall should be issued for the wholeitemslist, using theinoperator:missesshould be derived from the items whose value did not appear in the result set, preserving the current return contract.Why this matters
InMemoryDocumentStore: identicalhitsset andmisseslist for inputs with duplicates, missing values, and nestedcache_fieldpaths).Proposed fix
(Plus a
_get_field_valuehelper that mirrors the field-path handling inhaystack/utils/filters.py:304, and the same change inrun_async.)Acceptance criteria
CacheChecker.runissues exactly onefilter_documentscall regardless oflen(items).hitsandmissesare identical to the current implementation for: single item, many items, items with duplicates, items all missing, items partially missing, and a nestedcache_fieldlike"meta.foo.bar".run_asyncmirrors the change and still issues exactly onefilter_documents_asynccall.test_filters_syntaxis updated to assert the new single-call filter shape.Backward compatibility
run/run_asyncsignatures, return type, return value semantics) is unchanged.document_store.filter_documentscalls and thetest_filters_syntaxassertion, both internal.Risks
in-list size. If a back-end rejects large lists, falling back to the existing per-item loop is a straightforward mitigation.cache_fieldvalues are hashable. The current==filter requires the same, so no new constraint is introduced.