Describe the bug
LLMMetadataExtractor supports retrying documents returned in failed_documents. When such a retry successfully parses an empty JSON object ({}), the document is returned in documents but still retains the failure-only metadata keys metadata_extraction_error and metadata_extraction_response from the prior attempt.
The empty object is a valid JSON object when expected_keys is not configured. It can represent a successful extraction with no metadata to add.
Error message
No exception is raised. The successful retry returns a document with stale failure metadata:
{
"metadata_extraction_error": "Response is not valid JSON or missing keys. ...",
"metadata_extraction_response": ChatMessage(...),
}
Expected behavior
Once a retry successfully parses its response, the resulting document should not retain metadata_extraction_error or metadata_extraction_response, including when the parsed JSON object is empty.
Additional context
_parse_dict_from_json accepts {} when expected_keys is empty. In _process_results, stale failure fields are cleared inside the loop over parsed metadata keys. That loop does not run for an empty object. Both run() and run_async() use _process_results, so both paths are affected.
For comparison, a successful non-empty JSON response clears both fields as expected.
To Reproduce
from haystack import Document
from haystack.components.extractors.llm_metadata_extractor import LLMMetadataExtractor
from haystack.components.generators.chat import MockChatGenerator
extractor = LLMMetadataExtractor(
prompt="Extract {{ document.content }}",
chat_generator=MockChatGenerator(responses=["not json", "{}"])
)
first = extractor.run(documents=[Document(content="a")])
assert len(first["failed_documents"]) == 1
second = extractor.run(documents=first["failed_documents"])
retried_document = second["documents"][0]
assert "metadata_extraction_error" not in retried_document.meta # fails
assert "metadata_extraction_response" not in retried_document.meta # fails
The same failed-attempt followed by {} retry reproduces through await extractor.run_async(...).
FAQ Check
System:
- OS: Linux 5.15.0-117-generic x86_64
- Haystack version:
main at 99f2249a6 (3.2.0-rc0)
- Python: 3.12.13
Describe the bug
LLMMetadataExtractorsupports retrying documents returned infailed_documents. When such a retry successfully parses an empty JSON object ({}), the document is returned indocumentsbut still retains the failure-only metadata keysmetadata_extraction_errorandmetadata_extraction_responsefrom the prior attempt.The empty object is a valid JSON object when
expected_keysis not configured. It can represent a successful extraction with no metadata to add.Error message
No exception is raised. The successful retry returns a document with stale failure metadata:
{ "metadata_extraction_error": "Response is not valid JSON or missing keys. ...", "metadata_extraction_response": ChatMessage(...), }Expected behavior
Once a retry successfully parses its response, the resulting document should not retain
metadata_extraction_errorormetadata_extraction_response, including when the parsed JSON object is empty.Additional context
_parse_dict_from_jsonaccepts{}whenexpected_keysis empty. In_process_results, stale failure fields are cleared inside the loop over parsed metadata keys. That loop does not run for an empty object. Bothrun()andrun_async()use_process_results, so both paths are affected.For comparison, a successful non-empty JSON response clears both fields as expected.
To Reproduce
The same failed-attempt followed by
{}retry reproduces throughawait extractor.run_async(...).FAQ Check
System:
mainat99f2249a6(3.2.0-rc0)