mirror of
https://github.com/VectifyAI/PageIndex.git
synced 2026-10-02 07:44:37 +08:00
docs: doc_id is per-call table-setting — keep it identical across a conversation
The targeting block doc_id adds is re-set on every call and sits in the cached prompt prefix, so a round-trip that drops (or changes) doc_id silently diverges the prefix and loses the cache continuation. State the rule on all three chat surfaces' doc_id docs, and pin it with a prefix test that passes the same doc_id on both calls.
This commit is contained in:
@@ -368,6 +368,9 @@ class PageIndexClient:
|
||||
is appended to the managed system prompt.
|
||||
stream: Enable streaming responses.
|
||||
doc_id: Document ID or list of IDs to scope the conversation.
|
||||
Keep it identical across a conversation's calls — the
|
||||
targeting block it adds is re-set each call and is part
|
||||
of the cached prompt prefix.
|
||||
temperature: Sampling temperature, passed through to the model.
|
||||
stream_metadata: With stream=True, yield chunk dicts instead of
|
||||
text pieces.
|
||||
@@ -438,6 +441,9 @@ class PageIndexClient:
|
||||
tool outputs are emitted as ``response.output_item.done``
|
||||
events and the single final event is ``response.completed``.
|
||||
doc_id: Document ID or list of IDs to scope the conversation.
|
||||
Keep it identical across a conversation's calls — the
|
||||
targeting block it adds is re-set each call and is part
|
||||
of the cached prompt prefix.
|
||||
instructions: Appended to the managed system prompt.
|
||||
temperature / top_p: Passed through to the model.
|
||||
max_turns: Cap on agent turns per call.
|
||||
@@ -489,6 +495,9 @@ class PageIndexClient:
|
||||
(its native event objects, including SDK-synthesized
|
||||
convenience events), one message sequence per turn.
|
||||
doc_id: Document ID or list of IDs to scope the conversation.
|
||||
Keep it identical across a conversation's calls — the
|
||||
targeting block it adds is re-set each call and is part
|
||||
of the cached prompt prefix.
|
||||
system: Appended after the managed system blocks.
|
||||
temperature / top_p / top_k / stop_sequences: Passed through.
|
||||
max_turns: Cap on agent turns per call (default 10, like the
|
||||
|
||||
@@ -302,6 +302,25 @@ def test_responses_round_trip_extends_prefix(client, store_path, fake_model):
|
||||
assert second.inputs[0][:len(previous_final)] == previous_final
|
||||
|
||||
|
||||
def test_responses_round_trip_prefix_with_doc_id(client, store_path, fake_model):
|
||||
"""Same contract with doc targeting: re-passing the same doc_id re-sets
|
||||
an identical leading block, so the prefix still extends item-for-item."""
|
||||
seed_doc(store_path, "pi-a", "report.pdf")
|
||||
first = fake_model([
|
||||
[_call_item("get_document", {"doc_name": "report.pdf"})],
|
||||
[_msg_item("The answer")],
|
||||
])
|
||||
result = client.responses("What status?", doc_id="pi-a")
|
||||
|
||||
second = fake_model([[_msg_item("Done")]])
|
||||
follow_up = ([{"role": "user", "content": "What status?"}]
|
||||
+ result["output"]
|
||||
+ [{"role": "user", "content": "and now?"}])
|
||||
client.responses(follow_up, doc_id="pi-a")
|
||||
previous_final = first.inputs[-1]
|
||||
assert second.inputs[0][:len(previous_final)] == previous_final
|
||||
|
||||
|
||||
@needs_agents
|
||||
def test_responses_stream_passthrough(client, store_path, fake_model):
|
||||
seed_doc(store_path, "pi-a", "report.pdf")
|
||||
|
||||
Reference in New Issue
Block a user