docs: doc_id is per-call table-setting — keep it identical across a conversation

The targeting block doc_id adds is re-set on every call and sits in the
cached prompt prefix, so a round-trip that drops (or changes) doc_id
silently diverges the prefix and loses the cache continuation. State the
rule on all three chat surfaces' doc_id docs, and pin it with a prefix
test that passes the same doc_id on both calls.
This commit is contained in:
Ray
2026-08-12 04:15:44 +08:00
parent 02022df40f
commit adb2f1fddd
2 changed files with 28 additions and 0 deletions
+9
View File
@@ -368,6 +368,9 @@ class PageIndexClient:
is appended to the managed system prompt.
stream: Enable streaming responses.
doc_id: Document ID or list of IDs to scope the conversation.
Keep it identical across a conversation's calls — the
targeting block it adds is re-set each call and is part
of the cached prompt prefix.
temperature: Sampling temperature, passed through to the model.
stream_metadata: With stream=True, yield chunk dicts instead of
text pieces.
@@ -438,6 +441,9 @@ class PageIndexClient:
tool outputs are emitted as ``response.output_item.done``
events and the single final event is ``response.completed``.
doc_id: Document ID or list of IDs to scope the conversation.
Keep it identical across a conversation's calls — the
targeting block it adds is re-set each call and is part
of the cached prompt prefix.
instructions: Appended to the managed system prompt.
temperature / top_p: Passed through to the model.
max_turns: Cap on agent turns per call.
@@ -489,6 +495,9 @@ class PageIndexClient:
(its native event objects, including SDK-synthesized
convenience events), one message sequence per turn.
doc_id: Document ID or list of IDs to scope the conversation.
Keep it identical across a conversation's calls — the
targeting block it adds is re-set each call and is part
of the cached prompt prefix.
system: Appended after the managed system blocks.
temperature / top_p / top_k / stop_sequences: Passed through.
max_turns: Cap on agent turns per call (default 10, like the
+19
View File
@@ -302,6 +302,25 @@ def test_responses_round_trip_extends_prefix(client, store_path, fake_model):
assert second.inputs[0][:len(previous_final)] == previous_final
def test_responses_round_trip_prefix_with_doc_id(client, store_path, fake_model):
"""Same contract with doc targeting: re-passing the same doc_id re-sets
an identical leading block, so the prefix still extends item-for-item."""
seed_doc(store_path, "pi-a", "report.pdf")
first = fake_model([
[_call_item("get_document", {"doc_name": "report.pdf"})],
[_msg_item("The answer")],
])
result = client.responses("What status?", doc_id="pi-a")
second = fake_model([[_msg_item("Done")]])
follow_up = ([{"role": "user", "content": "What status?"}]
+ result["output"]
+ [{"role": "user", "content": "and now?"}])
client.responses(follow_up, doc_id="pi-a")
previous_final = first.inputs[-1]
assert second.inputs[0][:len(previous_final)] == previous_final
@needs_agents
def test_responses_stream_passthrough(client, store_path, fake_model):
seed_doc(store_path, "pi-a", "report.pdf")