Files
PageIndex/tests/test_package_surface.py
T
Ray 0400366d35 litellm terminal noise: quiet as a default, and the no-fetch stamp on every lane (#473)
* fix: flip litellm's banner switch at the provider lookups, not after them

_openai_agent asks litellm which provider serves a model twice
(_openai_protocol, _cache_extra_args) before _openai_model flips
suppress_debug_info, and openai_agent_config's litellm/ branch never
flipped it at all; a managed cloud client runs no preload, so each
failed lookup still print()ed the "Provider List:" banner. Quiet at the
two get_llm_provider call sites instead, which covers every lane.

Also: LITELLM_LOG at ERROR or above now clamps to that level (CRITICAL
used to skip the clamp and end up noisier than unset); the retry notice
is logged only when a retry follows; test_preload_stamps_litellm_log_level
no longer leaks LITELLM_LOG=ERROR into the session; the config-time
repair test covers _quiet_litellm with the preload thread pinned off.

Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9

* refactor: the preload thread only imports — every litellm entry quiets itself now

Since 30652ca each get_llm_provider caller and every completion helper
flips the banner switch and clamps litellm's logger before touching
litellm, so the background import's own _quiet_litellm() had no path
left that depended on it; it was also the one place that mutated
process-global logging at a moment the caller could not predict. The
env stamps stay: import-time records still need them. The config-time
repair test no longer needs the preload guard pinned.

Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9

* fix: litellm quiet is a default, not a clamp — and the no-fetch stamp covers every lane

_quiet_litellm() set litellm's loggers to ERROR on every entry, overriding whatever the host, litellm._turn_on_debug(), or LITELLM_LOG=DEBUG had set. It now sets a level only while a logger is still NOTSET: the default stays ERROR, anything set explicitly wins, and there is no per-call setLevel churn.

LITELLM_LOCAL_MODEL_COST_MAP moves from the client's preload to utils' import. The CLI pipeline and openai_agent_config on a cloud client import litellm without the preload and fetched the remote model map on first use (offline: a WARNING). Every litellm lane imports utils first.

The LITELLM_LOG=ERROR stamp is gone: it pinned litellm's stderr handler at ERROR for the whole process (silencing _turn_on_debug for good) while the logger level was what actually gated the chatter — output verified identical with and without it across plain, logging-configured, and debug hosts.

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* style: essential comments only

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* test: the utils-import stamp check lives with its package-surface sibling

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* refactor: drop the CLI's duplicate no-fetch stamp — utils' import sets it

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* fix: the thinking-budget ceiling imported litellm ahead of the stamp

_default_max_tokens is reached from anthropic_runner_config/messages(). On a managed cloud client the constructor never imports pageindex.utils, and local_chat deliberately keeps utils off its import path, so that lane imported litellm with LITELLM_LOCAL_MODEL_COST_MAP unset: the model cost map was fetched over the network (3517 entries vs the bundled 2982) and offline printed the very warning this branch removes. Verified end to end through the public API, before and after.

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
2026-09-03 18:37:11 +08:00

171 lines
7.1 KiB
Python

"""What `pip install pageindex` exposes: 0.2.8 helper compat and import cost."""
import os
import subprocess
import sys
from pageindex.utils import create_node_mapping, print_tree, remove_fields
TREE = [
{"title": "Root", "node_id": "0000", "page_index": 1,
"text": "root text",
"nodes": [
{"title": "Child", "node_id": "0001", "page_index": 3,
"text": "child text"},
]},
{"title": "Tail", "node_id": "0002", "page_index": 5, "text": "tail text"},
]
# ── the published 0.2.8 pageindex.utils surface, as the cookbooks call it ──
def test_remove_fields_max_len():
out = remove_fields({"keep": "x" * 50, "text": "gone"}, max_len=10)
assert out == {"keep": "x" * 10 + "..."}
assert remove_fields({"keep": "short"}, max_len=10) == {"keep": "short"}
def test_create_node_mapping_flat():
mapping = create_node_mapping(TREE)
assert set(mapping) == {"0000", "0001", "0002"}
assert mapping["0001"]["title"] == "Child"
def test_create_node_mapping_page_ranges():
mapping = create_node_mapping(TREE, include_page_ranges=True, max_page=9)
assert mapping["0000"] == {"node": TREE[0], "start_index": 1, "end_index": 3}
assert mapping["0001"]["start_index"] == 3
assert mapping["0001"]["end_index"] == 5
assert mapping["0002"] == {"node": TREE[1], "start_index": 5, "end_index": 9}
def test_print_tree_exclude_fields(capsys):
print_tree(TREE, exclude_fields=["text"])
out = capsys.readouterr().out
assert "Root" in out and "'text'" not in out
print_tree(TREE)
assert "[0000] Root" in capsys.readouterr().out
# ── import cost: the SDK must not pay for the indexing stack ──
def test_import_pageindex_is_lazy():
probe = (
"import sys; import pageindex; "
"heavy = [m for m in ('pageindex.page_index_classic', 'pageindex.flash', "
"'pageindex.utils', 'pageindex.tree_optimize', "
"'pageindex.local_chat', 'numpy', 'PyPDF2', "
"'agents', 'litellm', 'openai', 'anthropic') if m in sys.modules]; "
"print(','.join(heavy) or 'clean'); "
"print(type(pageindex.page_index_main).__name__)"
)
out = subprocess.run([sys.executable, "-c", probe],
capture_output=True, text=True, check=True)
assert out.stdout.split() == ["clean", "function"]
def test_public_method_type_hints_resolve_at_runtime():
"""Tools that introspect signatures at runtime (agents' function_tool,
pydantic, doc generators) evaluate the annotations: every public
method's hints must resolve, ChatStream included."""
import inspect
import typing
import pageindex
from pageindex import ChatStream, PageIndexClient
hints = {name: typing.get_type_hints(fn) for name, fn
in inspect.getmembers(PageIndexClient, inspect.isfunction)
if not name.startswith("_")}
assert len(hints) > 10, f"public-method walk collapsed: {sorted(hints)}"
assert ChatStream in typing.get_args(hints["chat"]["return"])
assert pageindex.local_chat.ChatStream is ChatStream, (
"the import path the class shipped under in 0.2.11-0.2.14")
def test_sdk_submodules_reachable_and_dunder_probes_stay_lazy():
"""The 0.2.10 modules resolve as attributes, and underscore probes (the
frequent unknown names: copy/pickle/inspect dunders) raise without
dragging in the indexing stack. A non-underscore unknown name still
raises AttributeError — after the compat fallthrough's one classic
import, which is the pre-0.2.10 behavior."""
probe = (
"import sys, pageindex\n"
"pageindex.agent_tools; pageindex.local_chat\n"
"pageindex.mcp_bridge; pageindex.integrations\n"
"assert not hasattr(pageindex, '__wrapped__')\n"
"heavy = [m for m in ('pageindex.page_index_classic', "
"'pageindex.flash', 'pageindex.utils') if m in sys.modules]\n"
"print(','.join(heavy) or 'clean')\n"
"try:\n"
" pageindex.definitely_missing\n"
" raise SystemExit('no AttributeError')\n"
"except AttributeError:\n"
" pass\n"
)
out = subprocess.run([sys.executable, "-c", probe],
capture_output=True, text=True, check=True)
assert out.stdout.strip() == "clean"
def test_classic_compat_surface_still_reachable():
"""The pre-0.2.10 catch-all made every classic/utils public name a
package attribute; dropping it broke `from pageindex import
ConfigLoader` on upgrade with no deprecation path."""
probe = (
"import pageindex\n"
"assert callable(pageindex.count_tokens)\n"
"assert isinstance(pageindex.ConfigLoader, type)\n"
"from pageindex import check_toc # noqa: F401\n"
"print('ok')\n"
)
out = subprocess.run([sys.executable, "-c", probe],
capture_output=True, text=True, check=True)
assert out.stdout.strip() == "ok"
def test_import_leaves_litellm_env_untouched(tmp_path):
"""Importing the package must not configure litellm for the host
process; constructing a local client (which will use litellm) does."""
env = {k: v for k, v in os.environ.items()
if k != "LITELLM_LOCAL_MODEL_COST_MAP"}
probe = (
"import os, pageindex\n"
"assert 'LITELLM_LOCAL_MODEL_COST_MAP' not in os.environ, "
"'stamped at import'\n"
f"pageindex.PageIndexLocalClient(storage_path={str(tmp_path / 's')!r})\n"
"assert os.environ['LITELLM_LOCAL_MODEL_COST_MAP'] == 'True'\n"
"print('ok')\n"
)
out = subprocess.run([sys.executable, "-c", probe], env=env,
capture_output=True, text=True, check=True)
assert out.stdout.strip() == "ok"
def test_chat_module_stamps_before_it_imports_litellm():
"""local_chat imports litellm without utils on its import path; the
stamp has to be in place by then anyway."""
probe = ("import os\n"
"from pageindex import local_chat\n"
"local_chat._default_max_tokens('claude-sonnet-4-5',"
" {'budget_tokens': 4096})\n"
"print(os.environ.get('LITELLM_LOCAL_MODEL_COST_MAP'))\n")
env = {k: v for k, v in os.environ.items()
if k != "LITELLM_LOCAL_MODEL_COST_MAP"}
out = subprocess.run([sys.executable, "-c", probe], env=env,
capture_output=True, text=True, check=True)
assert out.stdout.strip() == "True"
def test_utils_import_keeps_litellm_off_the_network():
"""utils' import sets litellm's no-fetch default; an explicit choice wins."""
probe = ("import os, pageindex.utils; "
"print(os.environ['LITELLM_LOCAL_MODEL_COST_MAP'])")
env = {k: v for k, v in os.environ.items()
if k != "LITELLM_LOCAL_MODEL_COST_MAP"}
fresh = subprocess.run([sys.executable, "-c", probe], env=env,
capture_output=True, text=True, check=True)
assert fresh.stdout.strip() == "True"
env["LITELLM_LOCAL_MODEL_COST_MAP"] = "False"
chosen = subprocess.run([sys.executable, "-c", probe], env=env,
capture_output=True, text=True, check=True)
assert chosen.stdout.strip() == "False"