Flash indexing spends most of its wall time in summaries, and until now that stage waited for expand to finish and then ran its calls in whatever order the tree recursion produced. This branch makes the summary stage run deepest node first and start while expand is still deciding, so the LLM channels never sit idle waiting on the expand chain.
**What changes**
- `_PriorityGate`: the summary semaphore admits the queued call with the most work still above it (depth = calls left on the node's path to the root, its own included), FIFO within a depth. Cancellation-safe like `asyncio.Semaphore`.
- Tasks are created deepest node first, so the first admissions are the deep leaves rather than whichever shallow leaves the recursion reached first.
- `summarize_tree` becomes a thin wrapper over `SummaryScheduler`: `mark_final(nodes)` says those nodes will not gain, lose or swap children and starts their subtrees; `finish()` awaits the roots. Same task order, gate and error semantics as before.
- `optimize(on_final=...)` reports which nodes are final as it goes: after each round's merges, at each expand candidate's decision (together with what it grew), and for the whole tree at the end. A node is final when it is collapsed under the trigger, collapsed and already judged by expand, or has children — the cost merge cannot fire on a surviving node after the first round (see the commit message for the argument).
- Same-page fusion moves to where duplicates arise (right after a collapsing merge, right after expand attaches children) instead of the next round's start, so no node waits a round for it. The nine corpus PDFs produce byte-identical merge-only trees; SpaceX just stops after two rounds instead of a third that did nothing.
- `page_index_flash` runs expand and summaries on one event loop when both are on; every other combination keeps the old path.
**Measured** (same hour, end to end via `submit_document`)
| | before | after |
|---|---|---|
| fed-2023 (222 p) | 97.9 s | 72.6 s |
| PRML (758 p) | 174.3 s | 136.8 s |
Summary-stage only (fed, 182 calls, 64 wide): FIFO 58–62 s → gate 50–57 s → gate + deepest-first 45 s.
Same calls, same prompts; outputs are order-independent. Peak in flight is now the expand cap plus the summary cap (32 + 64).
**Tests** cover the ordering, cancellation, scheduler, final-node reporting, immediate-fusion and one-loop overlap cases, and every knob's path from the client and the CLI to the model calls.
**Summary prompt and indexing knobs**
The summary prompts no longer ask for the `points` list that `parse_summary` discarded, and cap the summary at `summary_max_words` (default 150). Measured on gpt-5.6-luna, mirror A/B, summary stage only: per-call latency 9.7 → 5.3 s (−45%), fed-2023 47.5 → 30.7 s (−35%), PRML 71.1 → 38.1 s (−46%), output tokens −65%. Summaries come out ~1160 chars instead of ~670 and carry the specifics that used to sit in the discarded list; a blinded pairwise judge (claude-sonnet-5, source in view) prefers them 21-1-0 over the old ones. Deleting the list without a cap is not enough: the model then pours it into the summary (3× longer) and parents slow down more than the leaves gain.
Four indexing knobs are settable from the SDK (flat arguments or the `index=` slot) and the CLI: `summary_max_words`, `summary_concurrency`, `use_embedded_toc`, `optimize` (`"full"` / `"merge"` / `"off"`). `summary_concurrency` bounds both lanes: expand's gate becomes min(32, the cap), so one knob lowers the whole indexing lane on a tight quota (the lanes overlap, so up to cap + min(32, cap) calls run at once). Defaults are unchanged.
The two summary knobs are flash-only: `submit_document(mode="standard")` refuses them rather than index without the cap, as the CLI already does. Both must be positive integers, checked before the PDF is opened; a direct `page_index_flash` call that passed `0` (read as the default until now) or a whole-number float such as `8.0` now raises `ValueError`.
Flash returned an empty structure, and `submit_document(mode="flash")` and the CLI a hard error, for any PDF under 300 text weight, under 200 on its densest page, or with mostly-landscape pages. Both rules threw away documents the detector handles. Four more rules keyed on the document's script: the "other" script family (Arabic, Hebrew, Persian, Urdu, Devanagari, Bengali, Tamil, Thai, Khmer, Georgian, Armenian, Amharic, and numbers-only text) was refused as "no alphabetic text"; an unnumbered heading in a script other than the body's was dropped, so a Chinese report lost its English section titles; a kana-majority Japanese document had every detected heading discarded; a mostly-landscape document picked its title from page one without the body-paragraph check, so a slide deck's title became slide one's body text. These are Scholar's scope limits for an index of Latin and CJK papers; on PageIndex's default local mode they were silent refusals and silent losses.
**What changes**
- Layout decides, never script. The size and landscape bails, the script gate, the cross-script heading drop, the Japanese outline nullifier, the landscape title branch, the Cyrillic-only density threshold and the title scorer's cross-script penalty are deleted from this repo's copy of the port; the private `scholar/` tree stays a faithful port and the new tests guard the fork. Language now only decides which cues are available: case, keyword tables, numbering styles.
- When detection finds no hierarchy, `page_index_flash` returns one node per page titled `Page N`, covering every page, labelled `toc_source="pages"`. A flat tree over `FLAT_TREE_MAX_NODES` (10) pages comes back without the optimize and summary passes and is refused by the local client and the CLI through one shared `flash_rejection_reason()`, pointing at standard mode.
- Every page is in some node. A hierarchy that starts after page 1 (a memo whose first heading became the document title, a title slide, a report's cover and contents, a bookmark outline that begins on page 3) is preceded by a `Preface` node covering the pages before it, the node standard mode has always inserted for the same case; until now those pages were reachable from no node.
- `toc_source="unreadable"` means exactly that no page carries text; the refusal says so and points at OCR, not at standard mode, which would receive the same bytes.
- The character-level parser no longer raises on a glyph whose ToUnicode value is several code points (a Devanagari conjunct, a Thai cluster, an Arabic ligature); real Hindi and Thai PDFs used to fail with a `TypeError` before any rule ran.
- `toc_source` is present on every result: `detected`, `bookmarks`, `hybrid`, `pages`, `unreadable`. The README and the `page_index_flash` docstring list them, and describe a node as emitted: `node_id` on every node, `nodes` only on entries with children, `summary` only when summaries ran.
- `get_leaf_nodes` walks a flat page tree instead of raising `KeyError` on a node without a `nodes` key; it was the one tree helper reading the key unguarded.
**Behaviour change**
Small documents, slide decks, and Japanese, Arabic, Hebrew, Indic, Thai and mixed-script documents that used to fail flash indexing or lose headings now index; with the rules gone the same layout yields the same headings in every one of those scripts, and English is unchanged. A garbage text layer that still has layout structure now indexes as a garbage-titled tree instead of being refused. A Chinese-body report whose cover sets an English title over a Chinese subtitle now picks its title by layout; the deleted penalty could hand `doc_title` to a body paragraph. `extract_toc` yields the same nine example trees, node for node, before and after; `page_index_flash` adds the `Preface` node to the three whose hierarchy starts late (the two Federal Reserve reports, pages 1-4 and 1-2, and Four Lectures, page 1), the node standard mode already gives them, and leaves the other six identical.
**Tests**
Fixtures for Japanese, Chinese with English headings, Hindi and Arabic under `tests/data/flash/`, PyMuPDF-generated with open-licensed font subsets embedded; `make_fixtures.py` regenerates them byte-identically. Green on all three CI legs locally (with and without agent frameworks, pypdfium2 4 and 5).
* fix: flip litellm's banner switch at the provider lookups, not after them
_openai_agent asks litellm which provider serves a model twice
(_openai_protocol, _cache_extra_args) before _openai_model flips
suppress_debug_info, and openai_agent_config's litellm/ branch never
flipped it at all; a managed cloud client runs no preload, so each
failed lookup still print()ed the "Provider List:" banner. Quiet at the
two get_llm_provider call sites instead, which covers every lane.
Also: LITELLM_LOG at ERROR or above now clamps to that level (CRITICAL
used to skip the clamp and end up noisier than unset); the retry notice
is logged only when a retry follows; test_preload_stamps_litellm_log_level
no longer leaks LITELLM_LOG=ERROR into the session; the config-time
repair test covers _quiet_litellm with the preload thread pinned off.
Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9
* refactor: the preload thread only imports — every litellm entry quiets itself now
Since 30652ca each get_llm_provider caller and every completion helper
flips the banner switch and clamps litellm's logger before touching
litellm, so the background import's own _quiet_litellm() had no path
left that depended on it; it was also the one place that mutated
process-global logging at a moment the caller could not predict. The
env stamps stay: import-time records still need them. The config-time
repair test no longer needs the preload guard pinned.
Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9
* fix: litellm quiet is a default, not a clamp — and the no-fetch stamp covers every lane
_quiet_litellm() set litellm's loggers to ERROR on every entry, overriding whatever the host, litellm._turn_on_debug(), or LITELLM_LOG=DEBUG had set. It now sets a level only while a logger is still NOTSET: the default stays ERROR, anything set explicitly wins, and there is no per-call setLevel churn.
LITELLM_LOCAL_MODEL_COST_MAP moves from the client's preload to utils' import. The CLI pipeline and openai_agent_config on a cloud client import litellm without the preload and fetched the remote model map on first use (offline: a WARNING). Every litellm lane imports utils first.
The LITELLM_LOG=ERROR stamp is gone: it pinned litellm's stderr handler at ERROR for the whole process (silencing _turn_on_debug for good) while the logger level was what actually gated the chatter — output verified identical with and without it across plain, logging-configured, and debug hosts.
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* style: essential comments only
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* test: the utils-import stamp check lives with its package-surface sibling
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* refactor: drop the CLI's duplicate no-fetch stamp — utils' import sets it
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* fix: the thinking-budget ceiling imported litellm ahead of the stamp
_default_max_tokens is reached from anthropic_runner_config/messages(). On a managed cloud client the constructor never imports pageindex.utils, and local_chat deliberately keeps utils off its import path, so that lane imported litellm with LITELLM_LOCAL_MODEL_COST_MAP unset: the model cost map was fetched over the network (3517 entries vs the bundled 2982) and offline printed the very warning this branch removes. Verified end to end through the public API, before and after.
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
Indexing: dead credentials or a missing model fail the run instead of
storing a document with blank summaries; a 400 (context_length_exceeded)
skips the retry ladder — the prompt will not shrink — and stays a
per-prompt failure the run absorbs; all-empty model replies can no longer
store a retrieval-ready document; the one-sentence doc description
absorbs its own context overflow instead of discarding a fully indexed
document; the heading-less flash refusal points at mode='standard'.
Chat: messages() output is append-verbatim clean — unset response-only
defaults are dropped (no "caller": null the request schema rejects);
Claude cache marks follow the wire routing; model_settings and name are
openai_agent_config parameters; one Anthropic client per backend; lifted
thinking defaults are clamped to the model's output ceiling from
LiteLLM's capability map.
Store and inputs: lone surrogates are scrubbed from page text and the
stored basename, so the returned name is byte-for-byte the stored name
and the rename warning fires; NaN/Infinity metadata is rejected at the
gate; every cloud error now carries its HTTP status.
CLI: the flash lane resolves the summary model through ConfigLoader like
the standard and markdown lanes; an empty flash structure errors like the
SDK instead of writing "structure": [] with exit 0; --summary-model
reaches the markdown lane; the SDK page-spec surface keeps 0.2.10's
whitespace tolerance while the tool layer stays strict.
pypdfium2 stays on the 5.x line for every install; the 4.x code paths are
tested compatibility insurance with their own CI leg; process-pool
construction failure falls back to the sequential parse; a py3.10 GC
flake in text extraction is fixed.
Port of feat/local-chat 0667e3b..1993740 (28 commits); README and assets untouched.
Squash-port of feat/local-chat's post-dev5 wave (03ffab3..49a24e1): five
review rounds of reproduced-then-fixed findings. Highlights: standalone
agent_instructions back on the strict shadow check (the *_agent_config
bundles keep the relaxed in-set check they can prove); caller-owned
http_client survives the per-call backend closes; get_tree keeps
key_items under the flash merge default; OpenAI-protocol classification
follows litellm's own routing (azure/openrouter/deepseek/groq/xai);
thinking-aware max_tokens default shared by messages() and
anthropic_runner_config(thinking=); 401/403 re-raise instead of
retry-coaching envelopes, bridge errors carry status_code; cloud
discovery + instructions ride the ?tools=read endpoint matching the
tool gate (live-verified); the chat lane's litellm model resolution
deduplicated into utils._litellm_model; explicit optimize= wins over
the deprecated optimize_expand (DeprecationWarning added); mcp 2.0
compatibility; spawn-worker and doc-scope guard fixes; the suite is
.env-independent and pins the LitellmModel._fetch_response seam.
Per-finding rationale in the ported commit messages on feat/local-chat.
* docs: drop the demo's install step — openai-agents ships with the SDK now
* feat: the chat lane routes every model through LiteLLM — bare names included
The direct-OpenAI special case existed to dodge LiteLLM's import cost,
and it made OpenAI's own Responses-first models fail on the front door:
gpt-5.6-sol 400s on chatcmpl+tools while reasoning is on (server-side
policy — wire-captured with no reasoning_effort in our request).
LiteLLM 1.97 translates such calls onto /v1/responses; 1.84 does not,
so the sol-class 400 now carries its two exits (upgrade litellm /
responses()).
Routing after the flip: chat protocol — bare names are OpenAI-compatible
shorthand (wire form openai/<name>; OPENAI_API_KEY / OPENAI_BASE_URL
still select the backend, and the missing key stays a build-time
failure), litellm/ strips, openai/ opts out to the OpenAI SDK directly;
responses protocol unchanged (OpenAI-SDK native, LiteLLM refused).
The import cost is handled instead of dodged: local clients preload
litellm on a background thread (first call then perceives 0.0s), and
pageindex sets LITELLM_LOCAL_MODEL_COST_MAP=True via setdefault —
LiteLLM's import otherwise blocks on a network fetch of its price map
(fresh venv: 5.6s -> 1.3s; offline it hangs to the timeout).
Also restores prompt_cache_key delivery, found dead during the flip's
gating verification: openai-agents 0.20 no longer derives it from
RunConfig.group_id, so both lanes sent nothing. ModelSettings.extra_body
is the one channel all three model classes put on the wire (the bare
kwarg is dropped by LiteLLM; extra_args[extra_body] collides with the
responses model's own parameter — both wire-verified), and it is scoped
to OpenAI destinations: LiteLLM plants extra_body as a literal field in
other providers' bodies, and Anthropic rejects unknown fields — the
anthropic wire test now pins the absence.
Verified before landing: mock-server matrix (OPENAI_BASE_URL + bare
name works through LiteLLM; gpt-named self-hosted models are NOT
bridged off a custom base_url; prompt_cache_key on the wire in every
OpenAI lane with distinct per-conversation keys; anthropic body clean)
and live (sol answers through chat(), gpt-5.4 unchanged, responses()
bare unchanged with the key on its wire).
* refactor: no prefix-triggered direct lane — chat model names are LiteLLM's, verbatim
Ray's ruling on the flip's remaining carve-out: a routing decision must
never hide in a model-name prefix. openai/ now means what LiteLLM says
it means (its openai provider), like every other name on the chat lane —
the grammar is LiteLLM's with zero exceptions.
The two defenses for keeping a direct carve-out had no concrete victim:
debugging isolation (litellm is unavoidable in indexing anyway, and
responses() IS the OpenAI-SDK-native door), and endpoint determinism
(litellm sends chatcmpl for openai-provider models except the gpt-5
bridge, which never fires against a custom base_url — wire-verified).
If a direct escape is ever needed, it will be a declared parameter,
never name grammar.
openai/-prefixed names keep the build-time OPENAI_API_KEY check for
parity with bare names; responses() is untouched (bare and openai/
still drive the OpenAI SDK — LiteLLM cannot speak that protocol).
* fix: third-audit findings — extra_body naming drift, litellm floor hint
The prompt_cache_key delivery channel went through three iterations and
settled on ModelSettings.extra_body; the _conversation_cache_key
docstring still named extra_args from the middle iteration.
The litellm install hint said >=1.30, below both our own pyproject floor
(>=1.84.0) and the floor openai-agents' litellm extra declares (>=1.83).
A user in a broken environment following it would land on a version the
package itself rules out. The hint now matches the declared floor.
* test: pin the bundle door's Agents-SDK model grammar; rename the cache-key test to extra_body
The config bundle hands its model string to the Agents SDK's own
MultiProvider grammar, which refuses unknown prefixes (probe on 0.20:
'anthropic/x' -> UserError: Unknown prefix). _normalize_retrieve_model's
litellm/ spelling is what keeps that door working — a link the existing
self-referential assert (config["model"] == client.retrieve_model)
could not catch. Pinned with a provider-slashed name.
Also renames the cache-key delivery test to its real channel,
extra_body — the extra_args name survived from the superseded delivery
attempt.
* refactor: name the retrieve_model helper for its reason — the Agents SDK's grammar
_normalize_retrieve_model said what it does, not why. The litellm/
spelling exists because the Agents SDK resolves raw model strings with
its own prefix grammar and refuses unknown prefixes — the name now
points at that constraint.
* test: the bundle-grammar test skips without openai-agents, like its file's siblings
Every agents-dependent test in this file importorskips; without the
guard this one errors where the others skip.
* feat: index_model + chat_model — two-knob model surface with full legacy fallback
The documented surface becomes two role knobs: index_model builds the
index, chat_model answers on the chat surfaces. model turns into the
set-both umbrella (its 0.2.8 indexing semantics are a strict subset, so
old configs run unchanged); summary_model and retrieve_model stay
accepted as legacy role names.
Resolution lives in ConfigLoader.load(), the one seam every consumer
already passes through (client, CLI standard/md paths, flash's
summary fallback, tree_optimize's default_model): new names win over
old, specific over general, model sets every role, and code constants
close each chain. The packaged yaml no longer ships model keys — key
presence is what separates a user's explicit choice from a built-in
default, and _validate_keys accepts the five model names explicitly.
Consequences: with no config at all, classic-mode structure extraction
now uses DEFAULT_INDEX_MODEL (gpt-5.6-luna) instead of the yaml's old
gpt-4o-2024-11-20 line (ratified; flash-default users see no change).
client.retrieve_model becomes a read-only alias for client.chat_model.
The resolution matrix test pins one row per released generation:
0.2.8 (model), 0.3.0.dev (model+retrieve_model), 0.2.10.dev (all three
legacy names), the new pair, umbrella-only, and mixed.
* feat: --index-model on the CLI; README flag docs follow
The CLI leads with --index-model; --model stays as its legacy synonym
(the CLI only indexes, so the umbrella and the index role coincide).
The flash branch's summary fallback gains the index position, and the
standard branch now forwards --summary-model, which it had silently
ignored — the flag's help always claimed it worked there. The md
branch's unfiltered model=None no longer clobbers the default: the
resolver treats None as unset.
* feat: default chat model becomes gpt-5.6-sol
Ray's pick for the out-of-box QA default; indexing stays on luna. sol
runs tools on the Responses lane — current litellm bridges chat()
there automatically; older litellm gets the guided 400 naming both
exits.
* chore: trim the model-keys comment to the constraint
The per-generation history lives in 7244ee4's message.
* feat: per-door reasoning passthrough — reasoning_effort / reasoning / thinking
Each chat door gains its own protocol's native thinking control,
forwarded verbatim with no invented vocabulary and no default of ours:
chat_completions(reasoning_effort=...), responses(reasoning={...}),
messages(thinking={...}). Unset sends nothing, so backend defaults
(sol: medium, adaptive) are untouched. chat() stays answer-only.
Delivery channels, each verified: the chat door rides
extra_args["reasoning_effort"] — LiteLLM's own top-level kwarg on
every supported openai-agents version, admitting non-enum values
("none"); newer openai-agents promotes it to the top-level argument
and pops the duplicate. Wire-captured on a mock backend
(/v1/chat/completions body carries it) and coexists with the Claude
cache marker in one dict. The responses door rides
ModelSettings.reasoning — coerced to the typed openai Reasoning object
and forwarded verbatim by the Responses model; the envelope echoes the
caller's dict. The messages door joins the existing anthropic
passthrough dict, asserted through the real tool runner.
LiteLLM semantics observed and accepted as-is: unknown models refuse
the param loudly with LiteLLM's own remedies, and gpt-5.4+ names with
an explicit effort route to /v1/responses even against a custom
api_base (its documented pre-existing arm). The sol-class 400 guidance
now names the third exit — an explicit effort routes on older litellm
releases too.
Cloud chat_completions rejects the new parameter like model/max_turns;
responses()/messages() are local-only already.
* feat: extra_body escape hatch on the three protocol doors
The industry-standard per-request extension channel (openai/anthropic
SDK trio): a dict merged verbatim into the backend request, last, so
caller keys win over SDK-set ones. Routing per door: OpenAI-compatible
destinations get a true body merge (ModelSettings.extra_body / the
anthropic SDK's native extra_body); LiteLLM-routed providers take the
keys as LiteLLM's own top-level kwargs instead, since LiteLLM plants
extra_body as literal fields other providers reject. Cloud mode rejects
it like the other local-only knobs. chat() stays answer-only.
* chore: litellm floor 1.84 -> 1.97.0
1.97.0 is where the unset-effort chatcmpl->responses bridge landed
(responses_api_bridge_check's on_constraint_enforcing_endpoint arm,
A/B-verified against 1.96.2), so sol-class models work through the
chat lane out of the box instead of 400ing until a manual upgrade.
Three spots move together: the pyproject floor, the requirements.txt
CI pin, and the install hint.
* feat: named top_p/max_tokens on the chat door, max_output_tokens on responses
extra_body could not carry these: openai-agents' LitellmModel passes
every ModelSettings sampling field as an explicit keyword and unpacks
extra_args into the same call, so the common knobs collided with a bare
TypeError on the LiteLLM lane (reproduced against a stub — Python call
semantics, callee-independent). Named params ride ModelSettings fields,
the one channel clean on every lane; responses() uses the protocol's
own name (openai_responses maps ModelSettings.max_tokens to
max_output_tokens on the wire) and the envelope now echoes the real
value instead of a constant None. Both caps bound each backend call in
the agent loop, not the whole run — documented. Cloud rejects them like
the other local-only knobs; the long tail (frequency_penalty etc.)
stays extra_body-blocked-loudly on that lane by choice.
* chore: the missing-key error names chat_model as the other exit
The default chat_model is what put keyless users on the OpenAI lane,
so the error now points at the knob that picks a different backend.
* fix: break the phantom exception chain in _run_sync
Move asyncio.run(coro) out of the except RuntimeError block so real
errors no longer carry a bogus "no running event loop" context in
their traceback.
* feat: Flash with full optimization becomes the default local indexing mode
Every entrance now defaults to Flash with the full optimize pass
(deterministic merge, then LLM expand), replacing the standard LLM-built
tree as the default:
- submit_document(): mode=None now means "flash"; pass mode="standard"
for the LLM-built tree. _index_flash runs optimize="full" with the
expand model = summary_model, and fails fast with the missing key
name(s) via litellm.validate_environment before any work.
- page_index_flash(): optimize takes "full" (default) / "merge" / False;
True is accepted as "full" for compatibility, unknown values raise
instead of silently degrading to merge-only. optimize_expand stays
honored for legacy callers.
- CLI: --mode {flash,standard} replaces --flash (kept as a hidden
compatibility alias that forces flash). --optimize defaults to full in
flash mode with an `off` choice; explicitly passing it outside flash
still errors. Standard-only tuning flags (--toc-check-pages,
--max-*-per-node, --if-add-*) now error in flash mode instead of being
silently ignored, mirroring the existing flash-only flag errors. The
key pre-check runs only when an LLM will actually be called, so
--no-summary --optimize off|merge works keyless. Output drops the
_structure_flash suffix — always <name>_structure.json.
On the Disney earnings PDF the optimized default is also faster than
unoptimized flash (fewer nodes to summarize) and fixes hierarchy
mistakes; both modes emit identical schemas end to end.
Docs updated to match (mode flag, defaults, LLM usage honesty); tests
pin the new defaults: stored mode == "flash", optimize passthrough, and
the unknown-optimize rejection.
* Integrate litellm for multi-provider LLM support
* recover the default config yaml
* Use litellm.acompletion for native async support
* fix tob
* Rename llm_complete/allm_complete to llm_completion/llm_acompletion, remove unused llm_complete_stream
* Pin litellm to version 1.82.0
* resolve comments
* args from cli is used to overrides config.yaml
* Fix get_page_tokens hardcoded model default
Pass opt.model to get_page_tokens so tokenization respects the
configured model instead of always using gpt-4o-2024-11-20.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Remove explicit openai dependency from requirements.txt
openai is no longer directly imported; it comes in as a transitive
dependency of litellm. Pinning it explicitly risks version conflicts.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Restore openai==1.101.0 pin in requirements.txt
litellm==1.82.0 and openai-agents have conflicting openai version
requirements, but openai==1.101.0 works at runtime for both.
The pin is necessary to prevent litellm from pulling in openai>=2.x
which would break openai-agents.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Remove explicit openai dependency from requirements.txt
openai is not directly used; it comes in as a transitive dependency
of litellm. No openai-agents in this branch so no pin needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix an litellm error log
* resolve comments
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>