api/proxy endpoints (OpenRouter, other OpenAI-compatible aggregators) short-circuit _query_context_length: they only consult the static KNOWN_CONTEXT_WINDOWS table and otherwise return DEFAULT_CONTEXT (128000). Any model not in that table — e.g. a freshly listed OpenRouter model like Owl-alpha — was therefore capped at 128k even though the endpoint's catalog reports its true window (1048576), so the rest of the model context never got used. The short-circuit exists so a context lookup doesn't download a large proxy catalog on every call. Keep that property for the common case: known models still resolve from the table with no network. For a model missing from the table, read the window from the endpoint's /models catalog and cache the whole id->context map per endpoint, so the catalog is fetched at most once per endpoint (not once per model) and only for models that were broken anyway. On any fetch/parse failure or a model absent from the catalog, fall back to DEFAULT_CONTEXT exactly as before. Factor the per-entry field extraction the non-proxy path already used into _model_ctx_from_entry so both paths share it. Fixes #4886
11 KiB
11 KiB