fix: self-calibrate context window from API result event
the stream-json result event carries modelUsage.<model>.contextWindow which is the actual per-inference active window the model enforces. for claude-sonnet-4-6 this is 200k even though the full prompt cache can hold millions of tokens via accumulated cache reads. with the nix-configured sonnet = 1000000 the proactive compact watermark sat at 750k and was never reached. agents grew context until prompt_too_long at ~170k — reactive compact, no checkpoint turn. changes: - bus gains api_context_window field seeded from modelUsage.*.contextWindow in each turn's result event. authoritative; falls back to env var, then 200k. - new effective_context_window(bus) helper used by both watermark functions - compact_watermark (75%) and auto_reset_watermark (50%) call effective_context_window - context_tokens() docstring clarified: all three token fields (input + cache_read + cache_creation) count against the per-inference contextWindow limit. the large cache_read values seen in the result event are cumulative across all inferences in a turn, not per-inference. - /api/state context_window_tokens now reflects the calibrated window closes #129
This commit is contained in:
parent
3e94914569
commit
b0f6bd8ece
3 changed files with 131 additions and 43 deletions
|
|
@ -371,10 +371,10 @@ struct StateSnapshot {
|
|||
/// in flight). Mutable at runtime via `POST /api/model`.
|
||||
model: String,
|
||||
/// Effective context-window token budget for the current model.
|
||||
/// Derived from `events::context_window_tokens(&model)` — respects
|
||||
/// per-model and global `HIVE_CONTEXT_WINDOW_TOKENS_*` overrides then
|
||||
/// falls back to model-family heuristic. Consumers (e.g. dashboard
|
||||
/// badge) use this to render the ctx-usage percentage.
|
||||
/// Primary source: API-reported `modelUsage.*.contextWindow` from
|
||||
/// the last result event (authoritative per-inference active window).
|
||||
/// Falls back to `HIVE_CONTEXT_WINDOW_TOKENS_*` env vars, then 200 000.
|
||||
/// Consumers (e.g. dashboard badge) use this to render ctx-usage %.
|
||||
context_window_tokens: u64,
|
||||
/// Last-inference token usage from the most recent completed
|
||||
/// turn — represents the current context-window size at turn-end.
|
||||
|
|
@ -483,7 +483,10 @@ async fn api_state(headers: HeaderMap, State(state): State<AppState>) -> axum::J
|
|||
let inbox = recent_inbox(&state.socket, state.flavor()).await;
|
||||
let (turn_state, turn_state_since) = state.bus.state_snapshot();
|
||||
let model = state.bus.model();
|
||||
let context_window_tokens = crate::events::context_window_tokens(&model);
|
||||
let context_window_tokens = state
|
||||
.bus
|
||||
.api_context_window()
|
||||
.unwrap_or_else(|| crate::events::context_window_tokens(&model));
|
||||
let ctx_usage = state.bus.last_ctx_usage();
|
||||
let cost_usage = state.bus.last_cost_usage();
|
||||
axum::Json(StateSnapshot {
|
||||
|
|
|
|||
Loading…
Reference in a new issue