main
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bc3dc80200 |
perf(claude-native): Improve performance of claude native terminal typing, text streaming, etc. (#4582)
* perf(claude-native): stop taxing every hook spawn with the eager package init Claude Code blocks its TUI on command hooks — once per streamed text chunk (MessageDisplay), per statusline refresh, and per tool call — and every 'python -m omnigent.<hook>' subprocess re-ran omnigent/__init__, which eagerly imported the datamodel/executor/model-catalog graph. The deliberately stdlib-only hot-path hooks paid ~250 ms per spawn for imports they never use, capping visible streaming at ~4 chunks/s. The package init now re-exports lazily (PEP 562): the FIPS md5 patch and legacy-env mirror stay eager, every public name resolves on first attribute access (optional executors keep their import-failure->None contract), and submodule attribute access still works. Hot-path hook spawns drop to ~30 ms (~interpreter cost). A native_hook_spawn benchmark journey spawns the MessageDisplay hook exactly as Claude Code does and rides the release/nightly regression comparison; fresh-interpreter import-graph guards in the display-hook test suite pin what each hook entrypoint may import so the regression cannot silently return. Signed-off-by: dbczumar <corey.zumar@databricks.com> * perf(claude-native): keep the hook's hot path off the bridge's heavy imports The observer hook — Claude blocks on it at every prompt submit, tool call, Stop, and task event — imported claude_native_bridge, whose module-level tools/spec/pydantic imports cost ~450 ms of interpreter startup, plus httpx and the policy machinery besides. Enter and every tool call paid roughly a second of subprocess overhead per event even after the package init went lazy. The bridge now defers its tools graph to the one launch-path function that builds MCP tools (_build_tools) and its bundle-skills parse to the launch args builder; the hook imports httpx and the policy machinery inside the subcommands that actually speak HTTP. Module import cost: bridge 450 -> ~70 ms, hook 360 -> ~70 ms, and the hook's fresh-interpreter import graph now contains no third-party modules at all — the import guard pins the allowance at exactly that. Tests that reached httpx or create_os_environment through the hook's or bridge's module attributes now patch the owning modules directly. Signed-off-by: dbczumar <corey.zumar@databricks.com> * perf(claude-native): cache the ungoverned policy verdict at the relay Sessions with no policies at all still paid a full server round trip (~0.5-1.3s measured against a Databricks App) on every policy hook event — twice per tool call plus every prompt submit — with the server answering the same fast-path ALLOW each time. Typing during agentic turns stuttered in the gaps; vanilla Claude pays nothing there. The evaluate endpoint now stamps 'governed': false on its existing no-policies fast path (any_policies_apply's False is session-scoped — its only phase-scoped rule forces True), and the native-harness loopback relay caches that verdict for 30s, answering hook events instantly. A governed response of any kind drops the cache, a sys_add_policy call through the relay's own /tool path clears it before the policy lands, and expiry re-validates upstream — so enforcement for governed sessions is untouched and the attach delay for out-of-band policy edits is bounded at the TTL. Signed-off-by: dbczumar <corey.zumar@databricks.com> * perf(claude-native): keep blocking Claude hooks off Python and off the WAN Claude blocks its TUI on every command hook, and three of them still spawned a Python interpreter per event (~30ms floor, ~77ms under EDR): MessageDisplay once per streamed chunk, statusLine per refresh, and evaluate-policy twice per tool call — the last one also paying a 0.5-1.3s WAN round trip whenever its 30s ungoverned-cache window lapsed. - MessageDisplay: a /bin/sh one-liner appends the payload (newline- stripped, so any valid JSON lands single-line) straight to message_deltas.jsonl; the deltas reader already parses by key and skips malformed lines. - statusLine: the shim captures raw stdin to context_raw.json (atomic rename) and chains the user's own status command; the forwarder normalizes it into context.json on its poll loop (sync_raw_status_context), so the Python normalizer leaves the blocking path. The module entrypoint stays for older bridge dirs. - evaluate-policy: hooks try a curl against the relay's new /hook/claude/evaluate-policy endpoint (advertised via a shell-sourceable tool_relay.env); the long-lived runner process owns payload→EvaluationRequest mapping, retries, the ungoverned cache, and verdict→hook-output shaping. When the relay is absent or unreachable the same stdin replays into the Python hook, which keeps the direct-server path and the phase-aware fail-closed contract — exactly the pre-curl behavior. - The relay starts at session create (runner app) instead of at the first web-dispatched turn, so prompts typed directly in the TUI get the curl fast path too; it comes up in the background, and hooks that beat it use the Python fallback. Typing during a live 25-tool-call turn against a Databricks App measured 56.0ms median / 57.2ms p90 / 0 samples over 200ms, from 118ms median / 264ms p90 / 8 freezes before this branch. Also pins the relay-close ownership test's trusted-parent monkeypatch to tempfile.gettempdir() — the literal /tmp never contains the macOS fixture root, so the test only passed on Linux. Signed-off-by: dbczumar <corey.zumar@databricks.com> * perf(onboarding): cache harness CLI version and login probes Every readiness refresh on every host daemon execs vendor CLIs (--version / auth status) whose answers change only when the binary is swapped or a login flips; with a few dozen idle hosts that compounds into a constant machine-wide subprocess storm (~116 spawns/min observed) that competes with interactive terminals. --version output is a pure function of the binary bytes, so successful parses cache permanently against the binary's (path, mtime_ns, size) signature; failures keep re-probing. Login verdicts can flip without a binary change, so only positives cache, with a 120s TTL — negatives always re-probe so the setup wizard sees a fresh login immediately, and harness_logout invalidates its key so a successful logout is confirmed live. Signed-off-by: dbczumar <corey.zumar@databricks.com> * revert(claude-native): drop the ungoverned-verdict relay cache The cache required stamping 'governed': false on the evaluate response so the relay could tell which ALLOWs were safe to reuse — new response-field surface carried only by this optimization, which we don't need right now. Remove the stamp and the relay cache wholesale: every policy hook event consults the server again, the relay's /policies/evaluate proxy is a plain pass-through, and the evaluate response is byte-identical to its pre-branch shape. The sh-shim/curl hook path (no interpreter spawns) is unchanged. Signed-off-by: dbczumar <corey.zumar@databricks.com> --------- Signed-off-by: dbczumar <corey.zumar@databricks.com> |
||
|
|
b3fb64710c | Welcome to omnigent |