The 10s online-poll budget flakes when a loaded CI worker starves the runner
process. Hard cap only, not a behavior assertion: the loop exits the moment
the runner reports online, so only starved workers ever use the tail.
The interrupt-forward test this PR originally also touched was fixed better
in #2232 (direct awaits under pytest's global timeout); that hunk is dropped.
Signed-off-by: dosenr <robert.dosen@gmail.com>
* fix(ui): prioritize sessionModelOverride in AgentPicker display
* test(ui): cover session model override picker priority
* style(ui): format model picker e2e test
* fix(ui): preserve vendor model picker selection
---------
Co-authored-by: Pat Sukprasert <pattara.sk127@gmail.com>
Reinstall the bundled Python client and UI SDK non-editably in the host image so Landlock-sandboxed imports do not resolve through /build. Keep the existing root package reinstall and add a build-time check that .pth/.egg-link files no longer reference /build.
Co-authored-by: omnigent <noreply@omnigent.ai>
_fetch_search_snippets filtered and joined on conversation_id + position
but omitted workspace_id — the leading column of the only covering index
(workspace_id, conversation_id, position). Without it Postgres can't use
the index and full-scans every conversation_item to fetch the 20 snippet
bodies for a search page, so the snippet fetch alone roughly doubled
search latency and grew with total corpus size.
Add workspace_id to both the MIN(position) aggregate and the join-back so
both stay on the composite index. On a 5k-session / 1M-item Postgres
corpus this drops the snippet query from ~430-680ms (Seq Scan) to ~7ms
(Index Scan), and the search_sessions benchmark P50 from ~571ms to
~315ms. No behavior change — same rows, same earliest-match snippet.
Co-authored-by: Isaac
* fix(web): surface server error message in stop-session dialog
The stop-session dialog previously showed a hardcoded message on
failure. Now it displays the actual error from the API response
(e.g. "503 Service Unavailable") so users can diagnose the issue
without opening developer tools.
* fix(web): select-all only selects sessions in expanded sidebar sections
Previously, "Select all" in bulk-selection mode selected every loaded
session including archived and collapsed ones. Now it respects section
collapse state, matching the visible rows.
* fix(web): lift visibleConversations to Sidebar via ref getter
visibleConversations was defined inside ConversationList but referenced
in the parent Sidebar component, causing a ReferenceError at runtime.
Use the same ref-getter pattern as getVisibleIdsRef so the child
populates the getter and the parent calls it on demand.
A full-matrix native run spent minutes in dead waits: a broken vendor forwarder
burned the full 90s _FORWARDER_READY budget before SKIPping (kimi/hermes), and a
model that stalled a turn burned the full 180s _TURN/_TOOL budget. These are
"clearly stuck" ceilings, not expected durations — provisioning is local
(server/runner/host/forwarder boot, no model call) and a healthy native turn
streams within seconds, so a run that blows them is a cold-start on a slow CLI
or a connection/network problem, not normal latency.
Halve them, keeping cold-start headroom:
- _TURN_TIMEOUT_S / _TOOL_TURN_TIMEOUT_S 180 -> 60
- _FORWARDER_READY_TIMEOUT_S 90 -> 45 (and the terminal-ensure HTTP timeout now
references it instead of a separate hardcoded 90)
- _HEALTH_TIMEOUT_S 90 -> 45 (native + full_server)
- _HOST_ONLINE_TIMEOUT_S 45 -> 30
- _DENY_OBSERVE_S 30 -> 15 (post-tool-call grace window for policy_denied)
Worst case for a broken harness drops from ~90-180s to ~45-60s per stall; a
whole-harness provisioning failure now fails in ~45s instead of 90s. Healthy
runs are unaffected (they finish well under the new ceilings). Live gated
full-server tests keep their explicit timeout=180 (real gateway turns).
114 passed / 18 skipped; ruff clean.
Co-authored-by: Isaac
* feat(policies): show model checkboxes for expensive_models in policy dialogs
The expensive_models field in cost-budget policies was a free-text input
requiring users to type comma-separated model tokens. Populate it with
checkboxes from the existing model lists (CLAUDE_NATIVE_MODELS and
session-scoped codexModelOptions) so users can select models visually.
* style: fix prettier formatting in PoliciesPage
* fix: widen modelIds type to satisfy strict const array check
* fix: add missing useMemo import and type annotations in AgentInfo
* feat(policies): replace model checkboxes with dropdown + free-form input
Address reviewer feedback: show known models in a dropdown for quick
selection while also providing a free-form text input for adding custom
model IDs not in the predefined list. Selected values appear as
removable tags.
* feat(policies): themed multi-select combobox for model array params
Replace the native <select> + separate free-text box for array params
(e.g. expensive_models) with a single themed combobox. Users type a
free-form value or pick from a dropdown of existing models; selected
values show a checkmark and toggle on click, and render as removable
chips. The dropdown renders in normal flow inside the dialog so it
scrolls with the modal instead of overlapping the buttons or being
clipped.
The form still stores a comma-joined string and coerces to list[str]
on submit, so the wire format and free-form entry are unchanged.
Add tests covering the combobox in isolation and end-to-end through
both the per-session and global add-policy dialogs, guarding the
coerced list[str] payload against regression.
Co-authored-by: Isaac
---------
Co-authored-by: Serena Ruan <serena.rxy@gmail.com>
* feat(search): show matched-content preview in session search
Session search already matched on title OR conversation item content,
but GET /v1/sessions returned only session rows, so the command palette
could show only the title — a content match was invisible ("why did this
match?"). Surface a short excerpt of the matching chat text so the UI can
show *where* a session matched.
- build_search_snippet (db/utils): windows ~60 chars around the first
match, collapses whitespace, elides ends with "…"; never clamps the
match term out of the window.
- Conversation gains a transient search_snippet (never persisted).
- list_conversations, on a content search, bulk-builds one snippet per
matched conversation via a MIN(position) subquery join (earliest turn
wins; one row per conversation, no N+1). Title-only matches stay None.
- SessionListItem.search_snippet + populated in the shared list builder;
exclude_none keeps it off the wire for title-only matches.
- Command palette renders the snippet as a dimmed second line and bolds
the query term (regex-escaped) in both title and snippet.
Co-authored-by: Isaac
* fix(search): keep the palette match preview from flickering on stream ticks
search_snippet is a search-only field — only GET /v1/sessions?search_query=
computes it. But the WS /v1/sessions/updates stream patches the same cached
rows, and its dump had no query in flight, so it emitted search_snippet: null
and clobbered the snippet the search response had put in the cache. The preview
then vanished on the next stream tick (~60s or any session change), which is
why the highlight showed up only sometimes.
Exclude search_snippet from the watched-items dump so the key is absent from
the frame: the cache merge then leaves the cached snippet untouched. The GET
search path is unchanged (still emits it via exclude_none).
Co-authored-by: Isaac
The org requires all GitHub Actions to be pinned to a full-length commit
SHA; actions/checkout@v4 and actions/setup-python@v5 were rejected at
run time. Pin both to the same SHAs the repo's other workflows use.
Co-authored-by: Isaac
* feat(ci): add Discord watch rotation Slack reminder
Add a deterministic daily on-call reminder that pings the person on
Discord-watch duty in Slack at 08:00 their local time. A hosted GitHub
Actions cron runs the script; whose turn it is is a pure function of the
date, so there is no state to store.
- Weekday-only rotation that advances by workdays (Fri hands off to Mon).
- Per-person timezone: SF folks pinged at 8am PT, Singapore at 8am SGT.
- Manual OOO spans with skip-and-cover (next available person covers).
- Dry-run when SLACK_WEBHOOK_URL is unset (prints instead of posting).
Co-authored-by: Isaac
* fix(ci): restrict GITHUB_TOKEN to contents:read in rotation workflow
CodeQL flagged the workflow for not limiting GITHUB_TOKEN permissions.
The job only checks out the repo and runs a script, so grant the minimal
contents: read and nothing else.
Co-authored-by: Isaac
* fix(ci): redact webhook URL from rotation post errors
A bare urlopen lets urllib's exception stringify the full webhook URL,
which would land in the Actions log on any POST failure. Wrap the call
and re-raise a SlackPostError carrying only the HTTP status / reason, so
the secret never appears in logs or error output.
Co-authored-by: Isaac
* refactor(ci): simplify rotation morning check to a band
Replace the exact 7/8am hour check with a "morning band" (05:00–11:59
local): ping the day's assignee only when it's currently morning where
they live, otherwise the run for their timezone's morning covers them.
This drops the DST special-casing and, more importantly, tolerates
GitHub's frequently-delayed cron schedule — a run up to ~3 hours late
still lands in the band instead of silently skipping the day. The band
starts at 05:00 rather than midnight so a delayed cron from the other
timezone spilling past local midnight can't be mistaken for this
timezone's morning and double-ping.
Co-authored-by: Isaac
* feat(ci): always report today's watch on rotation runs
The morning-band check gated even the dry-run output, so a manual
workflow_dispatch outside anyone's window just printed "nobody's on
watch" — unhelpful for a button meant for testing. Log today's assignee
per timezone unconditionally before the gate, so a manual run is always
informative; pinging still only happens inside the morning window.
Co-authored-by: Isaac
* ci(images): make the Docker build check a required merge gate
The build-only PR check added in #2288 has proven fast (~1m28s cache-cold)
and reliable, so promote it from report-only to a blocking merge gate.
- required.sh: add "Docker build" to REQUIRED, and to ALLOW_SKIP with a
workflow_for() arm so a PR whose paths filter skips the build (nothing
image-relevant changed) doesn't strand the gate — a missing check is
treated green only when its workflow legitimately didn't run.
- merge-ready.yml: add "Docker build" to the workflow_run list so the gate
re-evaluates when the build completes.
Safe for fork / non-maintainer PRs: the check builds with push:false (no
secrets, no registry) and already runs behind the security gate, so it
behaves identically to a maintainer PR.
Co-authored-by: Isaac
* fix(tests): give each xdist worker its own snapshot_failures dir
The pytest-playwright-visual-snapshot plugin's session-scoped autouse
cleanup_snapshot_failures fixture runs in every pytest session — including
the non-visual unit shards — and rmtree->mkdir's a single static path. Under
xdist, all workers race on that one path: the non-atomic rmtree/mkdir lets
one worker's mkdir(exist_ok=True) re-raise FileExistsError when another
deletes the dir in the window, and that fixture error cascades to every test
on the worker (47 spurious failures in the runtime-core shard on CI run
29072231637).
Override the fixture in the root tests/conftest.py so it keys the failures
leaf off PYTEST_XDIST_WORKER (snapshot_failures/gwN). No two workers ever
touch the same directory, so the race is gone by construction — no retries
or sleeps. The shared parent is only ever created, never deleted, so the
plugin's delete-then-create-the-same-dir window cannot recur. Without xdist
(the serial ui-snapshot.yml gate) the worker id is unset and the base path
is used unchanged.
Co-authored-by: omnigent <noreply@omnigent.ai>
---------
Co-authored-by: omnigent <noreply@omnigent.ai>
Each omnidev dev pod now gets its own config.yaml under <pod>/config/,
pointed to by OMNIGENT_CONFIG_HOME (which omnigent's server/host/runner
already honor). On first create it is seeded from the developer's real
~/.omnigent/config.yaml so the pod works out of the box (keeps their
providers); thereafter the two are independent, so server-config edits
made while testing in a pod no longer leak into the real user config.
--clean wipes the pod dir, so the next run re-seeds.
Co-authored-by: Isaac
* ci(images): make the Docker build check a required merge gate
The build-only PR check added in #2288 has proven fast (~1m28s cache-cold)
and reliable, so promote it from report-only to a blocking merge gate.
- required.sh: add "Docker build" to REQUIRED, and to ALLOW_SKIP with a
workflow_for() arm so a PR whose paths filter skips the build (nothing
image-relevant changed) doesn't strand the gate — a missing check is
treated green only when its workflow legitimately didn't run.
- merge-ready.yml: add "Docker build" to the workflow_run list so the gate
re-evaluates when the build completes.
Safe for fork / non-maintainer PRs: the check builds with push:false (no
secrets, no registry) and already runs behind the security gate, so it
behaves identically to a maintainer PR.
Co-authored-by: Isaac
* Stabilize interrupt forward ordering test
Co-authored-by: omnigent <noreply@omnigent.ai>
---------
Co-authored-by: omnigent <noreply@omnigent.ai>
* feat(browser): embedded browser pane + design mode
Add a user-driven embedded Chromium browser as a right-rail Workspace tab
in the Electron desktop app: a native WebContentsView per conversation,
positioned over a measured placeholder, with a URL bar + back/forward/
reload/DevTools toolbar. Includes design-mode point-and-prompt — hover to
highlight an element, click to open an anchored input, Send routes the
element + a cropped screenshot to the agent through the normal chat path
(no backend route).
The renderer consumes the backend's `browser.action_request` SSE event by
string key and drives the view via a claim-first relay hook; the coupling
to the agent-tools half is this runtime event only — no compile-time
dependency, so this half builds and tests standalone.
Hardening: agent-issued navigation is gated by a scheme/host allowlist
(browserUrlPolicy.js — no file://, loopback, metadata, or private hosts);
design-mode submit markers require a real native input gesture within a
short window and carry a per-enable nonce, so a hostile page can't forge
unattended submits.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* refactor(browser): extract design-mode picker script to its own module
Move the ~270-line design-mode picker driver (the in-page IIFE injected
via executeJavaScript) out of the inline template literal in browserIpc.js
into web/electron/src/designModeScript.js, so it lints and highlights as
its own file instead of an opaque backtick string.
Behavior is byte-identical: the function is moved verbatim, keeping its
(nonce) signature and internal SELECT/SUBMIT/DISMISS marker derivation, so
the produced script string matches the old one exactly for the same nonce
(verified by diffing the output across several nonces). browserIpc.js now
imports buildDesignModeScript and re-exports it, so the existing tests that
require it from browserIpc keep working unchanged. No security logic
touched — the per-enable nonce, gesture gate, and console-marker channel
are all preserved as-is.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* docs(browser): tighten comments across the browser UI
Compress verbose multi-sentence comment blocks and JSDoc prose to terse
one-liners across the net-new browser UI files (normalizeTypedUrl,
browserActionBus, designModePrompt, browserUrlPolicy, BrowserPane,
useBrowserAgentRelay, browserViewBounds, railTabs). For the large shared
files (events.ts, sse.ts, chatStore.ts, AppShell.tsx, WorkspacePanel.tsx)
only OUR added comments were trimmed — every pre-existing upstream comment
is byte-identical.
Comments/docstrings only — no logic, identifier, JSX, or string changes;
JSDoc @param/@returns type tags preserved (tsc still parses). Load-bearing
WHYs kept as one-liners: the nav-allowlist SSRF rationale, the design-mode
gesture/nonce security note, the claim-first Risk-1 note, the rAF/layout
traps in BrowserPane.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* docs(browser): drop internal review-tracker references from comments
Remove internal security-review severity labels (P0/P1/P1-1/P1-2, "P1 fix")
and private design-doc citations (Risk-1/Risk-2/Risk-4) from browser-UI
comments, docstrings, the electron README, and test describe() names —
they're meaningless/leaky to a public reader. The security invariants
themselves are kept (nonce gating, isPinnedOriginSender gate, agent-nav
allowlist, execute trust boundary, single-winner claim) — only the
internal citation is dropped. Comments/test-names only; no logic change.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* docs(electron): fix browser-pane README terminology + split framing
Two accuracy fixes in the embedded-browser section:
- the browser_* tools are framework-owned BUILTIN agent tools, not MCP
tools — drop the "MCP" wording.
- post-split this README ships in the UI PR (the pane + toolbar + design
mode + renderer plumbing); frame the agent-facing browser_* tools as
landing in a separate PR, and the relay as receiving action requests
from it. Docs-only.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* docs(browser): drop redundant SECURITY labels from comments
The SECURITY: prefix was on 7 Electron comments; most just narrate normal
behavior. Drop it from the 5 narration ones (keeping the sentence) and keep
it on the 2 genuine do-not-regress invariants: the preload's deliberate
omission of a generic agent evaluate, and the console.log main-world
back-channel note the nonce gate depends on. Comments-only.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* docs(browser): drop internal phase reference from comments
Remove the internal "Phase 2" plan reference from 3 spots we added (README
heading, main.js browserRegistry docstring, ChatPage.tsx comment) — it cites
a private phased plan, meaningless on a public repo. Also reword the
normalizeTypedUrl header + the README URL-bar note to use neutral examples
(localhost) instead of internal intranet shortnames (go/ , jira/). Keeps the
technical point (dotless host → http, host-with-dots → https); comments/docs
only, code already generic.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* test(browser): use neutral hostnames in URL-normalization tests
Replace internal-convention fixtures (go/, glean, jira/PROJ) and the
"(corp shortname)" test name with neutral dotless hosts (myhost, wiki/…)
that exercise the same behavior. Assertions unchanged in intent — dotless →
http://, dotted → https://, explicit scheme preserved; test count stays 5.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* fix(deps): use public npm registry URLs in lockfile
The lockfile's resolved URLs pointed at an internal npm proxy
(npm-proxy.cloud.databricks.com), recorded when the lockfile was
reconciled after an upstream merge. That both leaks internal infra on a
public repo AND breaks npm ci for external contributors, who can't reach
the proxy. Swap all 137 resolved URLs to registry.npmjs.org; the
content-based sha512 integrity hashes are unchanged and still verify
(npm ci --dry-run: up to date, no integrity errors). Resolved-URL host
swap only — no version, integrity, or dependency-tree change.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* docs(browser): rename AP->server in comments (use codebase terminology)
"AP" was internal design-doc vocabulary; Omnigent's own terms are
server/runner/host. Rename the 6 relay-hook comment/JSDoc references to
"server". Comments only; identical meaning.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* docs(browser): add architecture diagram to the browser-pane README
Add a Mermaid sequence diagram to the embedded-browser-pane section
showing the action flow (agent → server → renderer/pane → local
WebContentsView → back), plus a one-line prose summary. Kept UI-PR-honest:
the diagram notes the browser_* tools ship in a separate PR and labels the
renderer/pane as "(this PR)". Docs-only.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* test(browser): add e2e_ui coverage for the browser pane tab
Add tests/e2e_ui/browser/test_browser_tab.py covering the desktop-only
embedded-browser rail tab, to satisfy the E2E UI Required gate on the UI PR.
The pane is gated on isElectronShell(); the e2e_ui harness runs plain
Chromium, so — following the sessions/test_pinned_session_hotkeys.py and
mobile/test_android_shell.py precedent — the test injects a minimal
window.omnigentDesktop electron stub via add_init_script before navigation.
Two cases: (1) under the stub the "Browser" tab appears in the Workspace
rail, is the LAST tab, and selecting it mounts the pane (aria-selected);
(2) in a plain browser (no stub) the tab is absent while Agents renders.
DOM-based assertions, no LLM turn; runs against the harness's mock-LLM
server. Verified locally: 2 passed.
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* fix(browser): prettier formatting + lockfile sync
Two CI-gate fixes, no logic changes:
- Prettier: reformat the 10 browser files that drifted from prettier
style (whitespace/wrapping only; jargon scrubs preserved). `npm run
format:check` now clean.
- Lockfile: regenerate web/package-lock.json exactly as the lint.yml gate
does (`npm install --package-lock-only --legacy-peer-deps`), which
prunes the extraneous peer-pulled entries the check flagged. Idempotent
(2nd regen = no diff); npm ci --legacy-peer-deps consistent. Kept the
registry public (0 databricks-proxy hosts).
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* test(browser): raise UI coverage for browser-pane modules
Add honest unit coverage for the under-tested browser modules that were
dragging aggregate UI coverage down:
- useBrowserAgentRelay.ts: 5.55% -> 97.22% — claim-first protocol (win /
lose / not-ok / throw), the full action-dispatch switch (navigate /
screenshot / snapshot / click-by-ref+selector / type), arg marshaling,
error + timeout branches, and result-POST resilience.
- browserActionBus.ts: 12.5% -> 100% — subscribe / emit / unsubscribe /
dedupe / throwing-listener isolation.
- BrowserPane.tsx: extend the existing RTL test with toolbar handlers
(reload / devtools / nav-state enable / url-bar reflect / dotless
navigate).
- WorkspacePanel.tsx: cover the Browser tab render + pane-mount branch.
Tests only; no source change. Aggregate UI line coverage 79.97% -> 80.59%.
(Still ~0.04% under the 80.63% baseline — see PR discussion re: baseline.)
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
* fix(browser): enforce agent-nav allowlist on redirects + deny child window.open (SSRF hardening)
B1 (blocking SSRF bypass): the agent-navigation allowlist was checked once,
before the initial loadURL. A server 302 / meta-refresh / location.href during
an agent nav then redirected the child view to an internal host (metadata /
loopback / RFC-1918) with no re-check, and browser_screenshot could exfiltrate
it. Wire will-navigate / will-redirect / will-frame-navigate on the child view
and preventDefault() any disallowed target, emitting a browser-nav-blocked
signal. Enforced only while the view is agent-locked (a per-entry flag set from
opts.agent on each navigation), so user-typed URL-bar browsing — including
legitimate auth-redirect chains to internal hosts — stays permissive.
S3: the child WebContentsView had no window-open handler, so a visited page
could spawn shell windows. Deny every window.open on the child view (safe
default; not routed to shell.openExternal — an agent page popping the user's
real browser is itself an abuse vector).
Tests: will-redirect/will-navigate to metadata/loopback/RFC-1918 on an
agent-locked view is preventDefault'd + signals blocked; a normal https→https
redirect is allowed; user-driven (non-agent) nav is NOT gated; a later user nav
unlocks a previously agent-locked view; the window-open handler denies popups.
Fast-follows noted, not in scope: S1 (DNS-rebinding, needs socket-level),
S2 (IPv6 fc00::/7 + IPv4-mapped hex holes in isBlockedHostname).
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
---------
Signed-off-by: Jackson Zheng <36802691+zhengwin@users.noreply.github.com>
The rich.Live progress table flickered and made the cursor jump around during a
run. Three causes, all fixed:
- refresh_per_second lowered 8 -> 4: fewer full repaints of a growing table.
- vertical_overflow="visible": a grid taller than the viewport now prints in
full instead of rich clipping + repositioning it each frame (the cursor-jump
thrash).
- whole-harness skip reason no longer appended to the row label: a long reason
(up to 60 chars) + transport tag could wrap the Harness cell, changing row
height mid-run and forcing a reflow. Rows are now always one line high. The
reason is unaffected in output — it still prints in the stdout Notes section
after the run (sourced from the matrix, not this sink).
Removes the now-dead self._notes state. Bench suite green; ruff clean.
Co-authored-by: Isaac
* feat(harness-bench): add policy_allow + policy_ask probes
Extends the policy axis beyond DENY toward Tomu's ALLOW/DENY/ASK matrix. The
DENY probe proved a policy can block a call; these prove the other two verdicts:
- policy_allow: an explicit action=allow tool_call policy lets the call proceed
(tool_call_allowed set from a non-blocked function_call_output).
- policy_ask: an action=ask policy parks the call on an elicitation
(response.elicitation_request), which the driver resolves with an approval
accept event so the turn settles instead of parking for the day-long ASK
timeout. elicitation_requested is the observed signal.
Mechanism (full-server, the transport where policy is observable): generalize
the spec-baked deny into a fixed-action policy — _build_bench_agent_config /
register_agent take policy_action ("allow"/"deny"/"ask"); the driver caches one
session per action (_ensure_policy_session) and adds policy_probe_turn /
run_policy_turn. _scan_tool_items now also sets tool_call_allowed.
Honest SKIP elsewhere (per the coverage decision): sdk-inproc (wrap-only, no
policy surface) and native-tui (CEL ALLOW/ASK attach is a follow-up) return an
unmeasured result, so the probes SKIP rather than assert a false verdict. Native
Policy DENY stays covered by run_tool_turn(deny=True). MCP-vs-native tool
distinction is the next PR (PR-B3).
Both probes are P1 and undeclared in the manifest (like cost_tracking): no
capability axis, verdict varies by transport, so declaring SUPPORTED would
manufacture false DRIFT. TurnResult gains elicitation_requested /
tool_call_allowed.
New test_policy_matrix.py (network-free) covers both probes' verdict branches.
Full bench suite 98 passed / 18 skipped; ruff clean; no uv.lock drift. Lands in
tests/harness_bench/ (not the parked package-move location).
Co-authored-by: Isaac
* docs(harness-bench): document Policy ALLOW / ASK
Add the two new policy verdicts to the README alongside Policy DENY: the
plain-terms table (ALLOW = the call actually goes through, not just
"wasn't blocked"; ASK = the call pauses for an approval prompt / elicitation),
the per-transport "what a ✓ verifies" table (full-server spec-baked allow/ask;
`·` on native-tui and sdk-inproc, where the attach is a follow-up), and Scope
(live on full-server; native ALLOW/ASK + MCP-vs-native distinction noted as
open items). Also updates the "what a ✓ means" narrative so the transport-`·`
cells include ALLOW/ASK, not just DENY-under-`--fast`.
Docs only.
Co-authored-by: Isaac
* refactor(harness-bench): address review notes on policy probes
Review feedback (Polly + code-quality bot):
- Document the two best-effort except blocks in policy_probe_turn's watcher
(code-quality: empty-except) — note when an unparseable elicitation id means
the turn parks to the deadline, and that an SSE read error must not fail it.
- Tighten the tool_call_allowed docstring: it's set for any non-blocked tool
output, not only under ALLOW; the probe's correctness comes from driving a
real action=allow session.
- Extend the manifest UNKNOWN-not-declared note to cover policy_allow/policy_ask
alongside cost_tracking.
- Trim verbose comments/docstrings per request (probes ~69->56 lines).
Stacking note from the review is already resolved: rebased onto main after
#2307 landed, so the cost feature reconciles to zero-diff here. Subscription-
race (time.sleep before ASK subscribe) left as a documented P1 live-flake.
100 passed / 18 skipped; ruff clean.
Co-authored-by: Isaac
* perf(harness-bench): policy_ask returns as soon as the elicitation fires
The ASK verdict is decided the moment response.elicitation_request arrives, but
the loop kept polling the turn to a terminal state — so a run where the model
never called the tool (no elicitation) burned the full 180s timeout before
SKIPping. Now: once elicitation_requested is set, resolve the elicitation (so no
park dangles) and break immediately. Also lower the timeout 180s -> 90s, so the
worst case (no tool call) is a bounded SKIP, not a 3-minute stall.
A real ASK success now returns with elicitation_requested=True but
completed=False (we don't wait for the turn to settle); added a unit test
locking that verdict shape.
Co-authored-by: Isaac
* fix(harness-bench): nest elicitation_id in data so the ASK resolve lands
Polly caught a real defect: _resolve_elicitation posted the approval event with
elicitation_id at the TOP LEVEL, but POST /v1/sessions/{id}/events deserializes
into SessionEventInput (no top-level elicitation_id field) and the handler reads
data.get("elicitation_id"). So the id was dropped, no Future matched, and the
resolve was a silent no-op — the parked ASK elicitation dangled until server
teardown.
Fix: send the canonical shape {"type":"approval","data":{"elicitation_id":...,
"action":"accept"}} (matches test_sessions_endpoints.py:4960). The ASK verdict
was already correct (decided when response.elicitation_request fires); this makes
the method actually settle the parked turn as intended.
Added a network-free test asserting the id is nested in data (guards the payload
shape a fake-client can verify without a live server).
102 passed / 18 skipped; ruff clean.
Co-authored-by: Isaac
* refactor(harness-bench): key ASK watcher on parsed event type, not substring
Per Polly's non-blocking note: the SSE watcher matched on the substring
'"response.elicitation_request"' in the raw frame, so an unrelated frame merely
mentioning that string (e.g. a mirrored/resolved event) could set the ASK
verdict early. Parse the frame once with json.loads and key on
frame.get("type") == "response.elicitation_request" instead — more robust, and
the parse was already happening right after to read the id.
102 passed / 18 skipped; ruff clean.
Co-authored-by: Isaac
* docs(readme): point to the harness test bench
The harness test bench (tests/harness_bench/) has no pointer from the
root README, so contributors adding or changing harness support can
easily miss it. Link to it from the Contributing section alongside
the design doc.
Signed-off-by: Pat Sukprasert <pattara.sk127@gmail.com>
* Apply suggestion from @PattaraS
---------
Signed-off-by: Pat Sukprasert <pattara.sk127@gmail.com>
The pi JS extension and the opencode policy plugin run OUT of the runner
process and POST to the omnigent server with a hand-rolled `Authorization:
Bearer` header, bypassing databricks_request_headers -- the single chokepoint
that folds in the server-routing selectors (X-Databricks-Org-Id and the opaque
OMNIGENT_DATABRICKS_EXTRA_HEADERS map that some Databricks deployments use to pin
a request to a specific server instance). Without those selectors their POSTs can
land on a different server instance than the one the runner and the web UI are
bound to, so on a multi-instance deployment pi's streamed items never reach the
browser's in-process event stream (they only appear on reload) and opencode's
policy evaluation hits a different instance.
- cli_auth: fold OMNIGENT_DATABRICKS_EXTRA_HEADERS into
databricks_request_headers (opaque JSON header map; no-op when unset).
- pi: build the extension config.authHeaders (launch + per-turn refresh) via
databricks_request_headers.
- opencode: bake the full routing header map as OMNIGENT_POLICY_HEADERS and merge
it in the policy plugin, replacing the bearer-only OMNIGENT_POLICY_AUTH.
- host: allowlist OMNIGENT_DATABRICKS_EXTRA_HEADERS in the host->runner env
builder so a host forwards the routing selectors to the runners it spawns.
Without it the host tunnel lands on the selected instance while its runners
fall back to the default one (their tunnel + callbacks register elsewhere), so
the session's runner is unreachable from the instance serving the UI and the
session reports runner_failed_to_start.
In-runner Python clients already route via _RunnerDatabricksAuth / _remote_headers;
the gaps were the two out-of-process posters and the host->runner env handoff.
Co-authored-by: Isaac
Signed-off-by: Edwinhe03 <41037314+Edwinhe03@users.noreply.github.com>
* feat(harness-bench): add cost_tracking probe
Cost tracking is the keystone for cost policies (Tomu): a cost_budget guardrail
is a no-op without usage to measure. This adds a P1 cost_tracking probe that
answers "can the operator see what a turn spent?".
- TurnResult gains total_tokens / total_cost_usd (both Optional; None = the
transport surfaced no usage).
- fill_snapshot_cost(result, snapshot) in driver.py reads the cumulative
totals the server records on the session snapshot (SessionResponse
total_cost_usd / last_total_tokens) — the uniform read point both
server-backed drivers already poll. full-server fills it on turn completion;
native-tui reads the snapshot post-turn (its usage arrives via
external_session_usage -> session.usage). sdk-inproc (wrap-only, no server)
fills from the completed turn's embedded usage when the wrap forwards it,
else leaves it None.
- Probe verdicts: SUPPORTED (priced cost), PARTIAL (tokens but no price =
unpriced model — usage visible, USD-cost policy can't price it), SKIPPED
(no usage surfaced / infra failure / timeout). Never a false UNSUPPORTED.
- Deliberately NOT declared in the manifest (left UNKNOWN): no backing
capability axis, and the observed verdict legitimately varies, so declaring
SUPPORTED would manufacture false DRIFT against a legitimate PARTIAL. The
P0-coverage test only requires declared verdicts for P0 dims, so a P1
probe with no declaration is allowed.
New test_cost_tracking.py (network-free) covers the verdict logic +
fill_snapshot_cost. Full bench suite 89 passed / 18 skipped; ruff clean; no
uv.lock drift. Lands in tests/harness_bench/ (not the parked package-move
location).
Co-authored-by: Isaac
* fix(harness-bench): cost probe requires positive usage, not just non-None
A completed turn always spends tokens, so a reported total_cost_usd == 0 or
total_tokens == 0 means the usage plumbing returned an empty default, not that
tracking genuinely measured zero. The `is not None` check would render a $0.00
turn as SUPPORTED — a false pass. Require a POSITIVE value:
- cost > 0 -> SUPPORTED
- tokens > 0 (cost None/0) -> PARTIAL (unpriced)
- both absent or zero -> SKIPPED
Readers (fill_snapshot_cost, sdk-inproc) still carry whatever the server
reported (including 0, distinct from absent); the >0 judgment lives in the probe
where interpretation belongs. Added tests for the 0/0 -> SKIP and
0-cost/positive-tokens -> PARTIAL cases.
Co-authored-by: Isaac
* docs(harness-bench): document cost_tracking; drop P0/P1 jargon
Add the Cost tracking dimension to the README: the plain-terms table (✓ priced
cost / ~ tokens-only / · no usage, and that it gates any cost policy), the
per-transport "what a ✓ verifies" table (snapshot read on server transports;
wrap-usage on sdk-inproc else ·), and the Scope section (now live).
Drop the P0/P1 framing from the public-facing doc — it's internal
(merge-gating vs reported) and doesn't help a reader. The Priority field stays
in code; the README just describes the dimensions.
Also corrects a stale Scope claim: native Tool calling / Policy DENY are
observed now (landed separately), not "not yet wired".
Docs only.
Co-authored-by: Isaac
* fix(electron): reload desktop window when workspace SSO session expires
A workspace-hosted Omnigent sits behind the Databricks SSO gate. When
that outer session's cookie lapses, the gate answers the SPA's API calls
with a 303 redirect to its own login.html instead of the expected JSON.
The SPA can't parse the login page as data and dies on a "Failed to
load: Fetch request failed due to expired user session" panel — and a
desktop user has no address bar to force a refresh out of it.
An earlier attempt handled this in the web SPA (identity.ts), but that
can't work here: the desktop app loads whatever bundle the remote server
serves, so an un-deployed SPA change never runs, and the host fetcher
rejects before any status/content-type check the SPA could inspect.
Handle it in the Electron shell instead. The shell sees the raw redirect
via session.webRequest.onBeforeRedirect regardless of which server bundle
is loaded, so it detects a 3xx redirect to login.html for a connected
server origin and reloads the affected windows. The reload re-issues the
top-level navigation the SSO gate inspects, so it can re-challenge and
re-mint the session. A per-window minimum interval caps reloads so a
persistently expired host can't reload-loop.
The detection logic lives in an Electron-free module (session-expiry.js)
so isLoginRedirect and the onBeforeRedirect wiring are unit-testable via
node --test without booting the app.
Co-authored-by: Isaac
* fix(electron): skip destroyed windows in the session-expiry reload loop
The reload loop in registerSessionExpiryAccess called win.webContents.reload()
without checking win.isDestroyed(). A BrowserWindow handle can outlive its
native window (the windows map keeps it reachable until the "closed" handler
removes it), so in the race between native destroy and map removal a
login-redirect callback could call reload() on a dead handle — which throws out
of the onBeforeRedirect listener and skips the remaining windows.
Fold the isDestroyed() check into the existing continue-guard, matching the
idiom used elsewhere in this file when iterating the windows map.
Co-authored-by: Isaac
---------
Co-authored-by: Amruth Sampath <amruth.sampath@databricks.com>
* fix(web_fetch): probe for bwrap at researcher-spec build time
A parent with no os_env hands the __web_researcher sandbox=None, which
resolve_sandbox fills with the platform default (linux_bwrap on Linux)
without checking the binary exists. The spawn then failed mid-run and
the error told the user to set os_env.sandbox.type, which a spawn-only
parent cannot apply without also registering OS tools on itself.
Probe shutil.which("bwrap") in build_researcher_spec for the no-os_env
case and fail at spec-build time with the remediation the operator can
actually use: install bubblewrap on the host. Parents that declare
their own os_env keep the inherit-verbatim path untouched.
Fixes#2068
Signed-off-by: Enes Yilmaz <115046343+EnesYilmazcode@users.noreply.github.com>
* fix(web_fetch): extend the seed-time sandbox probe to macOS
Review follow-up on #2097: darwin_seatbelt needs sandbox-exec on PATH,
mirroring the fail-loud check in SeatbeltSandboxBackend.resolve. The
Windows default windows_jobobject drives kernel Job Objects through
ctypes with no external binary, so there is nothing to probe there;
documented in the docstring.
Signed-off-by: Enes Yilmaz <115046343+EnesYilmazcode@users.noreply.github.com>
* test(web_fetch): keep seed-time sandbox probe host-independent
The new _ensure_default_sandbox_runnable() probe calls shutil.which
against the real host PATH for a no-os_env parent, so every existing
test that builds a researcher spec from such a parent now raises
OmnigentError on any runner without bubblewrap / sandbox-exec
installed (the unit-test CI job). Add an autouse fixture defaulting the
probe to "binary present"; the probe-specific tests override it with
their own monkeypatch.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SnpHpxeDkqfkrUEt3Sc3sj
---------
Signed-off-by: Enes Yilmaz <115046343+EnesYilmazcode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
* fix(smart-routing): enforce rationale consistency with selected model tier
Restructures the judge prompt to require explicit SIMPLE/MODERATE/COMPLEX
task classification, each mapped to a concrete model tier (haiku/sonnet/opus,
nano/mini/base), and enforces a structured rationale format so the explanation
always matches the chosen model.
* fix(smart-routing): restore Trade-off guidance label
* fix(electron): resolve lockfile from public npm registry
web/electron/package-lock.json pinned 286 of its 290 resolved URLs to the
internal npm-proxy.cloud.databricks.com mirror, which is unreachable from
public GitHub runners. npm ci fetches each tarball from its exact resolved
URL, so the Electron Build workflow stalled for ~8 minutes on the first fetch
and died with "Exit handler never called!" on both Linux and Windows.
Rewrite those URLs to registry.npmjs.org, matching web/package-lock.json
(already all-public) and the uv.lock normalization. The integrity hashes are
content-based and unchanged, so they still validate against the public
tarballs.
Co-authored-by: Isaac
* fix(electron): add publish provider and repository so build completes
After packaging the AppImage/deb/nsis artifacts, electron-builder 26.x crashed
in computeChannelNames with "Cannot read properties of null (reading 'channel')"
because it computes auto-update channel metadata but found no publish provider
and could not detect the repository (repeated "Cannot detect repository by
.git/config" warnings).
Add a github publish provider and a top-level repository field. Under
--publish never the metadata is generated locally without uploading, so the
build no longer throws.
Co-authored-by: Isaac
* fix(policy-hook): improve reauth logging and proactively refresh lapsed bearer
The baked one-shot hook token was silently failing: all exceptions in
_reauth() were swallowed with no stderr, making it impossible to tell
whether the factory import failed, no credential was available, or the
mint itself threw. Add distinct log lines for each failure path.
Proactively re-mint the bearer before the first evaluate POST when the
JWT exp claim shows the token is within 5 min of expiry (or already
lapsed). Handles the "runner older than ~1h" case without waiting for a
401/302 — the one-shot reauth fires before the request rather than as
a recovery.
* fix(policy-hook): drop proactive reauth — only improve failure logging
Proactive JWT expiry check was not fixing the actual failure pattern:
when reauth() returns None (the bug case), proactive fires first,
gets None, and the session still fails closed — same outcome as before.
Remove it.
Keep only the logging improvements: each _reauth() failure path now
prints a distinct stderr message instead of silently returning None.
* fix(policy-hook): treat 403 as re-auth signal alongside 401 and 302
Databricks Apps returns 403 "Invalid Token" for an expired bearer, not
401. Both _is_login_redirect_or_unauthorized implementations only
checked 401 and 302→/oidc/, so the 403 fell through as a final
non-retryable 4xx — the reauth callable was never invoked and the hook
failed closed on every call for sessions older than ~1h.
Extend both the hook and runner functions to treat status 401 and 403
as re-auth signals. Add a parametrize case for 403 in the classifier
test and an integration test that a 403 response triggers reauth and
retries with the fresh token.
* test(policy-hook): harness-level regression test for 403 reauth
Mirrors test_evaluate_policy_reauths_on_expired_token_instead_of_failing_closed
but with a 403 "Invalid Token" response instead of 302→/oidc/. Drives the
full claude_native_hook.main() → bridge dir → httpx → PolicyHookReauth →
retry path, asserting two attempts (stale token, then fresh) and that the
routing header survives the re-mint.
* fix(policies): apply DB-stored default policies to every session evaluation
PolicyStore.list_defaults() (policies created via POST /v1/policies with
session_id=NULL) was never consulted during engine construction — only
YAML-based caps.default_policies were included in admin_policy_specs.
Added _load_default_policy_specs() and call it in build_policy_engine so
DB-stored defaults are fetched fresh on every evaluation, inserted between
agent-spec policies and the YAML admin policies.
* feat(policies): cache DB default policy specs; add tests
- Add _DEFAULT_POLICY_SPECS_CACHE (TTLCache, 30 s, keyed by workspace_id)
in builder.py so list_defaults() is only called once per 30-second
window per workspace instead of on every tool-call evaluation.
- Add invalidate_default_policy_specs_cache() and call it in the
create/update/delete default policy routes so changes propagate
immediately rather than waiting for the TTL to expire.
- Add tests: _load_default_policy_specs (none store, filters disabled,
cache hit, invalidation), build_policy_engine DB-default inclusion,
and the full four-layer ordering (session → agent → DB default → YAML admin).
* fix(policies): guard against url-type default policies bricking all sessions
A single enabled url-type default policy would raise OmnigentError in
_load_default_policy_specs on every build_policy_engine call, taking
down session construction server-wide. Two-pronged fix:
- Reject type='url' at create_default route: default policies now only
accept type='python' (same restriction as session policies, but
enforced at API time so the bad state can't be persisted).
- Skip-with-warning in _load_default_policy_specs for any unsupported
type: a stale or manually-inserted row is logged and skipped rather
than raising, limiting blast radius to a warning log entry.
Adds test asserting the skip-with-warning path (url row skipped, python
row still included).
* test(policies): fix default policy route tests to use type='python'
The create_default route now rejects type!='python'. Update tests to use
a registered python handler, add test_create_url_policy_rejected to
assert the 400, and remove the stale url-type payload from _policy_payload.
* feat(policies): cache session policy specs with invalidation on mutation
Add _SESSION_POLICY_SPECS_CACHE (plain dict, no TTL) keyed by
(workspace_id, conversation_id). Unlike default policies (TTL cache),
session policies must be visible immediately after sys_add_policy, so
invalidation-on-mutation is used instead of TTL.
invalidate_session_policy_specs_cache() is called after create, update,
and delete in the session policies route. Tests cover cache hit and
invalidation behavior.
* test(policies): fix oidc default policy test to use type='python'
* fix(policies): bound session policy cache (LRU) and remove dead branch
- Switch _SESSION_POLICY_SPECS_CACHE from unbounded dict to
LRUCache(maxsize=4096), matching _SESSION_OWNER_CACHE and preventing
unbounded memory growth on long-lived servers.
- Remove the dead `if body.type == "python":` branch in create_default
(unreachable after the preceding `if body.type != "python": raise`).
* fix(host): re-exec via login shell to inherit full PATH on GUI launch
GUI-launched Electron inherits a minimal PATH from the desktop launcher
(launchd on macOS, systemd on Linux) that omits Homebrew, nvm, pyenv and
other user-installed tool directories. This meant claude, codex, tmux and
similar tools were missing when spawned from the Omnigent desktop app.
Extract loginShellPath.js to resolve the full login-shell PATH by spawning
`$SHELL -l -c 'echo $PATH'` and patch process.env.PATH at Electron startup.
Add Playwright browser-flow tests for the resolver's pure resolution logic
(trim, null-on-failure, colon-separated output) via dependency injection.
* fix(host): harden login-shell PATH resolution (-ilc, delimiter, merge, real test)
The login-shell PATH resolver worked for the simple case but missed the
edge cases that hit exactly the GUI-launch users #1933 targets:
- Use `-ilc` (interactive+login) instead of `-l`. A login-only shell sources
the profile but NOT the rc file (.zshrc/.bashrc), where nvm/pyenv and most
hand-rolled PATH exports live — so `-l` alone still missed those tools.
- Source the shell from the passwd DB (os.userInfo().shell), then $SHELL, then
a POSIX fallback list. $SHELL is typically unset in a GUI launch (the premise
of this bug), so relying on it fell back to /bin/bash for zsh users.
- Bracket $PATH in delimiter markers and strip ANSI before parsing, so an
rc-file banner / MOTD / version-manager greeting can't corrupt the result.
- Suppress hang-prone startup hooks (oh-my-zsh auto-update, zsh tmux plugin,
pagers) in the child env so a heavy rc file doesn't trip the timeout.
- Recover a delimited PATH from err.stdout when a shell exits non-zero after
already printing it.
- Add a fast-path skip when PATH already looks complete (launched from a
terminal), and merge (union, dedup) rather than replace process.env.PATH —
matching what the main.js comment already claimed.
Tests: replace the Playwright/Python test (which exercised a reimplementation
of the resolver in a browser, not the shipping module) with a node --test suite
that requires the real loginShellPath.js and injects execFileSync/os/env/platform
mocks, plus a source-guard pinning the main.js merge wiring. Full electron
suite: 76 pass.
Co-authored-by: Isaac
* style(host): prettier-format loginShellPath test
Collapse a chained .replace() onto one line to satisfy the repo's prettier
config (printWidth 100), matching the web-prettier pre-commit hook.
Co-authored-by: Isaac
---------
Co-authored-by: Zeyi (Rice) Fan <zeyi.f@databricks.com>
A gateway stream that ends without a finish_reason, no content, and no tool
calls means the worker turn died mid-stream. The executor yielded a silent
empty TurnComplete, so an aborted turn was sometimes accepted as a clean
completion and sometimes surfaced elsewhere as a reasonless failure. Emit an
ExecutorError with a clear message instead; a truncated stream that did
produce text still completes (with a warning).
Fixes#1118
Co-authored-by: ikatyal21 <ikatyal@terpmail.umd.edu>
Co-authored-by: Sabhya Chhabria <sabhyachhabria@gmail.com>
Resolving an elicitation through the resolve endpoint completes the
elicitation Future but never signals resolved_elsewhere, so a harness
turn parked on that elicitation stays parked until its timeout. Visible
symptom: approving an inbox card returns 202 and the approved tool call
never resumes.
Wire the resolve path to the existing resolved_elsewhere registry, the
same mechanism the terminal resolve path already uses. The new test
parks a harness elicitation, resolves it via the endpoint, and asserts
the parked wait wakes with the verdict; it fails before the fix.
Signed-off-by: Robert Dosen <robert.dosen@gmail.com>
A markdown file whose list has an item starting with a non-paragraph block
— a nested list (`- - x`), a fenced code block, a blockquote, a heading, or
a table — crashed the markdown editor's panel.
@tiptap/markdown (beta) parses those into a `listItem` whose first child is
that block, which violates the stock `paragraph block*` content model.
ProseMirror builds the initial document via `nodeFromJSON`, which does not
validate content, so the invalid doc loads silently — then the first
transaction that touches the list item (a user edit, or StarterKit's
TrailingNode appendTransaction that runs on load) calls `contentMatchAt` on
it and throws ("Called contentMatchAt on a node with invalid content"). The
viewer's React panel boundary catches the throw and renders a crash instead
of the file.
Relax the list item's content model to `block+` (SafeListItem) so a
non-paragraph first child is schema-valid. Same crash family as the
blockquote fix in #2004, but for list items — which agent-authored markdown
hits constantly.
Co-authored-by: Isaac
A final assistant row that lands while a poll's batch is still being
POSTed was picked up by the fresh completed-turn count at the end of the
same iteration, ringing the parent-waking idle edge before the row
itself was mirrored — a sub-agent orchestrator woke to a transcript
missing the final answer. Count only rows at or below the mirror's
high-water mark so the completion signal can never overtake the content
it announces.
Co-authored-by: tomsen-ai <230283659+tomsen-ai@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(harness-bench): derive creds like `omni run`; --profile now optional
The bench always minted its own bearer via a `databricks auth token` subprocess
(which does not handle OAuth `databricks-cli` profiles) and required --profile
for any live run -- a path entirely separate from how `omni run` authenticates.
Add tests/harness_bench/runtime_env.py with resolve_bench_env(), mirroring
`omni run`'s credential layering:
1. ambient OPENAI_BASE_URL + OPENAI_API_KEY win (skip resolution entirely, the
same short-circuit `omni run` has),
2. else the profile from --profile, else the ~/.omnigent/config.yaml
auth:/profile block (what `omni run` reads),
3. compose OPENAI_* via the canonical resolve_databricks_workspace()
(OAuth-aware, fail-loud on a typo'd profile) -- the resolver the runner uses.
So a no-flag run now derives creds exactly like `omni run`, and --profile
overrides. bench_creds_skip_reason() gives every driver's unavailable() a cheap,
token-free gate: a run skips cleanly when no creds are resolvable instead of
requiring a flag.
- SharedFullServer takes a BenchRuntimeEnv (was db_profile: str); __enter__
drops _mint_bearer + lookup_databricks_host and uses env.base_env.
- FullServerDriver / NativeTuiDriver / SdkInprocDriver resolve via
resolve_bench_env; databricks_profile is now Optional throughout (the
--profile override, None = derive). run_bench keeps the kwarg for back-compat.
- The full-server agent spec and the native provider-config omit
executor.profile / the auth: block when auth came from the ambient env.
- __main__: a live run no longer requires --profile; it turns on whenever creds
are resolvable, and --no-live forces the offline declared matrix.
This is deliberately independent of the package-move / `omni bench` work: it
stays in tests/harness_bench/ and is valid regardless of where the bench ends up
or what its user-facing entry point becomes.
Note: this drops the bench-only #1781 stale-token strip (env -u
DATABRICKS_TOKEN). Intentional -- `omni run` uses the same resolver and does not
strip either; aligning with omni is the point.
New test_runtime_env.py covers the layering (ambient wins, --profile overrides
config, config-derived, no-creds skip, hostless profile). 80 passed / 18
skipped; ruff clean; e2e still collects (376).
Co-authored-by: Isaac
* fix(harness-bench): resolve profile from providers: block, like omni run
The first cut of _profile_from_config only read the auth: block and a top-level
profile: key. But a machine configured through the provider wizard (rather than
`omni setup`) has neither -- its Databricks creds come from a
providers.databricks entry (default: true, profile: <name>). omni run resolves
that via default_provider_for_harness (runtime/workflow.py DATABRICKS_KIND
branch), so with no --profile it goes live; the bench went offline instead.
Add a third tier to _profile_from_config that reuses omni's own
default_provider_for_harness resolver (the same call resolve_credential and the
runtime spawn-env builder use) and reads .profile when it's a databricks
provider -- no reinvented selection logic, so the bench picks exactly the
profile a launch would. New test covers the providers:-block path.
81 passed / 18 skipped; ruff clean.
Co-authored-by: Isaac
A green cell is only as strong as the layer the probe drove it through, and that
differs by transport. Add a "What a ✓ actually means" section with a
per-dimension x per-transport table (full-server / native-tui / sdk-inproc)
spelling out exactly what each ✓ verifies, so a reader can tell whether a tick
implies end-to-end coverage for web-UI users.
Key points now written down instead of tribal:
- full-server (SDK default) and native-tui (native default) drive turns through
the SAME server API the web UI uses (POST /v1/sessions/{id}/events + the
/stream SSE), so a ✓ there is end-to-end through the server contract the
browser depends on -- minus the browser render layer (that's tests/e2e_ui).
- sdk-inproc (--fast) drives the harness wrap directly, below the server; a ✓
there does not imply the deployed server path works. Policy DENY is `·` there.
Also corrects two stale claims: native-tui now DOES observe Tool calling +
Policy DENY (landed in #2096/#2171), and sdk-inproc observes Tool calling (only
Policy DENY is missing there, not both).
Docs only.
Co-authored-by: Isaac
The build-only PR check added in #2288 has proven fast (~1m28s cache-cold)
and reliable, so promote it from report-only to a blocking merge gate.
- required.sh: add "Docker build" to REQUIRED, and to ALLOW_SKIP with a
workflow_for() arm so a PR whose paths filter skips the build (nothing
image-relevant changed) doesn't strand the gate — a missing check is
treated green only when its workflow legitimately didn't run.
- merge-ready.yml: add "Docker build" to the workflow_run list so the gate
re-evaluates when the build completes.
Safe for fork / non-maintainer PRs: the check builds with push:false (no
secrets, no registry) and already runs behind the security gate, so it
behaves identically to a maintainer PR.
Co-authored-by: Isaac
Resuming a claude-native session from the web UI could crash the
`claude` CLI at boot with `JSON Parse error: Unrecognized token '<'`.
Its input prompt never rendered, so the readiness gate timed out after
30s and the first message was never delivered.
On cold resume the wrapper rewrites Claude's local transcript from
committed Omnigent items, unconditionally storing the tool result string
as `toolUseResult`. Claude Code's `TaskOutput` renderer `JSON.parse`s
that field at resume time, so a plain display string (e.g. an
`isaac review` result starting with `<retrieval_status>...`) threw at
startup. The tool result content block was fine — only `toolUseResult`
is parsed.
Add `_json_safe_tool_use_result`: outputs that are already JSON (e.g.
image content-block arrays) pass through verbatim; anything else is
wrapped as a JSON string literal so the parse always succeeds. The
verbatim string still lives in the tool_result content block, so what
the model and web UI see is unchanged.
Co-authored-by: Isaac
Omnigent relay tools surfaced into Hermes (mcp_omnigent_* / mcp__omnigent__*)
are already policy-gated when the relay dispatches them back through the
server's tool path. The pre_tool_call hook evaluated them a second time, parking
a duplicate approval card per call; a human resolves one and the other's
long-poll never returns, wedging the turn after the approved tool runs. Skip
those prefixes in the hook, matching the guard the native claude/codex hooks
already apply. Hermes' own tools (shell, file) and non-Omnigent MCP servers lack
the prefix and stay gated.
Signed-off-by: rdosen <robert.dosen@gmail.com>
* feat(smart-routing): always route child sessions when parent toggle is on
Previously, smart routing was skipped for child sessions if the
orchestrator had already specified a model via sys_session_send (because
effective_runner_override was non-null). The routing verdict now always
wins over the LLM's own model choice when the parent toggle is on —
for both the SDK and native-terminal paths.
* fix: use conv.parent_conversation_id to detect child session in routing gate
* test: verify smart routing overrides orchestrator model for child sessions
Per-PR merges into main each triggered a full multi-arch image publish,
which is far more often than needed. Reduce the publish cadence and cover
the lost per-merge build validation with a build-only PR check.
- oss-publish-images.yml: drop the per-commit `push: branches: [main]`
trigger (keep `tags: ['v*']`). The daily cron now rebuilds main HEAD and
publishes :sha-<short> + :latest-nightly directly. Retire :latest-dev
(redundant with the daily :latest-nightly once per-commit builds are gone)
and the now-dead promote-nightly job + force_nightly dispatch input.
- docker-build.yml (new): on PRs touching image-relevant paths, build the
server image single-arch (amd64) with the GHA layer cache and run a
`omnigent --help` smoke, no push. Report-only for now; documented how to
promote it to a blocking merge-gate check later.
Co-authored-by: Isaac
* fix(goose): implement interrupt_session via ACP session/cancel (#1748)
The web Stop button was a no-op for the goose harness because
GooseExecutor.interrupt_session fell through to the Executor no-op.
Fix: override interrupt_session in GooseExecutor to:
1. Send ACP `session/cancel` to request a clean stop (gives Goose a
chance to close its own agent loop gracefully).
2. Fall back to SIGTERM on the subprocess when no session_id is
established yet (e.g. the process is still initializing), mirroring
the pattern used in KimiExecutor.
A dedicated `_interrupt_proc` helper (also used by the existing
asyncio.CancelledError path in run_turn) is added to avoid
duplicated terminate/suppress logic.
Tests added in tests/test_goose_executor_interrupt.py:
- interrupt with no live process → returns False
- interrupt before session established → terminates proc, returns True
- interrupt with live session → sends session/cancel RPC, returns True
- session/cancel error → falls back to SIGTERM, still returns True
* fix(goose): send session/cancel as an ACP notification
session/cancel is an ACP notification, not a request: the agent sends no
response and instead ends the in-flight session/prompt with a cancelled
stop reason. Dispatching it through _rpc() (which assigns an id and blocks
on a pending future) meant the graceful path always hit the timeout and
degraded to SIGTERM, adding latency to every Stop and never delivering the
clean partial-result cancel it was meant to.
Send it via _send() with no id, mirroring acp_executor.interrupt_session,
and let run_turn surface the cancelled stop reason. Drops the redundant
doubled asyncio.wait_for and the now-unused _CANCEL_TIMEOUT_SECONDS.
The interrupt test previously mocked _rpc to return a canned response goose
never sends, hiding the bug; it now asserts on _send and that the cancel
carries no id, exercising the real notification contract.
Co-authored-by: Isaac
---------
Co-authored-by: Daniel Lok <daniel.lok@databricks.com>
* ci: run store and db tests against PostgreSQL and MySQL
Adds two new CI jobs (stores-postgres, stores-mysql) that exercise
tests/stores and tests/db against real service containers, using a
fresh per-test database created via OMNIGENT_TEST_DB_URI. Updates the
db_uri fixture to support non-SQLite backends, adds pymysql to the
databricks extra, and fixes three SQLite-specific tests (PRAGMA
foreign_keys, FTS5 queries) to skip on incompatible backends plus one
SqlConversationItem insertion that used raw strings instead of encoded
SMALLINT values.
* fix(ci): MySQL PK fix for y1a2b3c4d5e6 widen_conversation_items_pk
MySQL PKs are unnamed; batch_alter_table can't drop then add without
erroring with 'Multiple primary key defined'. Use raw DDL for MySQL
matching the pattern from r1a2b3c4d5e6.
* fix(ci): fix remaining MySQL test failures
- conversation_store search: add MySQL dialect branch using
CONVERT(data USING utf8mb4) LIKE instead of the PostgreSQL-specific
'::text ILIKE' cast
- test_db_models + test_conversation_store: CHECK constraint violations
raise OperationalError on MySQL (code 3819), not IntegrityError;
update test_check_constraint_* and workspace-check tests to accept
both
* fix(ci): all store+db tests pass on MySQL
- permission_store: add MySQL dialect branch in grant() and ensure_user()
using ON DUPLICATE KEY UPDATE (mysql_insert) instead of PostgreSQL-
specific OnConflictDoUpdate/OnConflictDoNothing
- conversation_store search: replace 'ci.data::text ILIKE' (Postgres-only)
with CONVERT(ci.data USING utf8mb4) LIKE on MySQL
- test_db_models: CHECK constraint violations raise OperationalError on
MySQL (code 3819) not IntegrityError; accept both in check constraint tests
- test_conversation_store: same fix for workspace CHECK constraint tests
682 passed, 3 skipped locally against MySQL.
* style: ruff format
* perf(ci): session-scoped DB per worker + mysqlclient for MySQL tests
- conftest: add session-scoped _worker_db_uri fixture that creates one
database per xdist worker (not per test) and runs Alembic migrations
once. The per-test db_uri fixture truncates tables between tests for
isolation. This reduces migration runs from ~680 to 4.
- Remove FOREIGN_KEY_CHECKS toggles around TRUNCATE — all FKs were
dropped in p1a2b3c4d5e6 so the toggles are pure overhead.
- CI: install libmysqlclient-dev + mysqlclient (C extension driver)
instead of pure-Python pymysql, and switch dialect to mysql+mysqldb.
mysqlclient is significantly faster per round-trip.
* fix(policy-hook): improve reauth logging and proactively refresh lapsed bearer
The baked one-shot hook token was silently failing: all exceptions in
_reauth() were swallowed with no stderr, making it impossible to tell
whether the factory import failed, no credential was available, or the
mint itself threw. Add distinct log lines for each failure path.
Proactively re-mint the bearer before the first evaluate POST when the
JWT exp claim shows the token is within 5 min of expiry (or already
lapsed). Handles the "runner older than ~1h" case without waiting for a
401/302 — the one-shot reauth fires before the request rather than as
a recovery.
* fix(policy-hook): drop proactive reauth — only improve failure logging
Proactive JWT expiry check was not fixing the actual failure pattern:
when reauth() returns None (the bug case), proactive fires first,
gets None, and the session still fails closed — same outcome as before.
Remove it.
Keep only the logging improvements: each _reauth() failure path now
prints a distinct stderr message instead of silently returning None.
* fix(policy-hook): surface reauth failure reason in the UI error message
Hook subprocess stderr is discarded by the harness, so the reauth
failure reason was silently lost. Convert the inner _reauth() closure
to PolicyHookReauth — a callable class that records failure_reason on
each None return. Thread the reason through fail_closed_hook_output()'s
new detail param so it appears in permissionDecisionReason (the field
shown to the user in the UI) and in the block reason for
UserPromptSubmit.
Before: "Omnigent policy evaluation unavailable (could not reach or
authenticate to the Omnigent server); failing closed for this tool call."
After: "...failing closed for this tool call. Detail: no credential
resolved (no stored token and no Databricks SDK auth for '...')"
* fix(policy-hook): surface API error details in fail-closed UI message
post_evaluate_with_retry now returns (response, error) instead of
response | None. The error string captures the last failure reason
(4xx status + body preview, connection error, read timeout, budget
exhausted) so callers can include it in the deny/block reason shown
to the user — alongside the existing reauth failure detail.
Before: "...failing closed for this tool call."
After: "...failing closed for this tool call. Detail: server returned
403: <body>" / "connection error: ..." / etc.
All call sites updated (claude/kimi/codex/hermes/cursor). Cursor keeps
its fail-open policy on network error (no detail surfaced there since
nothing is blocked). Tests updated to unpack the tuple and assert on
the error field.
* test(policy-hook): relax fail-closed reason assertion to startswith
The reason now includes a "Detail: ..." suffix when an API error is
captured, so exact equality fails. Use startswith to check the base
message without coupling to the appended detail.
* feat(benchmarks): add fork, comment, and runner-file-read journeys
Extend the dev perf harness (dev/benchmarks/omnigent) with three more
user journeys:
- fork_session — POST /v1/sessions/{id}/fork then DELETE (pure HTTP)
- add_comment — POST /v1/sessions/{id}/comments (pure HTTP + DB)
- read_runner_file — GET .../environments/default/filesystem/{path},
the server → runner filesystem read proxy (needs a runner, no LLM turn)
fork and comment follow the existing runner-free journey pattern. The
runner-file read needs a bound runner: give runner-mode bundles an os_env
block so the runner can materialize the default filesystem environment
(without it the proxy 404s), and point the runner workspace at the temp
dir so planted files don't leak into the launch cwd.
Subagent spawn is left as a follow-up (recorded in the README) — it needs
mock-LLM tool-call scripting and parent/child auto-wake polling.
Co-authored-by: Isaac
* refactor(benchmarks): exclude fork DELETE from the timed span
The fork journey deleted each fork inline inside measure, folding the
DELETE into the timed op. Collect fork ids in the journey context and
delete them in teardown instead, so only the fork POST is measured.
Co-authored-by: Isaac
Add a "What each probe does" table describing the six P0 dimensions
(Basic turn, Streaming, Tool calling, Policy DENY, Model override,
Interrupt) in layman's language, plus a verdict-glyph key so a reader
who has never seen the bench can read a matrix. Also add an example
--rich run of the SDK harnesses on the oss profile, showing how a
diagnosed `·` SKIP (codex / Policy DENY) reads against the Notes line.
Docs only; no code change.
* feat(images): ship the kubernetes extra in the published server image
The kubernetes managed-sandbox provider is in the base package, but the
published omnigent-server image is built with no extras — the launcher's
lazy kubernetes-client import fails on the first managed launch, so no
official image can actually drive sandbox.provider: kubernetes. Default
OMNIGENT_EXTRAS to kubernetes (openshell variant becomes
openshell,kubernetes to stay a superset), and drop the sandbox-runners
overlay's mandatory self-built-image override now that the official
image works as-is.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(images): publish a kubernetes server variant instead of folding the extra into base
Keep the published omnigent-server image lean (OMNIGENT_EXTRAS stays
empty) and instead publish ghcr.io/omnigent-ai/omnigent-server-kubernetes,
mirroring the openshell variant end to end: tags, build step, SBOM,
nightly promotion, and floating-tag reconcile. The sandbox-runners
overlay swaps the base image for the variant via its images: block, so
`kubectl apply -k` works against official images with no self-build.
Co-authored-by: Isaac
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Bryan Li <15131870+btli@users.noreply.github.com>
Co-authored-by: Pat Sukprasert <pattara.sk127@gmail.com>
## Related issue
N/A
## Summary
- Add `NSCameraUsageDescription` and `NSSpeechRecognitionUsageDescription` usage strings (Debug + Release Info.plist) so iOS doesn't crash when the WebView requests camera or speech-recognition access.
- Gate WebKit media capture with `isAllowedMediaCaptureType`, allowing camera, microphone, and cameraAndMicrophone (previously microphone-only) and still only for the pinned app origin.
- Repair duplicate `PrivacyInfo.xcprivacy` object IDs in the Xcode project so the iOS target compiles.
## Test Plan
- Added `AppPrivacyInfoTests.testPrivacyUsageDescriptionsArePresent` asserting the camera, microphone, and speech-recognition usage strings are present and non-empty in the app bundle.
- Built the iOS target (duplicate object IDs previously broke the build) and exercised the camera/mic capture prompt via the WebView.
## Demo
N/A
## Type of change
- [x] Bug fix
- [ ] Feature
- [x] UI / frontend change
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Unit test verifies the required iOS privacy usage strings are present. Manual verification: built the iOS target and confirmed the camera/microphone capture prompt no longer crashes and is granted only for the pinned origin.
## Changelog
[UI] Fix iOS crash when granting camera or voice-dictation permission in the app
`omni run --harness acp:<slug>` (a configured ACP agent, e.g. acp:qwenacp)
failed at spec synthesis: _materialize_harness_launcher_file put the harness id
straight into the agent `name`, and the agent-name validator rejects the colon
("name must match [a-zA-Z0-9_-]+"). The generic ACP harness (#2152) intends
acp:<slug> as the run-time addressing form (canonicalizes to `acp`, command
resolved from the acp: config block at spawn), but this no-AGENT launcher path
was missed.
Fix: keep the FULL acp:<slug> in executor.harness (canonicalize_harness drops
the slug to bare `acp`, which would lose the agent selection), and sanitize the
colon (":" -> "-") for the agent NAME and temp filename only, which must be
[a-zA-Z0-9_-]+ / path-safe. Non-acp harnesses are unchanged: name still uses the
raw input (claude -> "claude"), executor/filename still canonicalize (claude ->
claude-sdk, kimi alias -> kimi). Added an acp:<slug> launcher test; existing
launcher tests green.
* feat(web): auto-fill a configurable default base branch for new worktrees
When naming a new worktree branch in the new-session composer, users had
to type the base branch every time. Add a "Default base branch" setting so
the base-branch field pre-fills automatically.
- New Settings › Git section with a "Default base branch" text input,
persisted per-device in localStorage (omnigent:default-base-branch),
mirroring the existing appearance/font preference modules. Blank = no
auto-fill (worktrees branch off current HEAD, unchanged behavior).
- The composer seeds its base-branch state from the stored default, so the
field appears pre-filled once a new branch name is entered.
Also reset the module-level landingDraft in the flow test's beforeEach to
stop composer state leaking across tests.
Co-authored-by: Isaac
* fix(web): stop stale base-branch auto-fill after clearing the default
The landing composer snapshots its fields into a module-level draft on
unmount. An auto-filled default base branch was captured in that snapshot
and, on remount, took precedence over the live setting — so clearing (or
changing) the Default base branch in Settings still left the old value
auto-filling the field.
Track whether the user actually edited the base branch. The draft now only
pins the base branch on a real edit; otherwise the field mirrors the current
default, so clearing or changing the setting takes effect immediately. A
user-typed base still survives a nav-away.
Co-authored-by: Isaac
* fix(web): refresh base-branch default when the worktree popover reopens
Changing the Default base branch in Settings and returning to the composer
didn't auto-fill until a full refresh: a same-tab settings change fires no
`storage` event, and the composer's mount-time seed can hold a stale value.
Re-read the configured default when the worktree popover opens, unless the
user has hand-typed a base. The field now reflects the current setting the
next time it's opened, without a refresh; a user-typed base is left intact.
Co-authored-by: Isaac
* fix(web): live-follow the base-branch default via a change subscription
The popover-open re-read missed same-tab settings changes when the composer
stayed mounted. Replace it with an explicit subscription: writeDefaultBaseBranch
announces same-tab changes on a custom event (the `storage` event only fires
in other tabs), and the composer follows the default while the user hasn't
taken over the field.
Encodes four rules, each covered by a test:
1. Nothing set → no auto-fill; the user types freely without side effects.
2. User already filled a base → a later setting change leaves it untouched.
3. Branch named, base empty → a setting change auto-fills it, still editable.
4. Once the user edits the base (even to blank), the default never touches it.
Co-authored-by: Isaac
* fix(web): re-seed the base branch from the default on each dropdown open
Simplify the model: the base-branch field is re-seeded from the Settings ›
Git default (or blank) every time the worktree dropdown opens, and never
remembers a value typed in a previous open. Within one open the user can
override it freely; reopening discards that and shows the setting again.
Drops the persisted baseBranch/baseBranchEdited draft state and the same-tab
change subscription — reading on open covers every case (change, clear, or
prior edit) without stale-state pitfalls.
Co-authored-by: Isaac
* fix(web): tie base-branch auto-fill to the branch-name lifecycle
Seed the base branch from the Settings › Git default when the user names a
new-worktree branch, then leave it to the user: any edit — including
explicitly clearing the field — stands, even when the worktree dropdown is
reopened. Clearing the branch name (starting the worktree over) re-arms the
auto-fill, so the next named branch seeds fresh from the current default.
Previously the field re-seeded on every dropdown open, so a base the user
had cleared came back on reopen.
Co-authored-by: Isaac
* fix(web): normalize the default base branch on read
Trim on read and treat a whitespace-only value as unset, so a hand-edited or
stale localStorage entry can't display un-normalized. Everything the app
writes is already trimmed; this closes the gap for values that bypassed the
writer. Addresses a non-blocking note from the automated PR review.
Co-authored-by: Isaac
Pytest (misc) had grown to ~9:52 wall, ~2x the next-slowest group and
the critical path of the matrix. Root cause (from JUnit + per-worker
progress artifacts of a main run): misc runs --dist=loadfile, which
pins a whole file to one worker, and tests/runner/test_app_sessions_native.py
alone (~506 cpu-seconds, 249 tests) set the wall floor -- 507 of 508s
on the critical worker while the other 7 finished in 264-310s and idled.
cpu breakdown of misc: tests/runner 36%, tests/stores 32%, tests/db 15%
(= 83%). The top-level *_native* coding-agent files everyone suspects
were only ~8% combined.
Carve tests/runner (runner-app) and tests/stores (stores) into their
own worksteal shards; misc ignores both and also gains worksteal so the
biggest remaining file can't re-pin a worker as the catch-all grows.
Both dirs' conftests are function-scoped, so fanning a file across
workers is safe. tests/db stays in misc (it's split by the databricks
marker, not by path).
Collection partitions exactly (-m "not databricks"):
misc_after 4425 + runner 1125 + stores 429 = 5979 = misc_before.
Also add the two new shard names to merge-ready/required.sh so they
gate. NOTE: required.sh is a generated file (replaced on internal sync)
-- the generator source needs the same two names or this hand-edit is
reverted on the next sync.
Co-authored-by: Isaac
* feat(cli): add `omnigent debug logs` command
Exposes runner, server, and CLI diagnostic log files via the debug
subgroup so operators can inspect them without navigating the
~/.omnigent/logs/ directory manually.
--type [runner|server|cli] which log category (default: runner)
--list list files with sizes and timestamps
-n / --lines N tail last N lines (0 = whole file)
-f / --follow stream in real-time (tail -f)
* feat(cli): filter runner logs by session id
Embeds the session id in each runner log filename
(runner-conv_abc123-<random>.log) so all relaunches for a session are
discoverable. Adds --session SESSION_ID to `omnigent debug logs` to
show all log files for a session oldest-first.
* fix(cli): address Polly review on debug logs command
- Separate runner into two types: runner (logs/runner/, local CLI) and
host-runner (logs/host-runner/, host daemon) — fixes the blocking bug
where the default type pointed at the wrong directory
- Broaden server glob to *server*.log to cover both server-*.log
(omnigent run) and local-server-*.log (background daemon)
- Scope --session to --type host-runner only (where session ids are
embedded in filenames)
- Guard --follow on Windows with IS_WINDOWS check
- Add min=0 bound to --lines to reject negative values
The kubernetes launcher forced kubernetes.io/arch: amd64 onto every
runner Pod because the host image used to publish amd64-only. The image
is now a multi-arch manifest list (amd64 + arm64), so the hard pin only
blocks scheduling on arm64 nodes. Keep amd64 as the default — existing
deployments keep their placement — but merge it first so an operator
kubernetes.io/arch entry in sandbox.kubernetes.node_selector wins.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(harness-bench): bind any registered harness passed by name
The bench could only probe an official profile (the 4 SDK harnesses +
auto-derived native-tui) or a dotted module:attr BenchProfile reference. A
harness registered in the omnigent registry but neither official nor native-tui
-- the in-repo generic ACP harness (`acp`, ACP_SUBPROCESS), or an entry-point
community plugin (`rovo`/`rovo-cli` from omnigent-rovo) -- KeyError'd on
resolve_profile, so `--harness acp` / `--harness rovo` could not run.
Add a registry fallback to resolve_profile: after the official + reference
checks, derive a BenchProfile for any harness in the omnigent registry
(_registry_profile in manifest.py). It resolves aliases (rovo -> rovo-cli),
keys off harness_modules() so it covers plugins that declare no capabilities
entry, maps integration_mode -> transport family (SDK/CLI/ACP subprocess ->
sdk-inproc family = the existing drivers; NATIVE_TUI -> native-tui), and
skip-gates on the harness's install-spec binary when present (rovo -> acli).
No new transport driver: an ACP harness registers as an omnigent agent
(config.harness=acp:<slug>) and runs on the existing SDK-wrap drivers. Both
harnesses are OWN_AUTH, so they run only where their vendor binary is installed
+ authed, and skip cleanly otherwise (verified live: rovo skips on missing
`acli`). tool_calling/policy_deny stay `·` for ACP (agent runs its own tools /
gates via session/request_permission) -- the same documented gap as native.
Tests: resolve_profile binds acp (sdk-inproc) and rovo/rovo-cli (alias, acli
gate); unknown still KeyErrors; plugin cases skip if omnigent-rovo absent.
Offline suite 71 passed / 18 skipped, ruff clean.
* fix(harness-bench): address review — NATIVE_SERVER refusal, own-auth model, ACP-login SKIP
Three fixes from PR review + a live rovo run:
1. (blocking, Polly) A MODELED integration_mode the bench has no driver for
(NATIVE_SERVER, e.g. opencode-native) was silently degrading to the
sdk-inproc default via `.get(mode, "sdk-inproc")` — binding a vendor-server
harness to the wrong driver and dropping its skip-gate. _registry_profile now
distinguishes: no caps (unmodeled plugin) -> assume SDK family; a modeled
mode NOT in the transport map -> return None so resolve_profile KeyErrors
(honest "unrunnable" rather than a wrong profile). resolve_profile("opencode
-native") KeyErrors again.
2. A live rovo run (acli absent) reported `!!✓>✗` DRIFT: the ACP-session /
vendor-login failure ("Ensure `acli` is installed and you are logged in",
"AcpProcessExited", "ACP subprocess/session") wasn't an infra marker, so it
read as a real UNSUPPORTED against the SUPPORTED declaration. Added those
markers + a reason so an own-auth harness with no vendor login SKIPs (env
gap), never drifts.
3. Registry profiles stamped a databricks-* placeholder model even for own-auth
harnesses (rovo/acp), which is misleading — the runner drops the gateway
model for them. Now: gateway-credential harness -> the databricks default;
own-auth or capless -> empty model (the harness owns it).
Tests: NATIVE_SERVER refusal; a plugin-independent happy-path (fake registered
CLI harness via monkeypatch) so the fallback's positive path isn't skip-gated
away in CI; rovo model=="" assertion. Offline suite 73 passed / 18 skipped.
* fix(harness-bench): registry profiles need a valid model to register
My previous "empty model for own-auth" change broke agent registration: the
omnigent executor spec mandates a model (spec/omnigent.py: "executor.type=
'omnigent' requires a model"), so model="" -> 400 "llm.model must be present
when llm block is present" on register_agent. Seen live: rovo got past auth +
skip-gate into provisioning, then failed registration.
A model is always required for registration, so stamp the databricks default in
all cases. For an own-auth harness it is inert: the generic ACP harness drops
databricks-* models (workflow.py::_build_acp_spawn_env), and rovo has no
spawn-env builder + reads HARNESS_ROVO_MODEL directly from env (which the runner
never sets for it), so rovo gets no model and lets Rovo Dev pick its own default
at session/new. The placeholder satisfies registration and never reaches acli.
Tests updated to assert a non-empty model (registration invariant) rather than
empty.
* feat(harness-bench): bind acp:<slug> ids to a specific ACP agent
`acp:<slug>` is a first-class omnigent harness id — the base `acp` harness is
registered and the slug selects a user-configured ACP agent at spawn (resolved
from the ~/.omnigent `acp:` block). The registry fallback now recognizes it:
look up caps/module/install-spec by the base `acp`, but keep the full `acp:<slug>`
as the profile harness so `config.harness=acp:<slug>` reaches the runner, and
sanitize the colon in the env-prefix/marker stem (acp:qwen -> HARNESS_ACP_QWEN_).
An empty slug ("acp:") is refused.
Lets `--harness acp:qwen` bind to a specific ACP agent for a live turn (qwen is
installed + authed), vs the bare `acp` which needs HARNESS_ACP_COMMAND. Test
added. Offline suite 73 passed / 18 skipped.
* fix(harness-bench): sanitize colon in bench agent name for acp:<slug>
The bench built its agent name as bench-<harness>, but an acp:<slug> harness id
has a colon, which the agent-name validator rejects ([a-zA-Z0-9_-]+). So a
--harness acp:qwen run would 400 at registration. Replace ":" with "-" in the
NAME only (bench-acp-qwen); config.harness keeps the real acp:<slug> id so the
runner still resolves the right ACP agent at spawn.
* chore: remove dead cost_advisor / cost_judge runner-side feature
No agent YAML ever used `executor.config.cost_optimize:`, making the
entire runner-side per-turn cost advisor a dead code path. The feature
was superseded by the server-side smart routing (OMNIGENT_SMART_ROUTING).
Deleted:
- omnigent/runner/cost_advisor.py
- omnigent/runner/cost_judge.py
- tests/runner/test_cost_advisor.py
- tests/runner/test_cost_judge.py
- tests/e2e/test_polly_cost_advisor_e2e.py
Cleaned up:
- omnigent/runner/app.py: remove AdvisorTurnResult import, _fetch_cost_control_mode_override,
_merge_advisor_note, _apply_advisor_to_body, _session_advisor_applied_model,
_run_turn_advisor, _emit_routing_decision, _apply_advisor_for_turn,
_advisor_spec_for_session, and both call sites in the turn paths.
- omnigent/spec/parser.py: remove cost_optimize from _STRUCTURED_EXECUTOR_CONFIG_KEYS.
- omnigent/cost_plan.py: strip to just COST_CONTROL_LABEL_NAMESPACE and
reserved_cost_control_keys (still used by sessions.py for the label
namespace guard); remove all advisor-only symbols.
- tests/runner/test_app_sessions_native.py: remove advisor integration tests.
* fix(ci): remove test_cost_plan.py, fix test_sessions_cost_labels imports
* fix: revert accidental Sidebar.tsx change; fix dangling cost_advisor doc refs
* chore: regenerate openapi.json for updated RoutingDecisionData docstring
* chore: remove tier from RoutingDecisionData and full frontend pipeline
* fix: re-delete cost_advisor.py (re-appeared in working tree)
* fix(test): remove routing_decision.tier assertion after field removal
## Related issue
N/A
## Summary
- Modals (e.g. Create custom agent) are `position: fixed`, centered with
`top-1/2 -translate-y-1/2`, and capped at `max-h-[85vh]`. On the iOS
shell the native app keeps the WKWebView layout viewport full-height
when the soft keyboard opens (`.ignoresSafeArea(.keyboard)`), so `vh`
and `50%` both resolve against the whole screen — the modal's lower half
(and any focused input) ends up hidden behind the keyboard.
- Fix in the shared `DialogContent` primitive so every modal benefits at
once: on the iOS shell only, an inline style pins the centering origin
and height cap to the keyboard-aware `--omnigent-viewport-height` (which
`useIOSViewportLock` already publishes on :root from
`visualViewport.height`), less the safe-area insets and a small margin.
The modal now shrinks and its inner content scrolls; nothing extends
behind the keyboard, notch, or home indicator.
- Inline style is deliberate: the several dialogs that pass their own
`max-h-[85vh]` would otherwise win, since `cn`'s twMerge keeps the
caller's class. Inline beats classes, so the keyboard-aware cap governs.
- Gated on `isIOSShell()` and carries a `100lvh` fallback, so web,
Android, and Electron keep the existing `85vh` / centered behavior
unchanged.
## Test Plan
- `npx tsc -b` — clean.
- `npx vitest run` on the new `dialog.test.tsx` plus dialog-consuming
suites (`PoliciesPage`, `NewChatDialog`) — 143 passing, including new
coverage that the iOS inline cap (top + maxHeight from
`--omnigent-viewport-height`) is applied inside the iOS shell and absent
off it.
- `src/components/ui` is excluded from oxlint (vendored shadcn), so no
lint applies to the changed primitive; prettier run on both files.
## Type of change
- [x] Bug fix
- [ ] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
The gating logic (iOS-shell-only inline cap wired to the keyboard-aware
viewport var) has unit coverage in the new dialog.test.tsx, and existing
dialog-consuming suites confirm no regression off iOS. The actual
keyboard-overlap behavior is WKWebView-specific and can't be reproduced
in jsdom (no soft keyboard / visualViewport resize), so final visual
confirmation on the iOS app — opening a tall modal with the keyboard up
and checking it stays fully on screen and scrolls internally — is still
recommended before release.
## Related issue
N/A
## Summary
- Add `.github/workflows/electron-build.yml`, a `workflow_dispatch`-only
pipeline that packages the Electron desktop shell (`web/electron`) for
Linux and Windows. A 2-way matrix builds each platform on its own native
runner (`ubuntu-latest` → AppImage + .deb, `windows-latest` → NSIS .exe)
since electron-builder does not reliably cross-compile installers, and
uploads the distributables as workflow artifacts (14-day retention).
- Reuses the repo's `./.github/actions/setup-node` composite action (pinned
to Node 22 per web/electron/README.md, npm cache keyed on the electron
lockfile), runs `npm ci` then `npm run build:linux`/`build:win`. Builds
are unsigned (`CSC_IDENTITY_AUTO_DISCOVERY=false` so a missing cert
doesn't fail the build) and never publish; macOS is omitted (its
signed/notarized build lives elsewhere). `fail-fast: false` so one
platform breaking still yields the other's installers.
- Fix `web/electron/package.json` metadata the Linux `.deb` build requires:
add `homepage`, expand `author` from a bare string to `{ name, email }`,
and set `linux.maintainer`. Without these, electron-builder's fpm packager
aborts the `.deb` target ("specify project homepage / author email /
.deb maintainer") — a pre-existing config gap the new Linux job would hit.
## Test Plan
- `actionlint .github/workflows/electron-build.yml` — clean.
- Validated the workflow YAML and package.json parse (yaml.safe_load /
JSON.parse).
- Locally in `web/electron`: `npm ci` resolves cleanly, and
`npm run build:linux -- --publish never` produces BOTH
`Omnigent-<ver>-<arch>.AppImage` and
`omnigent-desktop-electron_<ver>_<arch>.deb` after the metadata fix
(before it, the .deb target failed as described above). Confirmed the
workflow's artifact globs (`*.AppImage`, `*.deb`, `*.exe`) match the
real output names.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [x] Test / CI
- [ ] Breaking change
## Test coverage
- [ ] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
CI workflow + build-config change with no unit-testable surface; verified
by linting the workflow (actionlint) and by running the Linux build locally
end-to-end, which produced both the AppImage and .deb and proved the
package.json metadata fix. The Windows job could not be exercised locally
(macOS host), but it uses the same already-working `build:win` (nsis) script
on `windows-latest`; the first manual run from the Actions tab will confirm
it end-to-end.
## Related issue
N/A
## Summary
Three related fixes to the mobile / iOS chat surface:
- **Message copy button now works on mobile.** The user and assistant
bubble copy actions called `navigator.clipboard.writeText` directly and
silently no-op'd when it was absent (the iOS webview / non-secure
origins). They now route through the shared `copyText()` helper, which
falls back to an `execCommand` textarea copy. Deduplicated the two inline
handlers into a shared `useCopyMessage` hook.
- **Visual confirmation on copy.** On a mobile viewport the copy action
fires a "Copied to clipboard" toast in addition to the inline check icon
(which is easy to miss on a phone). Desktop is unchanged (icon + tooltip).
- **Native Chat/Terminal bar no longer disappears after copy.** The
`execCommand` fallback focuses a hidden textarea, which the iOS
keyboard-visible check mistook for the keyboard opening and hid the
native Liquid Glass bar — and WebKit doesn't reliably fire `focusout`
when the focused node is removed, so it stayed hidden. The helper textarea
is now marked `data-clipboard-helper` and excluded from editable-focus
detection.
- **iOS Chat/Terminal bar no longer overlaps the composer status line.**
The chat-view bottom spacer reserved 1rem less than the bar's footprint,
so the bar rode up over the host / harness / context-ring row. It now
reserves the full footprint (iOS-only, chat-view-only).
## Test Plan
- `npx tsc -b` — clean.
- `npx oxlint` on changed files — no new findings.
- `npx vitest run` on the affected suites (clipboard, keyboard-inset hook,
ChatPage user bubble) — 23 passing, including new coverage:
- clipboard-helper textarea is not treated as editable focus, while a
real textarea is;
- copy falls back to `execCommand` when the async clipboard is absent;
- a mobile viewport fires the copy toast;
- the fallback textarea carries the `data-clipboard-helper` marker.
- CSS + WKWebView-specific behavior verified by inspecting the Vite-served
compiled CSS; on-device visual confirmation still pending (see notes).
## Type of change
- [x] Bug fix
- [ ] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
The clipboard, keyboard-inset, and copy-button paths have unit coverage
(23 tests, listed in the Test Plan). The two behaviors that can't be
exercised in jsdom — the iOS status-line/bar overlap (CSS var math) and the
real WKWebview clipboard/native-bar interaction — were verified by reading
the Vite-served compiled CSS and by reasoning from the shell's focus/keyboard
hooks; final on-device visual confirmation in the iOS app is still
recommended before release.
## Related issue
N/A
## Summary
- Add a `--trust-lan-origins` flag to omnidev (the dev-pod supervisor) so a
phone or tablet on the same network can use the UI end to end when Vite is
bound with `--vite-host 0.0.0.0`. A device loads the UI at
`http://<lan-ip>:<vite-port>`, so its browser stamps that non-loopback
address as the `Origin` on every request. The pod's backend runs in
single-user local mode, where the origin guard trusts only loopback
origins — so multipart uploads get a 403 and the WebSocket stream is
refused. The flag closes that gap.
- New `lan.rs` enumerates this machine's LAN IPv4 addresses (private +
link-local, dropping loopback/public/broadcast/multicast via the
`if-addrs` crate) and builds the matching `http://<ip>:<vite-port>`
origins. They're fed to the server through its own exact-match allowlist
env var `OMNIGENT_WS_ALLOWED_ORIGINS`, merged with any value the developer
already exports (order-preserving, deduped). It stays exact-match — only
the enumerated origins are trusted, nothing is disabled — so it covers
both the upload guard and the WS handshake without weakening CSRF/CSWSH
protection. Off by default; a no-op unless the flag is passed.
- The trusted origins are printed in the combined log at startup; if the
flag is set but no LAN interface is found, a warning says so rather than
silently no-op'ing later.
- README documents the flag and a "Testing from a phone or tablet" section.
## Test Plan
- `cargo build`, `cargo test` (22 passing, incl. new unit tests for LAN IPv4
filtering, origin construction, and the env-merge onto an inherited
allowlist), `cargo clippy --all-targets` (clean), `cargo fmt --check`
(clean).
- Verified the real `if-addrs` enumeration on this machine produces the
expected `http://<ip>:5173` origins for the host's private/link-local
interfaces (loopback/public dropped).
- `--help` renders the new flag; `pre-commit` passed on the changed files.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Origin filtering, construction, and the allowlist env-merge have unit tests
(cargo test, 22 passing). The real interface enumeration and the
device-in-browser flow can't be asserted in a unit test, so they were
verified manually: the `if-addrs` call was run on this host and produced the
correct origins, and the resulting `OMNIGENT_WS_ALLOWED_ORIGINS` value was
confirmed to merge with an inherited value. Final confirmation from an actual
LAN device (upload + live stream over `--vite-host 0.0.0.0
--trust-lan-origins`) is recommended but not automatable in CI.
* ci(benchmark): default items-per-session to 200
Raise the seeded items-per-session default from 50 to 200 for a denser
per-session corpus. Update both the workflow_dispatch input default and
the ITEMS env fallback used by scheduled runs so manual and nightly runs
agree on the default.
Co-authored-by: Isaac
* ci(benchmark): rename workflow to "Benchmark", clarify iterations label
Rename the workflow from "Performance Benchmark" to "Benchmark" and reword
the iterations input label to "Requests per run" so it matches how the
harness drives the journeys.
Co-authored-by: Isaac
* ci(benchmark): cap full-turn journeys' iterations; HTTP default 200->100
The nightly benchmark timed out at 30 min inside the first full-turn
journey. `--iterations` applied uniformly, but the four runner journeys
cost ~1s+ per op (vs. ~ms for the HTTP journeys), so 200 iterations x 3
runs was ~20 min for `session_cold_start` alone.
Add a `max_iterations` field to `Journey` that clamps `--iterations` down
per journey (never up), and cap the four full-turn journeys at 5 samples
per run — `--runs` provides the repeats. Splitting samples across runs
vs. iterations doesn't change accumulation (all runs share one env), so a
small per-run count is the lever; it also keeps the cold-start session
drift (~2 ms/turn, sessions accumulate within a run) negligible. Lower the
HTTP iterations default 200 -> 100 to match run.py's own default.
The full runner suite now finishes in ~2.4 min locally (was 20+ min),
with meaningful cross-run percentiles.
Co-authored-by: Isaac
The omnigent-site PR was titled `docs: document omnigent-ai/omnigent#N`,
but the source PR number already appears twice in the body, so the title
carried no information. Title it after the actual docs change instead.
The doc-drafter now emits a `DOC_PR_TITLE:` line summarizing what the docs
cover; the workflow sanitizes it (untrusted LLM output) and falls back to
the source PR title, then the old `document #N` form, so a missing line
degrades gracefully. Also pass `--title` on the `gh pr edit` update path,
which previously never refreshed a re-draft's title.
Co-authored-by: Isaac
* fix(web): align project picker menu rows left with uniform height
The sidebar "Add to / Move to project" submenu had inconsistent rows: the
search box used px-2 py-1.5 while the project rows fell back to the
DropdownMenuItem default (px-1.5 py-1), so rows were indented differently
and slightly shorter than the search input. Give every row (project names,
"Create new project", "Remove from …", and the inline new-project input) a
uniform px-2 py-1 so they share one left edge and height.
Co-authored-by: Isaac
* style(web): fix prettier formatting in Sidebar.tsx
Restore the canonical multi-line union type on the drag-start cast that a
prior edit had collapsed onto one line, which prettier --check rejected.
Co-authored-by: Isaac
* feat(smart-routing): replace RoutingDecisionChip with collapsible RoutingDecisionCard
When auto-routing fires at first-message time (agent spec has no explicit
model), the UI previously showed a minimal muted chip. Replace it with a
collapsible card that mirrors the SmartRoutingCard style: same container
border, a model+tier pill, rationale text, and an expandable raw verdict
JSON block behind a chevron.
The chip remains exported for any downstream consumers but ChatPage now
renders RoutingDecisionCard for routing_decision bubbles.
* feat(smart-routing): mirror sub-agent routing decisions into the parent session
When sys_session_send spawns a child session without an explicit model,
the server routes it and emits a routing_decision item — but only into
the child's transcript. Orchestrators seeing the main session had no
visibility into which model was chosen for each sub-agent.
Changes:
- Add optional `agent` field to RoutingDecisionData so parent-mirrored
items carry the sub-agent name.
- _emit_server_routing_decision accepts a keyword `agent` arg.
- Both routing paths (_forward_event_to_runner SDK path, native terminal
path) now also emit into parent_conversation_id when _parent_routing_on,
passing the child's agent_name as the agent label.
- Thread `agent` through the frontend pipeline: RoutingDecision event,
RoutingDecisionBlock, RoutingDecisionItem, SSE reducer, blockStream,
itemsToBlocks, renderItems bubble, and RoutingDecisionCard.
- RoutingDecisionCard shows the agent name as the row label (replacing
"Session") when rendering a parent-mirrored decision.
* fix(smart-routing): remove tier label from RoutingDecisionCard pill
* chore: regenerate openapi.json for RoutingDecisionData.agent field
* refactor(db): enforce scoped uniqueness in app code, drop partial indexes
MySQL has no partial (WHERE-predicated) indexes. The four scoped indexes on
agents/policies/conversations leaned on dialect-scoped sqlite_where /
postgresql_where kwargs that MySQL silently dropped, yielding full unique
indexes that over-restrict on MySQL (session agents/policies could not reuse
names there). Replace them with plain indexes that behave identically on
SQLite, Postgres, and MySQL:
- ix_conversations_parent_title_unique: kept UNIQUE, predicate dropped. The
WHERE (parent_conversation_id IS NOT NULL) was redundant with NULL-distinct
semantics, so top-level conversations stay exempt. No behavior change.
- idx_conversations_parent: non-unique perf index, predicate dropped. Now
indexes every parented row; same query plan for child-session listing.
- ix_agents_template_name -> ix_agents_name (plain). Template-name uniqueness
moves to the store (SqlAlchemyAgentStore.create gains a workspace-scoped
pre-insert check; agents had no app-level check before).
- ix_policies_default_name_cksum -> ix_policies_name_cksum (plain). Default-
name uniqueness was already enforced in the store (add_default /
update_default); the index was just a backstop.
Migration z5a2b3c4d5e6 (index-only, off z4a2b3c4d5e6): drops the partials and
creates the plain replacements; downgrade restores the partials.
Co-authored-by: Isaac
* refactor(db): include kind in ix_agents_name for template lookups
Session agents can now share names, so (workspace_id, name) alone matches a
template plus every same-named session copy. Add kind to ix_agents_name ->
(workspace_id, name, kind, id) so get_by_name and the create() uniqueness
check seek straight to the template row instead of scanning session copies.
Co-authored-by: Isaac
MySQL's InnoDB does not compress TEXT/BLOB by default and SQLite never
does, so per-conversation JSON/text columns that PostgreSQL would TOAST
sat uncompressed on the other two backends. Compress them in the
application layer instead, for a uniform on-disk size across all three.
Add omnigent/db/compression.py: a `CompressedText` SQLAlchemy
TypeDecorator (LargeBinary impl) that zstd-compresses on write and
decompresses on read, transparent at the ORM boundary so the stores keep
reading/writing `str`. Values carry a NUL-sentinel + codec frame; sub-64B
payloads are stored uncompressed to avoid framing inflation. Rows written
before migration are unframed and decode unchanged (and on SQLite arrive
as `str`), so no backfill is needed — each re-frames on its next write.
Apply it to six columns never queried in SQL: conversations.session_usage
/ session_state / terminal_launch_args, comments.body / anchor_content,
and agents.description. Migration z4a2b3c4d5e6 flips them TEXT -> binary
via batch alter (PostgreSQL casts with convert_to/convert_from); the
downgrade decompresses every row before restoring TEXT.
Add zstandard as a dependency. Codec + migration + type-change tests
included; existing store suites pass unchanged.
Co-authored-by: Isaac
Projects are a "My sessions"-only surface — filing a session into a
project is owner-only, so the sidebar renders project folders only on
"My sessions". But the two backend surfaces that drive the project view
filtered by any access grant rather than ownership, so a session someone
shared with you, if it carried a project label, surfaced inside its
project folder under "My sessions" instead of under "Shared with me".
Scope both project surfaces to owner-level grants:
- list_projects / GET /sessions/projects: the folder names now come only
from projects that contain a session the viewer owns.
- list_conversations / GET /sessions?project=X: the sessions inside a
folder are now owner-scoped too.
The flat list (project=None) and Unfiled (project="") stay unscoped, so
shared sessions still surface for the "Shared with me" tab.
Co-authored-by: Isaac
Live instrumentation (temporary, reverted) proved the native Policy DENY chain
works end to end: the claude PreToolUse evaluate-policy hook fires, reaches
/policies/evaluate, the session-attached CEL deny loads, the server returns
POLICY_ACTION_DENY with our reason and publishes response.policy_denied. The
prior "hook not wired / ap_server_url not threaded" diagnosis was WRONG — it
came from searching $HOME instead of the real bridge root
(/var/folders/.../omnigent-502/claude-native), which HAS a valid
permission_hook.json.
The real bench bug was a reader race, and a first grace-window fix was still
flaky (passed 1 run, SKIPPED the next). Root cause: response.policy_denied is
published when the PreToolUse hook evaluates, and its timing relative to the
turn's output_item.done is highly variable — it can land after a SECOND
output_item.done and the session settle. A fixed grace window measured from the
first terminal event races that.
Deterministic fix: on a deny turn the reader no longer stops on the turn's
terminal events at all — it reads until it sees response.policy_denied (returns
immediately) or the caller signals stop after a generous observe budget
(_DENY_OBSERVE_S=30s). A real deny exits early; only a genuine no-deny waits the
budget then SKIPs. Non-deny turns are unchanged (stop on the terminal event).
Live: claude-native Policy DENY now SUPPORTED across repeated solo runs (was
flaky, then ·). Verdict semantics: SUPPORTED = "the tool call was routed through
policy and a DENY verdict returned"; vendor hard-enforcement (tool actually
blocked) is a separate axis noted in the driver. Offline suite 69 passed /
18 skipped; added a test for a policy_denied that lands after the terminal event.
Re-lands the benchmark harness (reverted in #2200) without the manual
seed-schema drift guard that caused the original merge friction.
The harness: HTTP/API journeys (list/create/get session, load history, search)
and full-turn journeys (session_cold_start, warm_turn, time_to_first_token,
interrupt) driven through server + runner + a zero-latency mock LLM, all via
the in-process openai-agents SDK harness. Seeds a deterministic corpus via the
store API; SQLite + Postgres backend matrix; nightly workflow uploads a
versioned JSON report for a workspace Databricks notebook to consume.
Drops the SEED_SCHEMA_REVISION constant, scripts/check_benchmark_seed_schema.py,
and the pre-commit hook. That guard was a false-positive tripwire — it failed on
every migration (even ones not touching the seed's tables) and its "fix" was
always just bumping a string; the seed never actually broke. Instead seed() now
reads the Alembic head at runtime (_get_head_db_revision) into the corpus reuse
marker, so an old corpus auto-reseeds with zero maintenance. The real invariant
— that seeding still works against the current schema — is covered by
test_seed_creates_listable_corpus, which seeds through the store (migrations run
to head on init) and so can't false-positive.
Verified: 8 smoke tests pass; seed auto-picked up the new head (x1a2b3c4d5e6)
with no code change; --print-head intact for the CI seed-cache key; ruff, mypy,
pre-commit clean.
Co-authored-by: Isaac
## Related issue
N/A
## Summary
- Add install-management subcommands to omnidev, for people who *run*
omnigent (installed from git via `uv tool install`) rather than develop
it. This fills a real gap: omnigent's own update notice only works for
PyPI-wheel installs and skips git installs, so a git-installed omnigent
never learns it is out of date.
- `omnidev install` — `uv tool install` from git, defaulting to the
`databricks` extra and `main`; `--ref`/`--extra`/`--no-default-extra`/
`--repo` override and persist to `~/.config/omnidev/install.toml`.
- `omnidev update` — reinstall the latest of the tracked ref/extras
(`--reinstall`, required for a moving git ref).
- `omnidev check` — the shell-hook primitive: reads a cache, refreshes
it detached when >24h stale (never blocks the shell), and on an
available update prints a notice and, on a TTY, prompts to update in
the foreground. A declined commit isn't re-nagged.
- `omnidev refresh` — the background `git ls-remote` probe.
- `omnidev shell-hook` — emits the `eval "$(omnidev shell-hook)"` snippet.
- These subcommands need no checkout and dispatch before repo-root
discovery, so they run from any directory; bare `omnidev` still launches
the pod supervisor. Installing from git builds the web UI from source, so
`install` fails early if `uv`/`npm` is missing.
- Lighten pod isolation: only omnigent's own state (`OMNIGENT_DATA_DIR`,
`OMNIGENT_DATABASE_URI`, `OMNIGENT_URL`) is isolated per pod. The pod now
inherits the real `HOME`, credentials, config, and uv/npm caches — which
the agents omnigent runs need — instead of the hermetic
`HOME`/`XDG_*`/`TMPDIR` sandbox that cut them off.
## Test Plan
- `cargo build`, `cargo build --release`, `cargo clippy --all-targets`, and
`cargo fmt` all clean.
- `cargo test` passes 13 tests (7 new): install-spec builder for default /
no-extras / custom ref+extras, install-config round-trip, missing-config,
update-availability logic including decline suppression, and the 24h
staleness window.
- Manually verified from a scratch dir with no git repo that `omnidev
check`, `shell-hook`, etc. run without a "missing checkout" error, while
bare `omnidev` still errors as expected; confirmed the CLI surface
(`--help`, `install --help`, `shell-hook` output).
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
The network- and install-driving paths (`uv tool install`, `git
ls-remote`, reading the installed tool's `direct_url.json`, the detached
refresh, and the TTY prompt) can't run in unit tests, so they were verified
manually. Pure logic — spec building, config round-trip, update-
availability and staleness decisions — is covered by `tests/install_mgmt.rs`.
* feat(db): include primary-key columns in every secondary index
The storage standard requires every index to contain the table's
primary-key columns. Each table's PK now leads with workspace_id (the
tenant partition key) then the entity id column(s), and every store
query filters workspace_id.
Rebuild each secondary index accordingly:
- Non-unique indexes lead with workspace_id and trail the remaining PK
id-columns, which double as the keyset tiebreaker / covering column the
queries already use.
- Unique indexes/constraints get workspace_id prepended only (appending
the entity id would make uniqueness vacuous), becoming per-workspace
unique. uq_hosts_token_hash is safe because resolve_launch_token
already filters workspace_id + token_hash.
Two orders are query-driven, not mechanical:
ix_session_permissions_conversation_id and
ix_conversation_items_response_id place the filtered PK column right
after workspace_id. ix_comments_created_at is dropped — no query sorts
comments globally by created_at (always conversation-scoped).
MySQL note: MySQL has no partial index, so the WHERE on the partial
unique indexes is dropped there and the unique spans all rows (more
restrictive; acceptable). Emulating partial-unique on MySQL is left to
the MySQL support work.
Co-authored-by: Isaac
* fix(comments): order list_for_conversation by (created_at, id)
created_at is seconds-granular, so comments added in the same second tie
under ORDER BY created_at and the listing order fell back to index scan
order. Adding id to every secondary index changed that implicit tiebreak
(rowid → id), surfacing the latent non-determinism. Sort by (created_at,
id) for a stable, deterministic order, matching the keyset convention
used by the other stores. The chronological-order test now advances the
clock per add so its "oldest first" assertion no longer hinges on the
same-second tiebreak.
Co-authored-by: Isaac
* feat(db): fold created_at into ix_comments_conversation_id
list_for_conversation now sorts by (created_at, id), so make the index
serve it: (workspace_id, conversation_id, created_at, id). This is
index-ordered for WHERE workspace_id + conversation_id ORDER BY
created_at, id and still contains the full PK. Re-adding the old bare
ix_comments_created_at would not help — the query filters conversation_id
first, so a created_at-leading index cannot serve it.
Co-authored-by: Isaac
The interrupt test awaited an already-unblocked task through
asyncio.wait_for(int_task, timeout=15.0). Under the misc shard's 8-worker
CPU contention the event loop can be starved past 15s, so the wall-clock
timer cancels the await even though the interrupt already returned 204 —
the traceback showed `int_task` finished with a 204 while wait_for raised
TimeoutError. This reddened the misc shard on main intermittently.
Drop the wall-clock timers: await the interrupt task and the post_seen /
fwd_seen events directly. The task is unblocked one line earlier
(fwd_gate.set()), so there is no correct reason to race it against a wall
clock; pytest's global --timeout=300 remains the genuine-hang backstop.
Widening the timeout only lowers the odds — a starvation spike past the
budget still trips it; plain await removes the race entirely.
Verified 5/5 green under all-cores-pegged + `-n 8` stress that reliably
reproduced the TimeoutError beforehand.
Co-authored-by: Isaac
* fix(tools): make in-process sys_timer builtin fail cleanly and share validation
sys_timer_set / sys_timer_cancel firing runs in the runner: execute_tool
intercepts both and owns the per-session timer registry. The in-process
builtin, however, still carried a _spawn_timer_workflow stub that raised
NotImplementedError on its success path, plus docstrings claiming timers
were "not yet re-implemented on the runner" — a misleading contract and a
latent crash for any future non-runner dispatch path.
Extract the shared argument validation into validate_timer_set_args so the
runner firing loop and the LLM-facing builtin reject the same inputs with
one delay ceiling, replace the raising stub with a structured "no timer
scheduled" error, and correct the stale docstrings.
* test(tools): remove unused type-ignore in timer validation test
`dict[str, object]` is assignable to validate_timer_set_args's
`dict[str, Any]` parameter, so the `# type: ignore[arg-type]` was an
unused ignore that a strict MyPy run flags. Drop it.
* fix(web): remember the last-picked host in the new-session picker
The landing composer only kept a host selection in an in-memory draft that
is dropped on create and lost on refresh, so every fresh visit re-ran the
auto-select default — the managed sandbox where it's offered, otherwise the
first online host — ignoring the host the user last picked. This is the
"always defaults to the sandbox / first host" complaint.
Persist the explicit choice in localStorage (mirroring the agent
preference) and restore it on mount: the auto-select effect now consults
the stored choice before defaulting, validating a stored host id against
the live list and falling back to the default when it's gone or offline.
The sandbox pick persists as a reserved sentinel.
Co-authored-by: Isaac
* test(web): add managed sandbox-default e2e + clarify seed comment
Address Polly review notes on the last-picked-host change:
- Add tests/e2e_ui managed variant: in a managed deployment whose default
is the "Databricks Sandbox" option, pick a connected host, reload, and
assert the host is restored rather than reverting to the sandbox default
— the original complaint, now covered end to end (the OSS test already
covered the first-online path).
- Note the intentional one-time-seed read of readLastHostChoice() so a
future reader doesn't add it to the effect's dependency array.
Left the pre-existing managed offline-host / info-load-race edge alone:
gating the default auto-select on the /v1/info probe regresses first-paint
host selection (and the flow tests model info as a steady "loading" state),
which isn't worth a rare, pre-existing corner.
Co-authored-by: Isaac
* feat(acp): generic ACP harness + Omnigent-tool MCP bridge for all ACP harnesses
Add a generic `acp` harness that connects Omnigent to ANY agent speaking the Agent Client Protocol (gemini --experimental-acp, @zed-industries/claude-code-acp, goose, qwen, custom in-house agents). Users register named agents in an `acp:` config block via `omnigent setup`; each surfaces as its own harness-picker row (`acp:<slug>`) and drives one well-tested ACP client. Generalized from the existing (duplicated) goose/qwen ACP executors; no new dependency.
Also expose Omnigent's builtin tools (sys_*, load_skill, web_fetch, policy tools) to ALL three ACP harnesses (acp, goose, qwen) via ACP's native session/new.mcpServers, reusing the shared serve-mcp stdio relay the native harnesses use — tool calls route through ctx.dispatch_tool so Omnigent policy is enforced. Shared helper omnigent/inner/_acp_omnigent_mcp.py; global kill switch OMNIGENT_ACP_MCP=0 (generic acp also has a per-agent omnigent_mcp flag).
Routing: the registry stays one `acp` harness; a configured agent is addressed as `acp:<slug>` (canonicalizes to `acp`), command resolved from config at spawn. Improvements over the goose path baked into the generic client: tool-call cards, reasoning (agent_thought_chunk), and a real interrupt via ACP session/cancel.
Tests: unit + a hermetic fake-ACP-agent e2e (handshake -> stream -> tool card -> permission -> completion, no vendor binary) + a real relay start/teardown; goose/qwen/claude_native_bridge/capabilities regressions green.
Co-authored-by: Isaac
* fix(acp): resolve CI failures + address AI-review comments
CI: ruff-format all touched files (pre-commit); move 'Custom ACP agent' to the end of the configure-harnesses list + update the position/priority tests; add 'acp' to the harness-readiness map expectations (config-gated, not CLI-gated); exclude the generic 'acp' harness from the no-agent live-binary matrix (it has no fixed binary).
AI review: comment the two expected-shutdown empty-except blocks in acp_executor; use module _logger instead of a redundant local 'import logging' in harness_plugins.harness_catalog; drop an unused fake_rpc in the acp tests.
Co-authored-by: Isaac
* feat(acp): list each configured ACP agent as its own configure-harnesses row
Previously the setup 'configure harnesses' overview showed a single 'Custom ACP agent' row and the individual agents were buried in the drill-in. Now each configured ACP agent gets its own top-level row (alongside the built-in harnesses), plus an 'Add custom ACP agent' row — matching the web picker, which already lists each acp:<slug>. All rows route to the shared ACP manager (add/edit/remove); a per-agent edit drill-in is a follow-up. No agents configured → unchanged single 'Custom ACP agent' row.
Co-authored-by: Isaac
* fix(acp): per-agent remove + straight-to-add in configure-harnesses
Addresses UX feedback on the ACP rows: (1) the Add row jumps straight into the add flow (prints examples, then prompts) instead of a second add/remove menu; (2) it renders with no ✗ glyph (new 'action' status kind); (3) Remove now lives on each agent's own row via a per-agent drill-in (_manage_acp_agent). Deletes the now-unused combined _manage_acp_harness / _remove_acp_agent.
Co-authored-by: Isaac
Reconnect/relaunch reconciliation looks up a runner's session(s) by
`runner_id` via `list_conversations_by_runner_id`. Four server call
sites drive that query (see omnigent/server/app.py), but `runner_id`
was unindexed, so each lookup was a full table scan of `conversations`.
Add `ix_conversations_runner_id` on `conversations.runner_id`, mirroring
the other single-column lookup indexes on this table, plus migration
z2a2b3c4d5e6 to create it. Extend the migration workspace test to assert
the index is present at head.
Co-authored-by: Isaac
## Related issue
N/A
## Summary
- `_wait_for_claude_prompt_ready` raised its "terminal did not become
ready" error with the tail of a **fresh** capture taken *after* the
30s deadline. That frame is a different moment than any of the ~200
poll decisions the loop actually made — it can show a healthy,
box-present composer while the real failure was 30s of box-absent (or
empty) captures. The mismatch makes the error actively misleading:
triaging one such failure sent us chasing footer-height, prompt-glyph,
and box-rule theories that the attached frame contradicted.
- Attach the **last non-empty capture the loop observed** instead, and
report the poll count and empty-capture count in the message. Those
counts separate the two failure modes that previously looked
identical: mostly-empty captures point at a torn read under a busy
mid-turn repaint (session alive, `capture-pane` came back blank),
while non-empty captures with no box point at Claude never rendering
the prompt (a boot crash whose text the tail then surfaces).
- Poll loop is now do-while so `timeout_s=0` still checks once and always
yields a capture to attach on failure.
- Observability-only: this does not change when the gate passes or fails,
so it does not by itself stop a dropped message — it makes the next
occurrence self-diagnosing instead of requiring reconstruction.
## Test Plan
- `pytest tests/test_claude_native_bridge.py -k wait_for_claude_prompt_ready`
— 3 passed (the pre-existing crash-tail test plus the two added below).
- Full file: 152 passed; the 3 failing tests are pre-existing MCP
channel-server tests unrelated to this change (verified by reproducing
them on the stashed clean tree).
- `pre-commit run --files omnigent/claude_native_bridge.py tests/test_claude_native_bridge.py`
— clean (ruff-format normalized one line).
## Type of change
- [x] Bug fix
- [ ] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Two regression tests added: one asserts the empty-capture count appears
in the error and no bogus "Last terminal output" tail is attached when
every capture was empty; the other proves the tail comes from an in-loop
capture and that a box-present frame arriving only after the deadline
never leaks into the error (i.e. no post-deadline re-capture happens).
Manually verified the live behavior earlier in the investigation by
driving real `claude` 2.1.203 under the production 80x24 tmux geometry
(idle, a 6-subagent fan-out, pane shrunk to 8 rows, all permission
modes) to establish which frames the detector sees.
Reorganize the Appearance page so its two orthogonal choices read
clearly. The single "Theme" block is split into labeled subsections —
"Mode" (System / Light / Dark) and "Color theme" — each with a one-line
helper; "Terminal theme" stays its own section.
- Mode cards now show a mini app-window preview (light / dark, and a
diagonally split tile for System) instead of a bare icon.
- Color theme moves into a dropdown (shadcn Select) with a swatch chip
per option; the trigger mirrors the current selection.
- One selection treatment across the card groups: accent border + a
corner checkmark badge, via a shared keyboard-navigable radiogroup
(roving tabindex + arrow keys). focus-visible stays distinct from
selected, and each group is labeled via aria-labelledby off its heading.
No available options or their names change — only organization, layout,
and interaction consistency. Unit tests + the Appearance e2e are updated.
Signed-off-by: SabhyaC26 <sabhyachhabria@gmail.com>
_RUNNER_ENV_ALLOWLIST forwards DATABRICKS_CONFIG_PROFILE and
DATABRICKS_CONFIG_FILE but not DATABRICKS_AUTH_STORAGE. The host daemon
inherits it (cli.py adds the DATABRICKS_ prefix to the daemon env), so
when the token store is selected via that env var (e.g. the plaintext
JSON cache while ~/.databrickscfg [__settings__] auth_storage=secure) the
host authenticates but every spawned runner falls back to the cfg
default, reads a different/stale token store, and the runner tunnel is
rejected with HTTP 401 even though the host is online.
Add DATABRICKS_AUTH_STORAGE to the allowlist -- a non-secret storage
backend selector, same rationale as the adjacent config selectors -- so
host and runner resolve the same credential store. Deliberately not
switching the runner to the daemon's blanket DATABRICKS_ prefix, which
would leak bearer secrets into (possibly hosted) runners.
Co-authored-by: Isaac
Co-authored-by: jtaylorisbell <jtaylorisbell@users.noreply.github.com>
Adds a color-palette axis to Appearance settings, independent of the
light/dark mode. Ships Omnigent (brand pink, default) plus four popular
palettes — Dracula, GitHub, Catppuccin, and Gruvbox — each with full
light + dark variants.
A palette re-points the existing CSS custom properties under a
`data-theme` attribute on <html>, so it composes with next-themes'
`.dark` class and re-skins the whole app without any component change.
The choice persists in localStorage and is applied before first paint
(no flash). Text selection now tracks the palette accent instead of a
hardcoded pink.
Covered by a themePalette unit suite, SettingsPage picker assertions,
and a Playwright e2e test for the Appearance palette picker.
Signed-off-by: SabhyaC26 <sabhyachhabria@gmail.com>
`list_conversations_by_host_id` had no production callers. Its docstring
claimed reconnect reconciliation used it, but the server mounts the host
tunnel without an `on_host_connect` callback, so that path is never
wired; the real reconnect/relaunch flow keys off `runner_id` via
`list_conversations_by_runner_id`.
Remove the store method (interface + SQLAlchemy impl) and the
`ix_conversations_host_id` index that existed solely to serve it.
`conversations.host_id` carries no FK, so nothing else depends on the
index. Add migration z1a2b3c4d5e6 to drop it.
Drop the two dedicated store unit tests and the
`test_reconnect_with_dead_runner_triggers_relaunch` integration test
(its synthetic callback was the only other caller, exercising the
never-wired host-id reconciliation path). Flip the migration test to
assert the index is absent at head.
Co-authored-by: Isaac
Widen the conversation_items primary key from (workspace_id, id) to
(workspace_id, conversation_id, id) so a conversation's items stay
contiguous under the workspace prefix for the per-conversation prefix
scans that dominate item reads.
Co-authored-by: Isaac
* fix(deps): drop mlflow from dev extras (accidentally added by #526)
mlflow was not in the dev deps on main before #526 merged. It was
inadvertently introduced via a conflict resolution that carried over a
stale comment block from the PR branch. Remove it and clean up the
now-orphaned comment fragment in the hindsight-client entry.
* chore(oss): regenerate public lockfiles against public PyPI/npm
* fix(deps): rename hindsight extra to memory (omnigent[memory])
The design steer on #526 asked for omnigent[memory] (capability-named,
not vendor-named) but the PR landed with omnigent[hindsight]. Rename
the extra key and update all user-facing references: the install hint in
the error message, the remy example, and the module docstring. Internal
names (hindsight.py, HindsightRetainTool, hindsight_retain tool names,
hindsight-client package) are unchanged.
* chore(oss): regenerate public lockfiles against public PyPI/npm
* chore: revert web/package-lock.json to main
The OSS lockfile-regen bot bumped prettier 3.8.4 -> 3.9.4 in
web/package-lock.json on this branch. Prettier 3.9 reformats multi-line
type unions, marking many untouched .ts files dirty and failing the
web-prettier gate. This PR only changes pyproject.toml + Python, so the
web lockfile should match main. Reverting drops the unrelated prettier
bump and its formatting churn.
---------
Co-authored-by: omnigent-ci[bot] <294685417+omnigent-ci[bot]@users.noreply.github.com>
* feat(tools): add Hindsight long-term memory built-in tools
Adds three first-party built-in tools — hindsight_retain / hindsight_recall /
hindsight_reflect — backed by Hindsight (https://github.com/vectorize-io/hindsight),
an open-source agent-memory system. Resolves issue #369.
- omnigent/tools/builtins/hindsight.py: Tool subclasses for retain/recall/reflect.
The memory bank resolves from config.bank_id, else ctx.agent_id, else
ctx.conversation_id, so a single declaration isolates memory per agent.
- Registry: lazy factories in builtins/__init__ that probe for hindsight-client
and fail with an install hint (mirrors the modal sandbox _ensure_sdk pattern).
- Packaging: optional 'hindsight' extra (hindsight-client); kept in the dev set
so the mocked tests can import it (same rationale as mlflow); mypy override.
- Manifests: registry frozenset lock + onboarding list_builtin_tools.
- Docs: tools.builtins example in AGENTSPEC.md.
- Example agent: examples/remy uses all three tools.
- Tests: tests/tools/builtins/test_hindsight.py (mocked client, no network).
hindsight-client is optional and lazily imported, so base installs are unaffected.
Signed-off-by: Ben <ben.bartholomew@vectorize.io>
* fix(tools): dispatch Hindsight memory builtins under wrapped harnesses
The registry entries alone only execute under the native llm executor. Under a
wrapped harness (claude-sdk / codex / cursor / pi) tool calls go through the
runner's local dispatcher, which only runs tools in _ALL_LOCAL_TOOLS — so
hindsight_retain/recall/reflect fell through to the harness and silently no-op'd.
Mirror the web_search wiring in omnigent/runner/tool_dispatch.py:
- add _HINDSIGHT_TOOLS to _ALL_LOCAL_TOOLS (runner dispatches them) and to
_NATIVE_RELAY_BUILTIN_TOOLS (native harnesses have no memory of their own)
- add _execute_hindsight_tool / _hindsight_config_from_spec: read the builtin's
spec config, build the tool, invoke with a ToolContext carrying agent_id so
the bank resolves correctly
- tests/runner/test_hindsight_local_dispatch.py covers dispatch + bank resolution
Full tests/runner suite green (927 passed).
Signed-off-by: Ben <ben.bartholomew@vectorize.io>
* docs(examples): pin a stable bank_id in the remy example
Memory now lands in a human-readable bank ('remy') instead of the opaque agent
id, so it's easy to find in Hindsight. A comment notes that omitting bank_id
falls back to per-agent isolation.
Signed-off-by: Ben <ben.bartholomew@vectorize.io>
* docs(tools): make Hindsight memory tools prompt the model to actually call them
Models tend to acknowledge a fact in chat without persisting it. Two levers:
- Tool descriptions (shown to every agent that enables the tools) now state that
context is lost between sessions and spell out when to call retain/recall.
- examples/remy prompt now mandates calling hindsight_retain and forbids claiming
a save without a successful tool call.
- AGENTSPEC notes that agent authors should prompt their agent to use the tools.
No behavior change to the tools themselves.
Signed-off-by: Ben <ben.bartholomew@vectorize.io>
* docs: drop AGENTSPEC.md edits from this PR
Leave the core spec doc untouched to keep the PR's review surface minimal — the
tools are documented via the examples/remy agent and the tool descriptions
instead.
Signed-off-by: Ben <ben.bartholomew@vectorize.io>
* chore(deps): regen uv.lock with hindsight-client and security fixes
Regenerates the lockfile to include hindsight-client 0.8.3 and its
transitive dependencies. Picks up cryptography 48.0.1 and
pydantic-settings 2.14.2 (fixes OSV advisories GHSA-537c-gmf6-5ccf
and GHSA-4xgf-cpjx-pc3j already present on main).
* test(remy): add structural e2e test for the Remy memory example
Satisfies the test_every_agent_has_a_dedicated_test_file coverage guard.
Checks name, harness, the three Hindsight builtins, and that they all
share bank_id 'remy'. Pure spec-load -- no credentials needed.
---------
Signed-off-by: Ben <ben.bartholomew@vectorize.io>
Co-authored-by: Pat Sukprasert <pattara.sk127@gmail.com>
The policies table enforced name uniqueness on the VARCHAR(256) name
column via a partial unique index (ix_policies_default_name, scope=default)
and a composite unique constraint ((session_id, name)). Both are now keyed
on a new name_cksum column holding sha256(name) — a fixed 32-byte digest —
so the index entries are compact and fixed-width instead of a wide varchar.
Uniqueness semantics are unchanged: two names collide iff their digests do.
The checksum is stamped on INSERT by an ORM column default and recomputed by
the store on rename; it stays store-internal and never appears in the Policy
entity or the HTTP/SDK schema. SQLite has no sha256(), so the migration
back-fills the digest in Python.
Co-authored-by: Isaac
* feat(benchmarks): add HTTP user-journey performance harness
Add a runnable benchmark under dev/benchmarks/omnigent/ that boots a real
omnigent server against a throwaway SQLite DB (no runner, no LLM), drives key
HTTP journeys under load, and emits a versioned JSON report of latency
percentiles + throughput. Modeled on MLflow's dev/benchmarks/gateway workflow.
v1 covers the server + DB request path: list_sessions, create_session,
get_session, and load_conversation_history (history seeded runner-free via the
external_conversation_item event). The report JSON is the contract a workspace
Databricks notebook consumes (artifact -> Delta -> AI/BI dashboard).
The environment is written as a superset: a with_runner flag (default off)
gates a mock-LLM + runner path so phase-2 full-turn journeys are additive, not
a rewrite.
Co-authored-by: Isaac
* feat(benchmarks): seeded corpus, backend matrix, nightly workflow
Make the benchmark meaningful and automated:
- seed.py: deterministic corpus seeder via the store API (no HTTP/runner) —
create_session_with_agent + "local" permission grant + batched append.
Idempotent (reuse marker), --reseed to force, SEED_SCHEMA_REVISION pinned
to the Alembic head.
- environment.py / run.py: accept --database-uri and stamp a `backend`
(sqlite/postgres) field into the report. None keeps the throwaway-SQLite
path; a seeded URI (SQLite file or postgresql+psycopg://) benchmarks a
realistic corpus.
- journeys.py: read journeys target an existing corpus session (self-seed
fallback when empty); add search_sessions (the unindexed LIKE path where
SQLite and Postgres diverge most).
- Schema-drift guard: scripts/check_benchmark_seed_schema.py + a pre-commit
hook fail when the DB schema head moves without the seed being refreshed.
- benchmark.yml: nightly + dispatch, backend matrix (sqlite + a postgres:16
service container), per-backend seed with an schema-keyed SQLite seed cache,
one artifact per backend.
Verified: seeded SQLite e2e shows list_sessions ~1.3ms -> ~6ms p50 and
search_sessions ~79ms p50 vs the empty-DB baseline. 9 smoke tests pass; ruff,
mypy, and pre-commit (incl. the new guard) clean. The Postgres leg's live run
is first exercised by CI (Docker is org-locked locally); the psycopg dialect
resolves and the URI passthrough is covered by the SQLite --database-uri path.
Co-authored-by: Isaac
* feat(benchmarks): full-turn (runner) journeys
Add four full-turn journeys that drive a real agent turn end-to-end through the
runner + a zero-latency mock LLM (with_runner=True), all using the openai-agents
SDK harness:
- session_cold_start: fresh session provisioning + first turn (runner spawn +
executor construction).
- warm_turn: steady-state per-turn dispatch overhead.
- time_to_first_token: post → first streamed output_text delta (subscribes the
session SSE stream; waits for connect rather than a fixed sleep so the delay
isn't in the measured window).
- interrupt: cancel a running (gated) turn; time to the cancellation marker.
Only measure what we control: full-turn journeys always use openai-agents, which
runs in-process (no vendor binary) — native harnesses launch the real CLI and
are excluded. The mock is zero-latency, so numbers are omnigent
dispatch/streaming/cancel overhead, not model latency. No delay knob added.
Excluded as agent-dependent: multi-turn, tool-calling, large-history turns.
run.py auto-boots with_runner=True when any selected journey needs it and stamps
harness=openai-agents. Adds a needs_runner flag on Journey; adds async
time_to_first_delta / drive_and_interrupt / _wait_idle to BenchEnvironment.
Extends the mock's /mock/set_fallback with an optional stream flag so a
reset-surviving fallback can emit deltas (needed for TTFT).
Verified: a with_runner smoke runs all four journeys once (first end-to-end
exercise of the runner path); manual e2e shows warm_turn ~235ms vs
session_cold_start ~1.6s. 10 smoke tests pass; ruff, mypy, pre-commit clean.
Co-authored-by: Isaac
Pre-fill the name field with the auto-derived slug and let users
override it. Also fix parameter description overflow in the dialog
with min-w-0 on the content container and break-all on long text.
The doc-drafter prompt was framed purely additively (extend a page, create
a page, document what the PR "introduced"), so a PR that removes or
deprecates a user-facing feature would nudge the drafter toward writing
prose rather than pruning the now-untrue docs. The classifier already
routes removals correctly, so the gap was only in the drafter.
Add a removal/deprecation path: classify the diff intent in Step 1, and in
Step 3 delete whole pages (git rm + drop the SECTIONS sidebar entry) or cut
sections/references for a removed feature, or mark deprecated-but-present
features in the site's usual style. Report deletions in the output summary.
The workflow already stages and detects deletions (git add -A /
git status --porcelain), so no workflow change is needed.
Co-authored-by: Isaac
Queued messages could reach the runner out of FIFO order when the user
navigated away mid-queue. The foreground flush (maybeFlushQueuedHead →
send()) serializes its POSTs on the module-level sendChain, but the
background flush (flushBackgroundQueues → postEvent) bypassed it. At the
navigate-away handoff, an in-flight foreground send() still awaiting its
chain slot could be overtaken by a background postEvent that fired
immediately — delivering messages out of submission order (observed on
cursor-native, whose instant turns make the window easy to hit; the runner
appends FIFO as received, so the scramble is entirely client-side).
Have flushBackgroundQueues join the same sendChain: take a slot (await
priorSend before the upload/post, release in finally), so every POST across
both paths is ordered through one primitive.
Also reset sendChain in initChatStore so a prior run's unresolved send
can't block the next (production calls it once at boot; tests per case),
and restore the real send action in the test beforeEach (a prior test's
setState({ send: spy }) otherwise leaks into later cases).
Test: a background flush fired while a foreground send()'s POST is held
open does not deliver until the foreground POST resolves. Verified it fails
without the fix (background overtakes) and passes with it.
Co-authored-by: Isaac
Replace a timing-based 0.5 s wait_for/shield assertion with a
fwd_seen Event set by _ForwardBlockingHarnessClient.post() the
moment the interrupt forward blocks on fwd_gate. The test now
waits for provable in-flight status instead of hoping 0.5 s is
long enough on a loaded CI machine.
* feat(web): split sidebar sessions into My sessions / Shared with me tabs
Sessions shared with the viewer previously sat in an inline collapsible
"Shared with me" section below the owned-session list. Move them to a
dedicated tab so the two scopes are visually distinct and the shared list
gets its own space (flat, headerless, with its own infinite scroll).
The "My sessions" tab keeps the full Pinned / Projects / Sessions
structure; "Shared with me" is a flat list of every non-archived session
the viewer doesn't own (computed from notArchived, so a pinned/filed
shared session never drops off it). New session snaps back to My sessions.
The tab strip only renders on a multi-user server — gated on
!isCurrentServerLocal(), the same predicate AppShell uses to disable the
Share affordance. A loopback-only local server has a single user and
can't share sessions, so the split is meaningless there; the list falls
back to the owned sessions. Keyboard nav and shift-select are tab-aware
and, on the shared tab, ignore the collapsed set (the list always renders
expanded), so a stale persisted "Shared with me" collapse can't empty them.
Co-authored-by: Isaac
* fix(web): keep pinned/filed shared sessions off My sessions; paginate empty tabs
Address two issues in the sidebar tab split:
- Pinned and project folders drew from all non-archived sessions, so a
shared session the viewer pinned (localStorage is ownership-agnostic) or
filed into a project (editable share) rendered under Pinned / a project
folder on My sessions AND on the Shared tab. Build both from owned-only
sessions so non-owned sessions stay on the Shared tab exclusively.
- The list is one paginated stream (owned + shared mixed, updated_at desc),
so a tab can be empty on the loaded window while its sessions live on a
later page. The pagination sentinel lived inside the non-empty render
branch, so an empty tab stopped fetching and stranded the user on a false
"empty" state (e.g. Shared tab when page 1 is all owned). Keep the
sentinel mounted in the empty branch when more pages exist.
Co-authored-by: Isaac
* refactor(web): reuse Pinned / Projects / Sessions layout for both sidebar tabs
Rather than rendering the Shared tab as a bespoke flat list, scope the
section-building to the active tab's conversations and render the same
Pinned / Projects / Sessions tree for both tabs. "mine" is the sessions
the viewer owns; "shared" is the ones others shared with them.
- Pins are localStorage and ownership-agnostic, so a pinned shared session
now floats to a Pinned section on the Shared tab, matching My sessions.
- Projects stay a My-sessions-only tool: filing into a project is now
gated on ownership (the row's "Add to project" / "Move session" menu
item is hidden for non-owned sessions), and the Shared tab renders no
Projects group. A shared session that already carries a project label
just lands in the flat Sessions list there.
- Collapses the special-case `showShared` render branch and the shared
special cases in keyboard-nav / shift-select ordering, since `sections`
is now tab-scoped.
Co-authored-by: Isaac
Fixes two MySQL incompatibilities: TEXT columns cannot have DEFAULT values,
and TEXT columns cannot be indexed without a key-prefix length.
- db_models.py: title → String(768); ix_conversations_parent_title_unique
gains mysql_length={"title": 512} so the index works on MySQL
- Migration w1a2b3c4d5e6: alters the column and drop/recreates the unique
index with the MySQL prefix hint; handles the case where the index is
absent on MySQL (TEXT was never indexable there)
- Tests: 4 new tests covering VARCHAR(768) column type, server_default,
data survival, and downgrade round-trip on SQLite; manually verified
upgrade+downgrade on PostgreSQL and MySQL
Reverts PR #1279. Model selection can now be done right after fork as a
first action for codex, so the dedicated codex-native --model launch flag
and the "Restart with model…" fork dialog are no longer needed.
Backs out:
- Backend: the OMNIGENT_CODEX_NATIVE_MODEL_FLAG opt-in flag, the
codex --help --model capability probe, and the explicit --model launch
plumbing in codex_native_app_server.py; the fork route's model_override
parameter, validation, and family-check (_agent_harness_id); the
SessionForkRequest.model_override schema field and its store plumbing.
- Frontend: the codex-only RestartWithModelDialog and the AgentInfo
"Restart with model…" trigger; forkSession's modelOverride param.
- The associated backend, store, vitest, and e2e-ui tests.
The always-on per-session config.toml `model =` pin and the pre-existing
session-level model_override field are untouched.
Resolved conflicts from the ap-web -> web frontend rename and later
main-branch changes to AgentInfo by re-applying the removal surgically on
top of current main rather than adopting the stale pre-PR text.
Verified: 202 backend tests (fork route, conversation store,
codex_native_app_server), 34 AgentInfo vitest, web tsc, and prettier all pass.
Co-authored-by: Isaac
Landing on a policy's config view in the add-policy dialog (the "+" in the
agent info popover, and the admin global-policies page) left no way back to
the policy list: both Cancel and the X closed the whole modal. Selecting the
wrong policy meant reopening the dialog from scratch.
Cancel now deselects back to the list when a policy is selected, and only
closes the dialog from the list itself. Closing via X/Escape resets the
selection so reopening always starts at the list instead of a stale config
view.
Co-authored-by: Isaac
* feat(android): add ktlint formatter to CI and pre-commit
Kotlin files had no enforced style — add ktlint 1.8.0 to close that gap,
mirroring the pattern already used for Swift (local wrapper that no-ops
when the tool is absent) but with full CI enforcement since Java is
available on ubuntu-latest.
Changes:
- web/android/.editorconfig: ktlint style config (4-space indent,
100-char line length, standard rule set)
- web/android/bin/ktlint.sh: wrapper script; exits 0 if ktlint is not
installed so developers without it don't get blocked at commit time
- .pre-commit-config.yaml: android-ktlint-format (auto-fix) and
android-ktlint-check (lint gate) hooks for *.kt / *.kts files
- .github/workflows/lint.yml: installs ktlint before pre-commit runs so
the check is enforced in CI
- web/android/**/*.kt: apply initial ktlint --format pass to existing
sources so the hook is green from the first run
* fix(android/ci): harden ktlint install step and scope editorconfig
Address review feedback on #2179:
- Add `curl --fail` so a 4xx/5xx response (e.g. wrong version tag) fails
loudly at the download step rather than silently installing an HTML body
- Verify the ktlint binary against the SHA-256 checksum published alongside
each release before marking it executable
- Add `root = true` to web/android/.editorconfig so a future repo-root
.editorconfig can't bleed Kotlin-unintended settings through EditorConfig
inheritance
Promotes host_id into the PK alongside workspace_id, demoting owner and
name to regular NOT NULL columns backed by a uq_hosts_workspace_owner_name
unique constraint. The old uq_hosts_host_id unique constraint is dropped
since uniqueness is now enforced by the PK.
- Migration u1a2b3c4d5e6: uses batch_alter_table with copy_from to
correctly rebuild the SQLite table from scratch with the new PK.
- HostStore.upsert_on_connect: primary lookup now keys on (workspace_id,
host_id). The W2-class boundary (reject foreign-owner host_id claim)
is enforced explicitly via IntegrityError when allow_host_id_reown=False
and the existing row's owner doesn't match the connecting owner.
- _rotate_host_id: already correct; kept as-is.
- Tests: update session.get() PK tuple in test_db_models; fix
test_unique_host_id to commit h1 before adding h2 so the PK violation
fires at the DB; update test_migration_workspace_id to handle the later
PK override for hosts; add test_migration_host_pk_workspace_host_id.
* fix(web): keep queued messages FIFO when status flickers idle
A follow-up sent while an earlier one waits in the client-side queue could
jump ahead of it: handleSend takes the direct send() path whenever the
session reads idle, and that path isn't ordered against the queue drain.
On harnesses whose sessionStatus flickers idle between quick turns
(cursor-native), a later message slipped onto the direct path mid-queue
and was delivered before the still-queued earlier one — scrambling the
order the agent received (verified in a runner log: the runner appended
messages FIFO as they arrived; the reorder happened client-side).
Funnel every send through the single FIFO queue once the conversation has
anything queued, even if it momentarily reads idle. enqueueMessage already
flushes immediately when genuinely idle, so this never stalls a message —
it only prevents the direct path from overtaking the queue.
Co-authored-by: Isaac
* test(web): unit-test the queue-vs-send decision
Extract handleSend's enqueue-vs-direct-send predicate into an exported
pure helper, shouldQueueSend, and unit-test it. The decision was inline in
handleSend (which reads the store) and had no coverage; the ordering fix
lives entirely in this predicate.
Tests: new chat sends directly; busy (streaming/running/waiting) queues;
idle with an empty queue sends directly; idle but with this conversation
already queued still queues (the ordering-race fix); a different
conversation's queue doesn't force this one onto the queue.
Co-authored-by: Isaac
* docs(web): trim shouldQueueSend comments
Co-authored-by: Isaac
The low-cardinality closed-set columns (conversations.kind,
conversation_items.type/status, comments.status, account_tokens.kind,
policies.type, policies.scope, hosts.status, agents.kind) were stored as
VARCHAR guarded by string CHECK constraints. Store them as compact
SMALLINT integer codes instead, matching the existing int-coded
session_permissions.level.
A new omnigent/db/enum_codecs.py owns the stable name<->int tables and is
the single translation point: conversion happens only at the store
row<->entity boundary, so entities, the HTTP API, the web client, and the
SDKs keep seeing the string names unchanged. A backfill migration
(u1a2b3c4d5e6) converts existing rows in place and is reversible, portable
across SQLite and PostgreSQL. The agents.kind and policies.scope partial
indexes are dropped and recreated around the column swap since SQLite
batch mode can't copy a partial-index predicate across a rename.
The comment-update route now rejects an unknown status with a 400 instead
of letting the enum codec raise into an opaque 500 — the column is now a
closed enum (draft/addressed), matching the validation the update_comment
tool already enforced.
Co-authored-by: Isaac
## Related issue
N/A
## Summary
- Native-harness sessions (`omnigent claude`/`codex`/`pi`/etc.) previously
always opened bash for "+ New shell"; they now open the user's login shell.
- `omnigent/_platform.py`: add `default_interactive_shell()` (basename of
`$SHELL` when it names a known shell on PATH, else bash) and
`installed_interactive_shells()` (that default first, then any of
bash/zsh/fish on PATH; always non-empty).
- `omnigent/native_coding_agents.py`: `native_shell_terminal_spec()` now
declares one unsandboxed caller-process terminal per installed shell, keyed
and commanded by the shell basename, `$SHELL` first. The 11 native wrappers
call this shared helper instead of a hardcoded `{"shell": {"command": "bash"}}`
block.
- `web/src/shell/NewTerminalButton.tsx`: branch on
`useTerminalFirst().isNativeWrapper` — native sessions with multiple shells
get a split button (primary click launches the `$SHELL` default; a caret opens
a picker of installed shells, default labeled). SDK agents with multiple
distinct-purpose terminals keep the existing plain dropdown unchanged.
- `examples/polly/config.yaml`: add a `zsh` terminal alongside the existing
bash `shell` for the builtin polly agent.
## Test Plan
- `uv run pytest tests/inner/test_proc_and_platform.py tests/test_native_coding_agents.py`
— new unit tests for shell detection and the multi-shell spec.
- `uv run pytest -k "native and (materialize or terminal or agent_spec)"` — 296
passed, including the runner create-session-terminal flow; updated 4 native
wrapper tests that asserted the old single-`shell` shape.
- `npx vitest run src/shell/NewTerminalButton.test.tsx` (+ related shell suites)
— split-button default launch, caret pick of a non-default shell, and SDK
dropdown-unchanged cases.
- ruff check/format, prettier, oxlint, and tsc clean on all touched files.
- Verified polly's YAML parses through `_parse_terminals` with both `shell`
(bash) and `zsh` terminals.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Shell detection, the native multi-shell spec, and the frontend split-button
behavior are covered by new/updated unit tests (pytest + vitest). Manually
verified that `default_interactive_shell()`/`installed_interactive_shells()`
resolve the host's shells, all 11 native wrappers import cycle-free, and
polly's edited YAML parses through Omnigent's real terminal parser. The live
end-to-end (clicking "+ New shell" in a running native session and confirming
the shell that opens) was not exercised here as it needs an interactive session.
* feat(web): show and manage the branch when starting in an existing worktree
Starting a session directly in a pre-existing git worktree previously
bound the workspace with no branch recorded, so the sidebar showed no
branch subtitle and the opt-in "Delete local branch" flow was
unavailable — the same worktree Omnigent would offer to clean up if it
had created it.
Thread the existing worktree's branch through as a new `workspace_branch`
field on both create paths (`POST /v1/sessions` and
`POST /v1/hosts/{id}/runners`). It persists as the session's `git_branch`
without creating a worktree, so the sidebar shows the branch and the
existing delete dialog (gated on `git_branch != null`) can remove the
worktree + branch. `workspace_branch` is mutually exclusive with `git`
(which creates a worktree) and requires a host; the server validates the
branch name since the host runs no git for this path.
Co-authored-by: Isaac
* test(e2e-ui): assert workspace_branch is sent for existing worktrees
The E2E UI Required judge flagged the existing-worktree start-session
change as needing Playwright coverage. Extend the existing
select-existing-worktree e2e_ui test to assert the create body now
carries workspace_branch (the picked worktree's branch), alongside the
existing no-git-spec / worktree-dir-workspace assertions.
Co-authored-by: Isaac
* fix(server): don't force-remove an existing worktree on create-rollback
The create-rollback in `_create_session_from_existing_agent` runs
`git worktree remove --force` + `git branch -D` when
`create_conversation` fails, to clean up an orphan worktree Omnigent
just created. It was gated on `git_branch is not None`.
The existing-worktree path (`workspace_branch`) also sets `git_branch`
but creates no worktree — the workspace IS the user's pre-existing
worktree. So a persistence failure on that path would force-remove the
user's worktree and delete their branch: data loss.
Gate the rollback on whether Omnigent actually created a worktree here
(new `created_worktree_path`), mirroring the `worktree is not None`
guard already used on the launch-runner path in hosts.py. Add two
integration tests: a failure on the workspace_branch path sends no
remove frame, and a failure on the git path still rolls back the
worktree Omnigent created.
Co-authored-by: Isaac
* refactor(server): fold existing-worktree bind into SessionGitOptions
Replace the separate top-level workspace_branch field with an
existing_worktree flag on SessionGitOptions, so the git block carries
both modes: create (default) makes a worktree, bind
(existing_worktree=true) records a pre-existing worktree's branch as
git_branch without creating one. base_branch is rejected in bind mode.
This keeps a single branch-name concept and puts the create/bind intent
on the git object itself. The create-rollback stays gated on whether
Omnigent actually created a worktree (created_worktree_path in
sessions.py, the worktree object in hosts.py), so a bind-mode
persistence failure still never force-removes the user's worktree.
Behaviour is unchanged; only the wire shape moves from
{workspace_branch: "x"} to {git: {branch_name: "x", existing_worktree: true}}.
Co-authored-by: Isaac
* refactor(server): dedupe branch validation across worktree modes
Fold the create/bind split into a single `if body.git is not None`
block on both worktree paths and hoist the shared
`validate_branch_name` call above the mode branch, so the name is
validated once instead of in each arm. Behaviour is unchanged; create
mode still creates a worktree and bind mode still records the branch
without creating one.
Co-authored-by: Isaac
Host names are short identifiers from config.yaml; 64 chars matches every
other short-identifier column in the schema. Adds migration t1a2b3c4d5e6
with upgrade/downgrade and a test verifying the column width after both.
Back-fills NULL titles to '' via migration s1a2b3c4d5e6 and alters the
column to NOT NULL with a server_default of ''. The store layer converts
'' ↔ None at the entity boundary so the Conversation.title field stays
str | None throughout the application layer.
* feat(server): publish response.policy_denied on a native tool-call DENY
A native harness (Claude Code, Codex, ...) routes each tool call through
Omnigent's policy engine via the vendor PreToolUse hook
(POST /v1/sessions/{id}/policies/evaluate). The DENY verdict is returned
synchronously to that hook, so unlike the SDK/wrap path nothing on the session
stream reflects that a native action was blocked -- observers could only infer
it from the blocked tool's absence.
Publish a positive signal instead:
- New PolicyDeniedEvent (type "response.policy_denied", fields conversation_id/
reason/phase) added to the ServerStreamEvent union. The wire name is
response-prefixed to match the web-UI wire decoder, which matches the raw
event: name literally (a bare "policy_denied" would be dropped).
- _publish_policy_denied helper mirrors _publish_collaboration_mode.
- Emitted from evaluate_policy on a tool_call-phase DENY, a sibling to the
existing request-phase blocked-notice forward. Observational (not gated on
write access); purely additive -- the synchronous hook response is untouched.
The web UI already handles this event type; the harness capability bench will
consume it to give native harnesses a real Policy DENY verdict.
Tests: PolicyDeniedEvent round-trips the union; the helper emits a typed,
union-valid event; _format_sse emits the response.policy_denied wire name.
* feat(harness-bench): observe native Tool calling + Policy DENY
The native-tui driver stubbed run_tool_turn, so every native harness row showed
`·` for Tool calling and Policy DENY -- a bench observation gap, not a native
limitation. Implement real observation:
- Tool calling (deny=False): post a per-vendor tool-provoking prompt (echo via
the vendor's own shell tool), then scan session items for the new
function_call the vendor bridge mirrors -> result.tool_calls.
- Policy DENY (deny=True): attach a tool_call-phase deny to the session via
POST /v1/sessions/{id}/policies using the registered cel_policy handler
(ternary expression targeting the provoked tool), then watch the stream for
the response.policy_denied signal -> result.tool_call_denied. Does not rely on
a blocked function_call_output (a native deny short-circuits at the hook and
may persist no output), which is why the server-side positive signal exists.
Per-vendor tool name + prompt live on NativeVendor (Bash for claude/pi, shell
for codex); a native with no mapping SKIPs. SKIP (never a false UNSUPPORTED) on:
no tool mapping, fail-open policy (policy_hook_disabled_reason captured at
terminal-ensure), or the CEL handler being unregistered (cel_expr_python absent).
The transport-agnostic probes are unchanged -- they read result.tool_calls /
tool_call_denied. Manifest keeps tool_calling/policy_deny SUPPORTED (now
live-probed on both transports; env gaps reconcile as SKIPPED).
Tests: offline driver tests with a fake client/stream cover tool-call
observation, the deny attach + denied-event, and every SKIP path; the probes
turn the native results into SUPPORTED verdicts.
* fix(harness-bench): check tool_call_denied before the no-tool-call guard
The policy_deny probe was written for full-server, where a denied tool still
surfaces a function_call item. On native-tui a tool_call-phase DENY short-
circuits at the vendor PreToolUse hook *before* the tool runs, so no
function_call item persists and result.tool_calls is legitimately empty. The
probe's first guard (`if not tool_calls: SKIPPED`) therefore swallowed a real
native deny before ever checking tool_call_denied.
Hoist the tool_call_denied check to the top: a confirmed DENY (from the
response.policy_denied stream signal on native, or the blocked function_call_
output on full-server) is enforcement whether or not an item persisted. The
"model never attempted the tool" and "wrap-direct, no evaluation" SKIP branches
now only apply when no deny was observed. No full-server regression: a denied
full-server call still sets tool_call_denied and completes -> SUPPORTED.
* fix(harness-bench): deny any tool call by phase; vary deny-turn command
Two refinements from the first live run, where both natives skipped Policy DENY:
- codex ran the tool but the deny didn't fire: the CEL targeted
event.data.name == "shell", but the wire tool_name in the policy-hook payload
is the vendor's raw name, which need not equal the forwarder's item name.
Deny on the phase alone (event.type == "tool_call") instead, so the block
lands whatever the vendor calls the tool. That is exactly what "is a
tool-call DENY enforced?" asks, and the bench-owned session makes a
blanket tool-call deny harmless.
- claude called no tool on the deny turn: the deny turn reused the allow turn's
session with an identical echo request, so the model saw it already done.
Vary the echo token per turn (omnigent-bench-allow vs -deny) so the deny
turn is a fresh request the model must actually call the tool to satisfy.
* docs(harness-bench): scope the manifest note to what is live vs wired
tool_calling is live-probed on both transports; policy_deny is live on
full-server and wired (but native enforcement is a follow-up) on native-tui.
Keep the note honest so a reader doesn't assume native DENY is confirmed.
* docs(harness-bench): record the root cause of unenforced native deny
Live diagnosis (temporary instrumentation, now removed) confirmed the native
Policy DENY gap: the deny policy IS attached to the correct session and the CEL
DENYs a tool_call event, but the tool runs anyway with NO policy evaluation on
the stream. Root cause: the bench's native terminal-ensure launch does not
thread ap_server_url into claude_native_bridge.build_hook_settings, so the
evaluate-policy PreToolUse hook (gated on `if ap_server_url:`) is silently
omitted -- no permission_hook.json is written and native tool calls are never
gated. Not a session-scoping issue (ruled out: policies=['bench_tool_deny'] on
the right session) and not a harness that ignores policy. Wiring the hook on the
bench launch path is the follow-up; the probe SKIPs cleanly meanwhile.
* feat(harness-bench): map tool provocation for every in-repo native
Extend _NATIVE_TOOL_PROVOCATION from 3 natives (claude/codex/pi) to all
in-repo ones: adds kiro (shell), qwen (run_shell_command), goose
(developer__shell), hermes (terminal), antigravity (run_command), kimi (Bash).
Tool names sourced from omnigent/policies/builtins/safety.py::ask_on_os_tools
and each vendor's native module, so each entry is a grounded claim, not a guess.
Now that the deny gates on the tool_call phase alone (name-agnostic),
``tool_name`` is only a descriptive non-empty gate, so a shared shell-tool
prompt covers the vendors uniformly. Comments/docstring updated to match (the
old "must equal the raw PreToolUse tool_name" note was stale). cursor-native is
deliberately left unmapped (lazy-chat; add once it provisions reliably), which
the skip test still relies on. SKIP-safety unchanged: a wrong prompt skips,
never a false verdict. Verification of the new entries is a live follow-up.
* feat(web): add terminal theme preference module
A persisted light/dark palette choice for the terminal, independent of the app
chrome theme. Mirrors codeFontPreferences, localStorage-backed with an in-module
pub/sub so a Settings change re-themes mounted terminals live. "auto" follows the
app's resolved theme, while "light"/"dark" pin it.
* feat(web): choose a terminal theme in Appearance settings
Adds a Terminal theme radiogroup (Match app / Light / Dark) under Settings ->
Appearance. TerminalView resolves the chosen mode against the app theme and
pushes the result to the live xterm through the existing setTheme path, so a
light terminal can sit under a dark app and vice versa. The resolved palette is
exposed as data-terminal-theme on the terminal view for observability.
* test(e2e_ui): terminal theme is independent of the app theme
Drives the Appearance control and a live shell to assert a light terminal under
a dark app and a dark terminal under a light app, plus the match-app default and
persistence across reload.
* fix(e2e_ui): scope theme-toggle locators to the app Theme radiogroup
The new "Terminal theme" radiogroup shares the "Theme" substring and reuses the
Light/Dark radio labels, so test_theme_toggle's unscoped get_by_role locators
matched two elements under Playwright strict mode. Scope every lookup to the
exact app Theme radiogroup so the app-theme test stays unambiguous.
* feat(db): add workspace_id to all tables as leading primary-key column
Add a NOT NULL workspace_id column (BigInteger, server_default 0) to all
twelve tables and fold it into each primary key as the leading column,
laying the groundwork for per-workspace tenancy. Behaviour is unchanged:
every row lives in workspace 0 (DEFAULT_WORKSPACE_ID).
Migration r1a2b3c4d5e6 backfills existing rows to 0 and rebuilds each PK
to (workspace_id, <existing pk cols>) via SQLite-safe batch recreate /
explicit PK drop on PostgreSQL. Store and server primary-key lookups
(session.get) and dialect upserts (on_conflict index_elements) are
updated for the composite key.
Co-authored-by: Isaac
Signed-off-by: aravind-segu <aravind.segu@databricks.com>
* feat(db): scope all store queries to the default workspace_id
With workspace_id now the leading primary-key column, queries that
filtered only on the old key columns (e.g. WHERE id = ?, WHERE user_id
= ?, WHERE owner = ?) could no longer seek the primary-key index — the
unconstrained leading workspace_id degraded them to scans.
Add workspace_id == DEFAULT_WORKSPACE_ID to every store/server query on
these tables — selects, updates, deletes, subqueries, joins, the legacy
Query.filter paths, and the raw-SQL ILIKE search fallback — so
primary-key lookups seek the composite PK again and every access path is
workspace-scoped (forward-correct for multi-tenancy). Behaviour is
unchanged: all rows live in workspace 0.
Co-authored-by: Isaac
Signed-off-by: aravind-segu <aravind.segu@databricks.com>
* feat(db): resolve workspace_id through a context seam, not a constant
Introduce ``current_workspace_id()`` (a ContextVar defaulting to
DEFAULT_WORKSPACE_ID) plus a ``workspace_scope`` context manager, and
route every store/server access through it: reads and filters call
``current_workspace_id()`` instead of the hardcoded constant, and the
workspace_id column's insert default is now that callable (so ORM
inserts stamp the active workspace).
This is the single injection point a multi-tenant deployment needs.
OSS leaves the ContextVar at 0, so behaviour is unchanged; a deployment
like universe binds a real workspace id per request via ``workspace_scope``
in middleware — an additive change that touches none of these files, so
the code stays byte-identical across deployments and syncs cleanly.
Adds tests covering the default, scope set/reset, insert stamping, and
cross-workspace read isolation.
Co-authored-by: Isaac
Signed-off-by: aravind-segu <aravind.segu@databricks.com>
---------
Signed-off-by: aravind-segu <aravind.segu@databricks.com>
Standard Okta tiers (without custom API Access Management) omit the
email_verified claim from id_tokens for directory-provisioned users,
so the OIDC callback's hard reject breaks SSO for those deployments.
Add OMNIGENT_OIDC_SKIP_EMAIL_VERIFICATION (default off): when set,
accept the signed id_token email claim without requiring
email_verified. Default path unchanged — absent/false claims still
hard-reject. Enabling logs a startup warning plus an info line per
bypassed login. GitHub OAuth unaffected.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Related issue
N/A
## Summary
- Add `dev/omnidev/`, a standalone Rust TUI that replaces the
three-terminal local dev flow (`omnigent server`, `omnigent host`,
`npm run dev`) with one long-running supervisor.
- Each checkout runs as an isolated "pod": its own state dir under
`~/.cache/omnidev/<repo>-<hash>/`, its own SQLite DB / artifacts /
logs, and auto-allocated server + vite ports (probed from 6767/5173,
persisted in `pod.toml`). Isolation reuses the env-var contract proven
by `scripts/backend-smoke.sh` (`OMNIGENT_DATA_DIR`,
`OMNIGENT_CONFIG_HOME`, `OMNIGENT_DATABASE_URI`, `HOME`, `XDG_*`,
`OMNIGENT_URL`).
- Supervises the three processes in their own process groups with
health-gated startup ordering (server `/health` then host) and crash
auto-restart with backoff; tears the whole tree down cleanly on quit.
- Restarts the backend (server then host) on debounced `omnigent/**/*.py`
changes; the frontend is left to Vite HMR and is not watched.
- Log inspection: per-process ring buffers with scrollable panes
(`server | host | vite | all`), follow-tail, and write-through to
`<pod>/logs/*.log`.
- TUI styling reads on both light and dark terminals: a light neutral
chrome bar with dark text, mid-tone per-service accent colors, and the
log body left on the terminal's default background so ANSI colors
render naturally. Header shows clickable `localhost:<port>` URLs while
functional connections stay on `127.0.0.1`.
- Ignore `dev/omnidev/target/` in `.gitignore`.
## Test Plan
- `cargo build`, `cargo clippy --all-targets`, and `cargo fmt` all clean.
- `cargo test` passes 4 integration tests covering repo-root discovery,
per-repo pod-dir stability, and port probe/persist/override.
- Verified `--help` and the out-of-repo error path, and confirmed
`uv run omnigent --version` (the exact spawn path) resolves from the
repo root.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
The TUI process-supervision loop needs a live terminal and real
child processes, so it isn't unit-tested. Pure logic (paths, ports,
pod-dir keying) is covered by `tests/pod_setup.rs`; the interactive
behavior (backend reload on a `.py` edit, Vite HMR without restart,
crash recovery, clean teardown) was verified manually per the README's
verification steps.
* feat(web): add code font size + family setting for editor and terminal
Settings → Appearance gains a "Code font" size stepper and family input
that drive the Monaco code editor and the xterm terminal, separate from
the chrome/UI font (which #2040/#2047 already handled and deferred code
widgets on).
Unlike the rem-based chrome — which scales off the --ui-font-scale /
--ui-font-family CSS variables — Monaco and xterm are fixed-pixel
widgets: they read an absolute size + family once at construction and
only re-measure when told to. So codeFontPreferences.ts exposes an
in-module pub/sub (subscribeCodeFont) that the write helpers fire after
persisting; mounted editors/terminals re-apply the change imperatively
(editor.updateOptions / term.options + refit) with no reload or
reconnect.
Size defaults to 13 (range 10-24); an empty family falls back to the
shared mono stack. Persisted under omnigent:code-font-{size,family}.
* feat(web): label code-font controls in full instead of a shared heading
Drop the "Code font" subheading and rename the two rows to "Code font
size" and "Code font family" so each reads unambiguously next to the
UI-font rows above. Labels only — the test-ids and the role="group"
aria-label ("Code font size") are unchanged.
* fix(web): code-font — emit intended value on write; unify empty-family default
Addresses review feedback:
- writeCodeFontSizePx / writeCodeFontFamily now broadcast the intended value
instead of having emit() re-read storage. A failed persist (quota/denied)
still live-applies to mounted editors/terminals rather than snapping them
back to the stale/default stored value.
- codeFontFamilyForEditor resolves an empty family to the shared mono stack for
Monaco too (not just the terminal), so the editor and terminal share one
default look instead of Monaco falling back to its own built-in mono.
- Tests: a MonacoDiffViewer case asserts a mounted editor live-re-fonts via
updateOptions; the TerminalSession setFont test asserts the refit
(sendResize) and tolerates a down socket; module tests cover emit-on-write
failure.
* test(e2e_ui): disambiguate font-group locators; keep comment anchor visible at 13px
The new code-font controls' aria-labels ("Code font size" / "Code font
family") contain the chrome-font labels as substrings, so the existing UI-font
e2e locators — get_by_role("group", name="Font size"/"Font family"), which match
by substring — resolved to two elements. Add exact=True to those (and the
code-font locator, defensively).
The non-markdown comment test seeded its anchor word in a trailing comment on
the longest line; at the code editor's new 13px default that line scrolls
off-screen, so the double-click word-select couldn't reach it. Move the anchor
to a short leading comment line so it stays visible at any code-font size.
The OpenAI Agents SDK (`openai-agents`) was a selectable brain harness in the
composer / new-chat / create-agent pickers for bundle YAML agents (polly, debby,
and others). Remove it as a pick by dropping its `harness_labels` entry from the
built-in harness catalog (so `/v1/harnesses` no longer lists it) and from the
static `BRAIN_HARNESS_LABELS` fallback the web merges on top — the web merge only
adds server rows, so both sources must drop it.
It stays a fully valid harness for YAML specs and remains the credential-free
mock harness the integration/e2e suites and the required `Integration
(openai-agents)` CI check depend on: only the UI picker option is removed
(valid_harnesses / harness_modules / capabilities are untouched).
Also update the e2e_ui picker assertion and the unit-test mock seeds to match.
Co-authored-by: Isaac
prepare_claude_cli_path binds part of ~/.claude into the sandbox but not
.credentials.json, where the Claude CLI keeps its OAuth token on Linux. A
host-authenticated user's sandboxed claude-sdk harness saw the account
metadata in ~/.claude.json but not the token, so the CLI reported "Not
logged in". Bind the credential file alongside ~/.claude.json so a host
login works inside the sandbox.
Closes#1922
Signed-off-by: Enes Yilmaz <enesyilmaz5157@gmail.com>
* fix(runner): cancel pending futures after asyncio.wait in _spawn_async_tool
When the cancel event or exec coroutine won first in asyncio.wait(),
the losing future was never cancelled, leaking tasks in long-running
sessions.
* test(runner): regression guard + caveat comments for async-tool future leak
Adds a unit test that drives the real _spawn_async_tool with a stubbed
execute_tool and asserts no asyncio task is leaked on either race outcome
(success: the orphaned cancel_event.wait(); cancel: the orphaned tool coro).
Fails on the pre-fix code, passes with the fix.
Also comments both cancel sites: the cancel-branch note records that
cancelling the task cannot interrupt an underlying asyncio.to_thread, so
that thread may still run to completion.
Co-authored-by: Isaac
---------
Co-authored-by: Dhruv Gupta <dhruv.gupta@databricks.com>
Intelligent routing (`databricks.mas.omnigent.intelligentRouting`) worked for
codex but not claude: claude sessions stayed pinned to Opus instead of being
routed by the judge. The server contract is correct (`if model_override is
None: route()`); two client spots re-pinned a `model_override` and tripped
that guard.
- bindStream: skip the sticky-model handoff PATCH when the session has routing
enabled (`costControlModeOverride === "on"`), so a routing-enabled session
isn't silently re-pinned to the last-used model.
- setCostControlMode: when routing is turned on and a model is pinned, clear
`modelOverride` in the same PATCH (mirrors the new-chat dialog's mutual
exclusion); skip the clear for model-less sessions so no spurious model_change
fires.
Adds tests for the claude-native repro, the same-PATCH clear, and the
no-spurious-clear case.
Co-authored-by: Isaac
* feat(sessions): add server-side (tool, session_name) filter to child-session lookup
Both _find_open_child_by_title and _find_existing_child_session were
fetching all children (100–1000 rows) and scanning in Python to match
by title. Thread the existing title column through a new exact-match
filter so the DB resolves the target in a single indexed query.
* chore: regenerate openapi.json for new child-session query params
* fix(antigravity-native): re-scan on bridge clear to surface deferred gates (#1472)
agy only surfaced the FIRST approval in a conversation; a subsequent gate — e.g.
the 2nd segment of a chained `a && b` run_command, each permission-gated — never
rendered an approval card and the agent hung.
Root cause: the single-in-flight guard in `_maybe_handle_interaction` skips any
new WAITING step while an interaction bridge is in flight, assuming a later
WAITING step is only ever a timeout RETRY of the gate the bridge already owns.
That holds for retries, not for a genuinely-new distinct gate. The deferred step
is never recorded in `state.interacted`, so it could surface later — but only the
poll fallback re-reads the full snapshot; the primary stream path acts only on
frames, and agy emits none while parked awaiting the gate, so the deferral is
permanent.
The guard's one-at-a-time invariant is necessary: `bridge_interaction` delivers
to the freshest WAITING step of a kind (no per-step pinning), so two concurrent
same-kind bridges would mis-target. Rather than weaken it, the bridge done-callback
now RE-SCANS the freshest steps (`_resurface_pending_interaction`) and re-dispatches
them, so a deferred gate surfaces without waiting for a stream frame.
`state.interacted` makes an already-surfaced step a no-op, so the re-scan surfaces
only the not-yet-seen gate and self-terminates, draining a chain of sequential
gates one at a time. Teardown drains the bridge + any chained re-scan tasks to
quiescence.
Tests: a deferred 2nd gate is surfaced via the clear's re-scan; the re-scan
swallows a transient steps-read error; existing guard/clear/teardown tests updated
for the no-op re-scan. Reader suite 80 pass; broader antigravity (by path) 242
pass; ruff + source mypy(strict) clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Isaac
* fix(antigravity-native): pin verdict delivery to the surfaced gate + harden teardown (#1472 review)
Adversarial-review (Codex + Opus) follow-ups on the re-scan-on-clear fix:
- Per-step DELIVERY PIN (Codex BLOCKER / Opus recommended). `bridge_interaction` now
delivers the verdict to the step it was SURFACED for when that step is still
WAITING (new `_waiting_step_at`), falling back to `_freshest_waiting` only when the
captured step is gone — the genuine same-gate timeout-retry. This removes the
unverified "agy never parallel-gates same-kind" assumption: a verdict can no longer
land on a different higher-index gate. The timeout-retry path is preserved
(`test_freshest_waiting_overrides_stale_captured_index` still green).
- Teardown callback flush (Codex). The drain loop yields once per pass
(`await asyncio.sleep(0)`) so a bridge that completed NORMALLY just before teardown
has its `_clear_slot`-scheduled re-scan land in `interaction_rescans` before the
snapshot, instead of escaping the drain and running post-teardown.
- Tests. Add the stream-backstop "case B" (re-scan finds nothing -> a later live
frame surfaces the gate with the slot open), the delivery-pin test (captured-WAITING
beats a distinct higher gate), and an auto-allowed-segment edge case (an
already-allowed command in a chain is DONE / never WAITING -> transparent to the
re-scan, the next real gate still surfaces). Clarify the dedup-race test's intent.
- Docs. Make the sequential-gating assumption explicit in `_resurface_pending_interaction`.
Gemini review was unavailable (Google retired the Gemini Code Assist free tier the CLI
authenticated against). Verified: ruff + mypy(strict, both source modules) clean; the
antigravity suite + tests/runner/test_app_sessions_native.py (229) green; no regressions.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Isaac
* fix(antigravity-native): drain teardown suppresses all task exceptions (#1472 review)
The interaction-bridge teardown drain awaited each cancelled task under
contextlib.suppress(asyncio.CancelledError) only. A drained task that had
already finished with a REAL exception (before the cancel landed) would re-raise
it on await, aborting the drain and leaving the remaining inflight tasks
uncancelled/unawaited (a resource leak). Each task's done-callback already logs
its exception, so the drain now suppresses (asyncio.CancelledError, Exception)
to guarantee it always runs to completion. Surfaced in adversarial review (agy).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Isaac
* fix(antigravity-native): retry the bridge-clear re-scan poll so a transient blip can't strand a deferred gate (#1472 review)
The bridge-clear re-scan is the sole backstop that surfaces a deferred
chained-&& gate on the healthy-stream path (agy emits no frame while parked
and the poll loop is only the stream's failure fallback), so a single
swallowed poll error would re-introduce the permanent hang. Retry the
snapshot read a bounded number of times before giving up.
Co-authored-by: Bryan Li <bryan.li@gmail.com>
Co-authored-by: Isaac
* docs(antigravity-native): trim verbose comments in interaction re-scan code
Condense multi-paragraph inline comments and docstrings in the new
_resurface_pending_interaction / _waiting_step_at / teardown drain
code to the essential why. No logic change.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: SabhyaC26 <sabhyachhabria@gmail.com>
Co-authored-by: Pat Sukprasert <pattara.sk127@gmail.com>
* feat(spawn): add file_ids to sys_session_send schema (#900)
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* feat(server): add lineage-scoped file copy endpoint for subagent file passing (#900)
Add POST /v1/sessions/{session_id}/resources/files:copy. The destination
(child) session copies parent-owned files authorized by spawn lineage:
the source must be the destination itself or an ancestor up the
parent_conversation_id chain. Each file is re-stored as a new
child-scoped row so the child reads its OWN copy — no cross-session read
grant is created, preserving the session-scoping invariant.
Co-authored-by: Isaac
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* feat(runner): forward file_ids from parent to subagent via copy-at-spawn (#900)
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* test(e2e): file passing from parent agent to subagent (#900)
Co-authored-by: Isaac
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(#900): harden file copy — strict-ancestor source, rollback partial copies, delete phantom child
Address codex review findings:
- Reject self as copy source; require a strict parent_conversation_id ancestor.
- Prefetch blobs during validation + roll back created rows/blobs on mid-batch
storage failure, restoring true all-or-nothing semantics.
- Delete the freshly-created server child session when copy-at-spawn fails, so a
failed spawn cannot leave a phantom child that poisons a same-(agent,title) retry.
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* test(#900): update sys_session_send schema assertions for new file_ids field
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(#900): regenerate openapi.json for copy endpoint schema
Docstring reformatting (rst -> markdown) and the sessions ->
session_resources tag move drifted the committed spec from the
generator output, failing the openapi-drift gate. Regenerate to match.
Co-authored-by: Isaac
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(#900): tear down child + defer resource events on copy-at-spawn failure
Two partial-failure bugs surfaced by cross-model (codex) review of the
copy-at-spawn path:
P1 (tool_dispatch): a named send that copied files successfully but then
failed to POST the child message only unregistered runner-local state —
it did not delete the freshly-created child like the copy-failure branch
does. That left a phantom child (poisoning a same-(agent,title) retry)
and orphaned the already-copied child-scoped file rows. Extract the
teardown into `_teardown_failed_child` and call it on every post-copy
failure path so they undo identically.
P2 (sessions copy endpoint): `files:copy` published and persisted
`session.resource.created` inside the per-file loop, before the batch
was known to succeed. A later write failure rolled back the file
rows/blobs but not those events, so clients saw phantom files. Defer all
resource events to a second loop that runs only after every write lands.
Tests: send-failure-after-copy deletes the child; mid-batch write
failure persists zero resource events and no orphan rows.
Co-authored-by: Isaac
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(#900): bound copy-at-spawn — cap files/bytes + stream one at a time
Address PattaraS's blocking review finding on PR #1041: copy_session_files
prefetched every source blob into memory before writing, so a send with many
or large file_ids was an unbounded memory spike on a shared server.
- Cap file count and summed StoredFile.bytes during metadata validation,
BEFORE any blob is read, rejecting an over-limit request with 400 so a
rejected request never buffers a blob.
- Limits are parameterized config knobs (copy_max_files / copy_max_total_bytes
in server_config, defaulting to MAX_COPY_FILES=20 / MAX_COPY_TOTAL_BYTES=256
MiB in content_resolver), overridable per deployment via the YAML config.
- Copy one file at a time (get -> create -> put) so peak memory is a single
blob, not the whole batch; the existing rollback still gives all-or-nothing.
- Tighten the CopyFilesRequest/endpoint docstring to state the source must be
a strict ancestor (self rejected).
Tests: over-count and over-total-bytes rejections assert 400 with ZERO blob
reads (artifact_store.get never called) and nothing copied; at-limit boundary
succeeds. Existing lineage/rollback/self-rejected coverage stays green.
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(#900): enrich copy response + CopyResult dataclass (PR #1041 nits)
Two non-blocking nits from PattaraS's review of PR #1041:
nit #1 — the copy response returned only an id mapping, so the runner
dispatch path did an extra metadata GET per file and guessed content-type
from the filename, even though the true content_type is preserved at copy
time. CopyFilesResponse.mapping now carries {new_id, filename, content_type}
per file (new CopiedFile model); _build_subagent_message_content reads the
type straight from the response — dropping N round-trips — and only falls
back to a filename guess when the source row had no recorded type.
nit #3 — _build_subagent_message_content returned a clunky
tuple[list, None] | tuple[None, str] (value, error) union. Replace it with a
small frozen CopyResult(content, error) dataclass; the single dispatch call
site branches on result.error.
Also regenerated openapi.json for the tightened CopyFilesRequest/endpoint
docstrings (strict-ancestor wording).
Tests: dispatch asserts the content type comes from the copy response with
ZERO per-file metadata GETs, plus a no-content_type→filename-fallback case;
endpoint tests assert the enriched {new_id, filename, content_type} mapping.
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(#900): probe artifact_store.exists during copy validation
Codex review of the cap-and-stream change flagged a regression: moving to
metadata-only validation dropped the original "missing source blob surfaces
before any child row is created" guarantee. A blob that failed mid-stream
(dangling row: metadata present, blob gone) would only surface after earlier
files were already written, leaning on best-effort rollback.
artifact_store.exists() is a cheap metadata probe (S3 HEAD / local stat / DB
row) — NOT a blob read — so calling it in the validation pass restores the
fail-before-any-write guarantee without reintroducing the batch prefetch or
spiking memory.
Test: a source whose blob was deleted (row intact) → 404 with nothing copied.
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(files): address review feedback
---------
Signed-off-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
Co-authored-by: praneeth_paikray-data <praneeth.paikray@databricks.com>
* fix(web_fetch): run __web_researcher on the parent leg's harness
web_fetch does not fetch directly: it dispatches a synthetic __web_researcher
sub-agent that runs curl via sys_os_shell. build_researcher_spec built that
child as a bare ExecutorSpec(max_iterations=5), copying only the parent's llm
and dropping the parent's executor harness, auth, and model. With executor.type
defaulting to "omnigent" and an empty config, every fetch broke on every leg:
- Layer 1 (active): executor.harness_kind (config["harness"] or type) resolved
to the literal "omnigent", so the runner aborted the researcher spawn with
`RuntimeError: unknown harness 'omnigent'` before any model routing.
- Layer 2 (latent): with the parent's harness and auth gone, a gateway model
such as z-ai/glm-5.2 fell through to the in-process native router
(`Unknown provider 'z-ai'`), and the codex/claude legs failed on missing
credentials.
PR #817 reconstructs the researcher on a resolve-miss but calls the same
build_researcher_spec, so the bug persisted.
Fix: inherit the parent executor fields the harness spawn-env builders actually
read on the claude-sdk/codex/pi legs — config["harness"] (selection;
runner/app.py:8691,18601), model (_resolve_spec_model; workflow.py:1115), and
auth (_resolve_provider_for_build; workflow.py:1040) — plus type, the executor
discriminator. connection (rides on llm), context_window (auto-detected), and
the deprecated Databricks profile (subsumed by auth) are not read on these legs
and are omitted. os_env carried inside executor.config is an inline-sub-spec
artifact superseded by the explicit os_env, so it is dropped.
A parent's real harness can also live only in resolved session state (an API
harness_override on a spec with no config["harness"]); that is not visible at
the build_researcher_spec call sites (WebFetchTool.__init__ and the
_find_spec_by_name resolve-miss), and the researcher child never carries an
override. Rather than emit a child that the runner aborts with the cryptic
unknown harness 'omnigent', fail loud at build time with an actionable
OmnigentError naming the parent leg.
Add regression tests: the reconstructed spec carries the parent's
harness/auth/model (not the bare type=="omnigent"/no-harness spec); the inline
executor.config os_env is dropped; a no-harness parent raises the clear error.
Signed-off-by: Vadim Comanescu <vadim984@gmail.com>
* docs(web_fetch): trim verbose build_researcher_spec comments
The inline commentary in build_researcher_spec had grown to multi-paragraph
blocks with file:line references. Condense to the essential why (inherit the
parent leg's routing fields; fail loud on no bootable harness) per the repo's
comment guidance. No logic change.
---------
Signed-off-by: Vadim Comanescu <vadim984@gmail.com>
Co-authored-by: Pat Sukprasert <pattara.sk127@gmail.com>
Reorder is no longer an optional follow-up — drag-to-reorder (grip handle,
within-conversation) shipped, so the actions table reflects it.
Update the per-harness steer table from this session's code audit: cursor-,
pi-, hermes-, opencode-native all report supports_live_message_queue = True
(opencode via supports_enqueue=True through NativeServerHarness), so the steer
button is honored on all of them. opencode-native is settled — its app server
has no live-steer endpoint, so a steered message is admitted as a new prompt
and promoted by the server's own queue at the next turn boundary.
Narrow the TODO: the delivery mechanism is now code-confirmed for every native
harness; what remains is upgrading the app-defined (mid-turn vs next-turn)
rows via a LIVE steer per harness — confirmed live only for claude-/codex-
native so far.
Co-authored-by: Isaac
A web-UI message injected while Claude Code is mid-turn still rendered a
spurious "terminal did not become ready within 30s" runtime-error card
when many subagents ran concurrently. The readiness gate scans for the
`❯` input glyph; PR #2001 widened the scan to an 8-line box-rule-framed
window to clear a one-subagent footer, but a subagent fan-out adds one
`○ Explore …` row per concurrent subagent, so the footer height is
unbounded — five subagents push `❯` to the 12th line from the bottom,
past the fixed window, and the gate times out.
Drop the fixed framed window: scan all visible non-empty lines for a `❯`
that has a box rule below it. The box rule (the input box's closing
`────` frame) is a reliable structural signal at any depth, and
`capture-pane -p` returns only the visible pane, so the scan stays within
one screen. The scrollback-echo false positive stays rejected — an echoed
`❯` never has a box rule beneath it.
Co-authored-by: Isaac
A sparkle button inside the "Git worktree branch" input fills a unique
"worktree-<hex>" name (crypto.randomUUID), so users can spin up a
throwaway worktree without inventing a branch name.
Co-authored-by: Isaac
Adds an explicit policies.scope column ('default' | 'session') so queries
can filter by column value instead of checking session_id IS NULL — the same
pattern used for agents.kind (o1a2b3c4d5e6). Includes a SQLite-safe Alembic
migration (q1a2b3c4d5e6) with back-fill, a partial unique index on default
policy names, and corresponding store, entity, and test updates.
* feat(web): make sidebar Search open the command palette
The sidebar's "Search sessions" box was an inline filter that only
narrowed the visible list. Session search (title + chat content) already
lives in the ⌘K command palette, so point the box at it instead of
duplicating a weaker filter.
- Sidebar: replace the search input with a "Search" button that opens the
palette, showing a ⌘K badge on hover/focus. Drop the inline
searchQuery/debounce state; the list is now unfiltered.
- CommandPalette: list Sessions above Actions (the palette doubles as the
session-search entry point). Cap the session list to 5 while the query
is empty so Actions stays visible without scrolling; typing lifts the
cap. Indent session rows to align with the icon-prefixed actions.
Placeholder → "Search sessions or run a command".
- AppShell: wire the button to the palette; mount the palette in embedded
mode too (the ⌘K hotkey stays disabled there).
Co-authored-by: Isaac
* test(e2e-ui): regenerate visual baselines
* test(e2e-ui): retarget sidebar search tests to the command palette
The sidebar's "Search sessions" input became a "Search" button that opens
the command palette, so the two E2E tests that located the old searchbox
were failing.
- test_sidebar_hotkeys: probe sidebar collapse/expand width via the
"Search" button (data-testid=sidebar-search-button) instead of the
removed search input.
- test_sidebar_search: drive the server-side search round-trip through the
palette (opened from the Search button) — matching query lists the
session, non-matching empties it — the same chain the old inline filter
exercised.
Co-authored-by: Isaac
* test(e2e-ui): fix sidebar search tests for the palette (verified locally)
The first retarget pass had two real bugs, both now reproduced and fixed
against a local live server + Chromium:
- test_bracket_chord: the collapse probe measured the search control's
width, but the new Search button (a flex item, min-width:auto) floors at
its content width and stays 260px on collapse — the old input shrank to
0. Probe the sidebar <aside> width instead; it's what the chord animates.
- test_sidebar_search: the session title also renders in the chat header
(the test is on /c/{id}), so a page-wide text match never reached zero.
Scope both palette assertions to the dialog.
Co-authored-by: Isaac
---------
Co-authored-by: omnigent-ci[bot] <294685417+omnigent-ci[bot]@users.noreply.github.com>
* feat(web): select an existing git worktree when starting a session
The new-session worktree field previously only created a new worktree
off a branch name, and picking a directory that was already an existing
worktree errored ("branch already exists"). This adds first-class
support for starting a session directly in an existing worktree.
The branch input is now a combobox: focusing it lists the repo's
existing worktrees, typing filters them, picking one starts the session
in that worktree (no git opts sent — so no branch-already-exists guard),
and a name matching none creates a new worktree as before. A concise
warning flags that the session starts in an existing worktree.
Backend adds a read-only list_worktrees host git op, the matching
list_worktrees tunnel frame pair, a server proxy, and
GET /hosts/{id}/worktrees (owner-scoped; non-git path → 400 → empty
list in the picker), mirroring the existing create/remove worktree
plumbing.
Co-authored-by: Isaac
* fix: prettier-format worktree UI + regenerate openapi.json
CI caught two gaps: the new worktree combobox files weren't
prettier-formatted, and the new GET /hosts/{id}/worktrees route made
the checked-in openapi.json stale. Regenerated via scripts/dump_openapi.py.
Co-authored-by: Isaac
* test(e2e-ui): cover selecting an existing worktree in start-session
Drives the branch combobox end-to-end: focusing it lists the repo's
existing worktrees (stubbed GET /hosts/{id}/worktrees), selecting one
points the workspace at that dir and sends no git spec on create.
Mirrors the existing test_start_session_add_worktree harness.
Co-authored-by: Isaac
Native Claude sessions stayed "busy" in the web UI (composer stuck on
Stop) after a /model switch, even though the terminal was idle. It
self-healed only on the next real message.
A surfaced CLI built-in (/model, /effort) becomes a slash_command
transcript item that opens its own response id but runs no LLM turn, so
no Stop hook ever fires to close it. The forwarder's turn-start edge
still published an id-bearing running for it, which opened a streaming
activeResponse in the web store; the store suppresses the trailing bare
PTY idle while a response is streaming, so nothing cleared it.
Gate the turn-start running edge on the turn actually having assistant
output (a function_call or assistant message) — the exact turns a later
Stop/StopFailure hook will close. Turns that produce no LLM output
(slash_command, or terminal_command from !cmd) no longer strand the UI
busy. A skill that does trigger an LLM turn shares its id with the
assistant text it produces, so running still fires one poll later when
that output appears.
Co-authored-by: Isaac
* refactor(db): remove all FK constraints; application owns relationship cleanup
Drops all 9 FK constraints (8 CASCADE + 1 SET NULL) from the SQLAlchemy
models and adds a new Alembic migration (p1a2b3c4d5e6) to remove them from
the live schema, following internal DB standard Rule R032.
- db_models.py: remove ForeignKey() from session_permissions.user_id,
session_permissions.conversation_id, conversations.parent_conversation_id,
conversations.root_conversation_id, conversations.agent_id,
conversations.host_id, conversation_items.conversation_id,
conversation_labels.conversation_id, and policies.session_id.
- migration p1a2b3c4d5e6: upgrade drops all FKs via batch_alter_table
(recreate="always" on SQLite); downgrade re-adds them.
- delete_conversation: now collects the full conversation subtree via a
recursive CTE and explicitly deletes items, labels, comments, policies,
and session-permissions for all descendants before deleting conversation
rows, replacing the previous reliance on ON DELETE CASCADE.
- switch_conversation_agent: removes the defensive null+flush of agent_id
before deleting the old session-scoped agent, since there is no longer
a CASCADE constraint that would destroy the conversation row.
* test(db): update tests for FK removal; fix migration and ORM cascade assertions
- Fix migration p1a2b3c4d5e6 to correctly drop all FKs on SQLite by
reflecting actual constraint names (including unnamed/None FKs that get
convention-derived names during batch rebuild) and drop_constrainting each.
Restore host_id FK in downgrade as fk_conversations_host_id_hosts to match
the original name so subsequent migrations can find it.
- Restore row.agent_id = None + flush before deleting old agent in
switch_conversation_agent so SQLAlchemy ORM identity map stays consistent.
- Update ORM cascade tests to assert new no-FK behavior (children survive
parent deletion; app must clean up explicitly).
- Update migration_workspace test to document that host deletion no longer
auto-nulls conversations.host_id without a DB FK.
- Update permission store cascade test to document that permissions persist
after conversation deletion without DB FK cascade.
- Update agents migration FK test to document that referential integrity is
now the application's responsibility.
* fix(db): explicit cleanup in delete_user and delete_host after FK removal
delete_user now explicitly deletes session_permissions rows before
removing the user row — without the DB CASCADE, orphaned permissions
could grant access to a re-created account with the same identifier.
delete_host now explicitly nulls conversations.host_id for any sessions
still bound to the host before deleting the row — replaces the removed
ON DELETE SET NULL FK behavior. Also updates stale FK-reference comments.
* feat(harness-bench): rich live progress, --jobs parallel, --report file
Three CLI/output improvements, built on a structured progress-event seam.
- Structured events (events.py): the orchestrator now emits typed BenchEvents
(HarnessStarted/Skipped, ProbeStarted/Finished, HarnessFinished) to a
ProgressSink, instead of pre-rendered strings. The old per-line output is
preserved via LineSink, and a bare-callable `progress=` is auto-adapted to
it — back-compat, no caller change required.
- Rich live table (richreport.py, --rich/--no-rich): a ProgressSink backed by
rich.Live draws one row per harness with per-dimension cells that fill in as
probes finish (spinner while running → verdict glyph). Auto-selected on a
TTY when rich is available; falls back to LineSink under a pipe/CI or when
rich is absent (rich_sink_or_none returns None). Most useful with --jobs.
- Bounded parallel (--jobs N / -j, default 1): run up to N harnesses
concurrently via an asyncio.Semaphore. Probes WITHIN a harness stay
sequential (they share one driver/session with a single in-flight turn);
concurrency is only across harnesses, each of which owns its own
server/runner. gather preserves input order, so the matrix stays in
--harness order regardless of finish order. The cap keeps process/port and
gateway load bounded rather than spawning every harness at once.
- Report file (--report PATH): write the final matrix to a file; format from
--json/--markdown, else inferred from the extension (.json/.md), else a
plain (un-colored) grid.
Tests: structured-event emission + LineSink adaptation, --jobs order
preservation under staggered finishes, and --report file writing (md + json).
Offline suite 55 passed / 14 skipped, ruff clean. rich renders live when
present; the plain path is unchanged.
* feat(harness-bench): share one server+runner across parallel full-server harnesses
Folds the shared-server optimization into the parallel path. Previously each
full-server harness spawned its own server + runner; under --jobs > 1 that was
N server boots + N runners. The Omnigent server is multi-agent/multi-session
and a single runner resolves the harness per session from its agent spec, so N
SDK harnesses can share ONE server+runner, each registering its own agent +
session.
- New SharedFullServer (full_server_driver.py): owns the server+runner
lifecycle + agent/session registration, extracted from FullServerDriver.
- FullServerDriver takes an optional `shared=`: injected → registers on the
shared server and spawns nothing; None → owns a private SharedFullServer
(back-compat, exactly the old one-server-per-harness behavior for --jobs 1).
- run_bench stands up one SharedFullServer for a live, parallel run with >1
full-server harness (via _maybe_shared_full_server), passes it to each, and
tears it down after. native-tui harnesses still self-provision (each needs
its own host daemon).
Cuts the heaviest, slowest part of full-server startup (server boot +
health-wait) from N times to once, and roughly halves the process/port count
for a parallel SDK run. Gateway load is unchanged (same total turns).
Test: a parallel full-server run builds exactly one SharedFullServer and all
harnesses register on it. Offline suite 56 passed / 14 skipped, ruff clean;
solo full-server path unchanged (back-compat).
* refactor(harness-bench): split shared server into its own module; hoist imports
Readability/structure cleanup requested in review, no behavior change.
- Split full_server.py out of full_server_driver.py: the server+runner
lifecycle and agent/session registration (SharedFullServer + spawn/wait/
config helpers + the shared _find_free_port/_mint_bearer/spawn_omnigent_server
that native-tui also uses) now live in full_server.py; full_server_driver.py
keeps just FullServerDriver and its probe/item-scan helpers. Clear seam:
"the server" vs "the driver that runs probes against it".
- Hoist function-body imports to module top across the package (Any, shutil,
cli_unavailable_reason, omnigent.harness_capabilities/plugins, LineSink,
SharedFullServer, socket/io/tarfile/yaml). The only inline imports left are
intentional and now commented: the optional `rich` dependency (richreport +
its lazy load in __main__) and two documented cycle-avoidance imports
(transport→drivers, profile→manifest).
- Update consumers (native_tui_driver, bench) to import the shared helpers
from full_server; fix the shared-server test to patch bench's namespace
(bench now imports SharedFullServer at top).
Offline suite 56 passed / 14 skipped, ruff clean, no import cycle.
* feat(harness-bench): default SDK harnesses to full-server; add --fast
Full-server is a strict coverage superset for SDK harnesses: it observes
everything sdk-inproc does (basic / streaming / interrupt / model-override)
*plus* the two dimensions sdk-inproc physically cannot reach — Tool calling
and Policy DENY, as server-dispatched, policy-gated calls. The only cost is
the server boot. So make full-server the default and offer --fast as the
opt-out, rather than a per-harness --best selector.
Transport is now resolved from the harness *family* + flags
(resolve_transport_name):
- SDK family (sdk-inproc/full-server) -> full-server by default; --fast picks
sdk-inproc (skips the boot; Tool calling + Policy DENY then report SKIPPED,
which those probes already emit on the wrap-direct path -- no false DRIFT).
- native (native-tui) -> single transport; --fast does not apply.
- --transport NAME still overrides the family for any harness, and is mutually
exclusive with --fast.
The profile's `transport` field stays the family marker (the _is_native
applicability gate keys on it), so nothing about probe applicability changes.
--list now prints the resolved default transport so it matches what runs.
Both driver gates already agree with this: FullServerDriver.unavailable only
rejects native profiles (not sdk-inproc-family), and SdkInprocDriver accepts
its own family -- so neither default nor --fast self-rejects.
Docs (harness-bench-design.md) updated: transport-selection prose, the
which-transport-exercises-what table, and the run examples now lead with the
full-server default and --fast opt-out.
Offline suite 57 passed / 14 skipped, ruff clean.
* fix(harness-bench): quiet expected provisioning skips; keep tracebacks for bugs
A parallel live run dumped three full tracebacks for the own-auth natives
(goose/kimi/hermes) whose forwarder never wires up — an expected, already-
handled skip (they show as skipped in the matrix), but the stack dumps break
up the --rich table and read like failures.
Introduce ProvisioningError (in driver.py) for an *expected* provisioning
failure: a known-unrunnable environment through no fault of the bench, e.g. an
own-auth native whose vendor CLI is installed but not logged in. native-tui's
forwarder-timeout now raises it instead of a bare RuntimeError.
run_harness splits on it: an expected ProvisioningError logs one INFO line
(reason only, no traceback), while any other exception keeps exc_info=True so a
genuine driver bug (e.g. an AssertionError) can't vanish behind a green skip.
The matrix output is unchanged either way — the harness is still a
capability-neutral skip with the reason shown in its row.
Offline suite 58 passed / 14 skipped, ruff clean.
* feat(harness-bench): label each matrix row with its resolved transport
Show which transport actually produced each row, e.g. `claude-sdk
[full-server]`, `kimi-native [native]`. This matters now that transport is
resolved from family + flags: an SDK harness's profile.transport is the
`sdk-inproc` family marker, but it runs on `full-server` by default -- so the
label reflects the *resolved* transport, not the marker, or it would mislabel
exactly the rows worth clarifying.
- HarnessReport carries the resolved `transport` (the driver class's transport,
or the resolve_transport_name result offline). Populated at every report site
(success, unavailable-skip, provisioning-skip, offline).
- report.py labels the harness column in both the terminal and Markdown
renderers (native-tui abbreviated to `native`); render_json adds a distinct
`resolved_transport` field alongside the family `transport`.
- The rich live table labels its rows too: HarnessSkipped gained a transport
field (HarnessStarted already had one), and the sink tracks harness→transport.
Offline suite 58 passed / 14 skipped, ruff clean.
* docs(harness-bench): refresh README for phase-2 state
The README still described the phase-1 MVP (sdk-inproc only, four SDK
harnesses, Markdown/JSON output). Bring it current:
- Run examples lead with --jobs + --rich; add a Flags section covering
--fast, --transport, --jobs, --rich/--no-rich, --report.
- New "Transport selection" section: full-server is the SDK default (fullest
coverage), --fast opts down to sdk-inproc, natives use native-tui.
- Note the per-row transport label and that Tool calling / Policy DENY only
get a real verdict on full-server.
- Layout table lists the current modules (transport.py, full_server.py split
from full_server_driver.py, native_tui_driver.py, events.py, richreport.py).
- Scope reflects what is live (3 transports, all natives auto-derived) vs the
remaining open items, instead of "phase-1 MVP".
* docs(harness-bench): clarify native Tool calling / Policy DENY is a bench gap
A reader skimming the matrix could misread the `·` in the native rows'
Tool calling / Policy DENY cells as "native harnesses can't do this". They
can -- the bench just cannot observe it on native-tui yet.
Sharpen both docs to say so unambiguously:
- A `·` always means "the bench did not measure this here", never "the harness
lacks it".
- The native-tui `·` for those two dimensions is a driver/observation gap, not
a native-harness limitation: a native tool call is the vendor's own
(Bash/Read/...) and a native deny is a vendor permission decision, neither of
which is the server-dispatched, policy-gated call the probe watches for.
- The which-transport table cells now read "bench can't observe vendor tools/
deny yet" instead of the terse "not yet wired"; the open-items entries lead
with "bench observation ... a driver gap, not a native-harness limitation".
No behavior change; docs only.
* fix(harness-bench): treat any native provisioning failure as a quiet skip
The earlier quieting only covered the forwarder-timeout RuntimeError. A native
harness can fail provisioning other ways -- goose-native's terminal-ensure
returns a 500 (the vendor cannot start a thread), which raised a raw
httpx.HTTPStatusError and still dumped a full traceback.
Native provisioning drives a live vendor CLI plus a server-native terminal, so
any HTTP failure there is an environment/server-state gap, not a bench bug.
NativeTuiDriver.__aenter__ now converts httpx.HTTPError into ProvisioningError
so the orchestrator skips the harness quietly (reason shown in its row). A
programming error (AssertionError, etc.) is not an HTTPError, so it still
propagates with its traceback. The deliberate readiness-timeout and
agent-not-seeded raises in the provisioning path also became ProvisioningError
for consistency.
Test: an httpx 500 in provisioning surfaces as ProvisioningError. Offline suite
59 passed / 14 skipped, ruff clean.
* test(harness-bench): single import style in test_bench (review)
Code-quality review flagged tests.harness_bench.bench being imported both as
`from ... import run_bench, run_harness` (top level) and `import ... as
bench_mod` (in three test bodies). Drop the in-function module aliases and
patch module attributes via monkeypatch's string-target form
(`"tests.harness_bench.bench.resolve_driver_class"`), which the file already
uses elsewhere -- so there is one import style throughout.
No behavior change. Offline suite 59 passed / 14 skipped, ruff clean.
* fix(harness-bench): don't reprint the grid under --rich on a terminal
Running `--rich` interactively showed the matrix twice: the rich live table
(progress, on stderr) and then the plain report grid (deliverable, on stdout),
which land on the same terminal and look like a duplicate.
The report is not pure duplication -- it carries the legend, per-cell Notes,
and any Drift section the rich table omits. So the fix keeps the footer and
drops only the grid, and only when it would actually duplicate:
- render_table gains grid=True/False; grid=False emits just the footer
(legend/drift/notes/skips), no heading or glyph rows.
- Sinks expose drew_grid (rich live table True, LineSink False). The CLI prints
grid=False only when the sink drew the grid AND stdout is a TTY (same
terminal as the stderr progress). Redirect stdout to a file and the report
keeps the full grid, so the file stays self-contained.
Tests: grid=False drops the grid but keeps the legend; _grid_already_shown is
True only for a grid-drawing sink. Offline suite 61 passed / 14 skipped, ruff
clean. README output-format note updated.
Drop the back-pointer `agents.session_id` column (FK to
`conversations.id`) in favour of the forward pointer
`conversations.agent_id`, which was already the canonical source of
truth. An agent is now classified as session-scoped if any conversation
row references it via `conversations.agent_id`, discovered at query time
with a NOT EXISTS subquery rather than a nullable FK column.
- Remove `session_id` from `SqlAgent`, `Agent` entity, and the
`sql_agent_to_entity` converter.
- Rewrite `get_by_name` and `list` template-agent filters from
`session_id IS NULL` to `NOT EXISTS (SELECT … FROM conversations …)`.
- Drop the partial unique index `ix_agents_template_name` (was scoped
to `session_id IS NULL`) and recreate it as a plain unique index;
drop `ix_agents_session_id`.
- Add Alembic migration `o1a2b3c4d5e6` with upgrade/downgrade paths.
Queued messages could be steered, edited, or deleted, but not reordered —
the queue drained strictly in enqueue order. Add drag-to-reorder so the
user can change the order their held follow-ups will send in.
Each strip row gains a grip handle (shown only when reordering is wired);
dragging it reorders via @dnd-kit/core primitives — the same pointer
sensors the sidebar uses (5px mouse activation, so a grip click still
reaches the row's steer/edit/delete buttons). A dedicated handle rather
than a whole-row drag keeps those buttons clickable.
New reorderQueuedMessage(queueId, beforeQueueId) store action does the
move. queuedMessages is one flat array interleaving conversations, so it
reorders only within the dragged message's own conversation and refills
that conversation's absolute slots — other conversations' entries keep
their positions. No-ops on a missing id, a self-move, or a cross-
conversation target.
Tests: store reorder (before/end, no-op identity, interleaved-queue slot
preservation, cross-conversation guard) and the strip's grip affordance
gating on onReorder.
Co-authored-by: Isaac
Polly flagged a duplicate-upload leak on #2065 that also pre-exists in
send(): when a message with attachments retries after a post-phase failure
(background flush re-queues on a cooldown; send() is retried by the caller),
the retry re-uploads every File from scratch, orphaning the blobs the first
attempt already stored server-side.
Add a shared uploadFileBlock(sessionId, file) helper that memoizes each
File's successful upload (WeakMap keyed by File, then by session) and
returns the cached content block on a retry instead of re-uploading. Wire
both send() and flushBackgroundQueues through it. The WeakMap auto-releases
once the File is dropped from the queue/pending state.
Tests: a send() retry after a failed post reuses the cached file_id (one
upload, not two); the background-flush retry does the same and the posted
message still carries the original id.
Co-authored-by: Isaac
* feat(web): background-flush queued messages with attachments
Background cross-session flush previously skipped any queued message that
carried files, leaving it for the foreground flush — so an image queued in
a navigated-away conversation sat until the user returned.
Mirror send()'s two-phase sequence in flushBackgroundQueues: upload each
attachment via uploadFile (→ real file_id), build input_image/input_file
blocks, then post the message referencing them via postEvent. Both awaits
sit under the one in-flight guard and the one catch, so a failure in either
the upload or the post phase re-queues the head (FIFO-preserving) and sets
the same cooldown — no separate guard, no double-send.
Removing the files skip also closes the head-blocking edge: an image at the
head of an idle conversation's queue now drains instead of stalling the
text messages behind it.
Tests: upload-then-post emits an image block with the real file_id and
clears the queue; an upload-phase failure posts nothing and re-queues.
Co-authored-by: Isaac
* test(e2e): background-flush a queued image to its origin session
Adds a cross-session e2e alongside the text one: attach an image + text to
B while B is busy (held POST), switch to idle A, release B. Asserts the
background flush uploads the image to B then posts an input_image block
carrying the returned file_id — and that neither the upload nor the message
leaks into the active session A.
Covers the two-phase upload→post path end-to-end (the unit tests cover it
at the store level); shares the seeded_session_pair fixture and route-mock
harness with the text test.
Co-authored-by: Isaac
* feat(web): keep the working indicator lit for the whole turn, rotate its label
The Otto + shimmer "Working…" indicator was hidden the moment an assistant
bubble began streaming, so long tool runs and reasoning gaps looked stalled.
Keep it lit for the entire busy turn (only a trailing compaction spinner still
suppresses it), and rotate its label through a short pool for variety.
- shouldShowWorkingIndicator no longer hides on a streaming bubble; drop the
now-unused hasInProgressAssistantBubble helper.
- Add useWorkingLabelTick: one shared wall-clock timer (useSyncExternalStore)
so both render sites rotate in lockstep. ROTATE_MS = 1 minute.
- workingIndicatorLabel(bgCount, tick) cycles WORKING_MESSAGES (7 labels,
index 0 = "Working…"); background-task counts still take priority.
- Keep the pinned pill's aria-live announcement stable at "Working…" while
only the visible tab text rotates, so screen readers aren't re-announced.
Reduced motion needs no change: the shimmer sweep and Otto bob already freeze
via CSS, and the label is a JS text swap so it keeps rotating.
Co-authored-by: Isaac
* fix(web): address PR review — drop "Thinking…" label, fix e2e assert
Review follow-ups on #2006:
- Remove "Thinking…" from WORKING_MESSAGES — it carries a specific
reasoning/thinking meaning in the LLM context (per @daniellok-db).
- Update the background-task e2e (test_background_task_indicator_label_lifecycle)
now that the running-turn label rotates: assert on the trailing ellipsis
every rotating label shares (the background-task text has none) instead of
the literal "Working", so it's robust to which pool entry the wall-clock
bucket lands on.
Co-authored-by: Isaac
* test(e2e): match working label against the pool, not the ellipsis
Per review follow-up: assert the running-turn indicator shows one of the
actual rotating labels (regex alternation over the WORKING_MESSAGES mirror)
rather than the trailing ellipsis. A commented _WORKING_LABELS constant
mirrors the web pool and must stay in sync if it changes.
Co-authored-by: Isaac
---------
Co-authored-by: Anthony Ivan <anthony.ivan@example.com>
* feat(cli): add omni session export --id <session_id> command
Closes#1623
* test(cli): add unit tests for omni session export
* fix(test): rename l -> line to fix E741 ambiguous variable name
* feat(cli): switch session export to use server API via --server
* fix(cli): pass auth headers to session export HTTP client
The `build codex-parity sidecar` job recompiles the Rust sidecar (~1100
crates, ~7 min cold) on nearly every PR run. The old `Cache Rust build`
step cached the whole 1.6 GB `--target-dir` keyed on `Cargo.lock`, but:
- The job triggers only on `pull_request`, so every cache is scoped to
`refs/pull/NNNN/merge`. GitHub only lets a PR restore caches from its
own ref or the base branch (main), and this workflow never writes a
main-scoped cache -- so no PR can ever restore another's. Every first
run is a guaranteed cold miss.
- Each 1.6 GB entry churns out of the 10 GB repo cache under LRU, so
even same-PR re-runs frequently miss.
- Even on a target-dir hit, Cargo re-fingerprints and rebuilds anyway.
Mirror the fix#2016 applied to ci.yml's codex-parity job: cache just
the ~10 MB binary, keyed on `sidecar/**` + the rustc version, and skip
`cargo build` on a hit. This uses the SAME key as ci.yml, which runs on
push to main -- so the main-scoped `codex-parity-bin` cache ci.yml
produces is now restorable by this PR-only workflow. Warm runs drop from
~7 min to the artifact download/upload (~15-25s). The key self-
invalidates when the source, Cargo.lock, or toolchain changes.
Co-authored-by: Isaac
* feat(web): background cross-session flush of queued messages
A message queued in conversation B now flushes when B goes idle, even
while the user is viewing a different conversation A — previously it sat
until the user returned to B (navigating away aborts B's SSE stream, so
the foreground flush couldn't see B's status).
New flushBackgroundQueues store action: for each conversation with queued
messages that isn't the active one, read its status from the live
["conversations"] cache (kept fresh by the WS session-updates overlay +
poll) and, if idle, POST the head via postEvent — a stateless primitive
that touches no active-session state (no optimistic bubble; it re-hydrates
on return). One message per idle conversation per call (FIFO); re-queues
on POST failure to retry. Text-only for now — attachments are left to the
foreground flush (tracked in the code comment).
A new app-wide QueueFlushProvider triggers it on queue changes and on any
["conversations"] cache change (the signal a navigated-away conversation
went idle). The foreground maybeFlushQueuedHead still owns the active
conversation; the two are complementary.
Updates the cross-session routing e2e: it now asserts the queued message
is delivered to its origin B via background flush (never leaking to the
active A) — closing the loop the pre-queue test guarded.
Co-authored-by: Isaac
* fix(web): bound background-flush retries on persistent POST failure
Polly review flagged an unbounded retry storm: on a persistent POST
failure the head is re-queued, which mutates queuedMessages and re-fires
QueueFlushProvider's effect; the failed POST leaves the conversation idle
in the cache, so it flushes → POSTs → fails → re-queues → … with no
backoff, hammering /v1/sessions/{id}/events.
Add a module-level throttle (kept out of store state so it can't
re-trigger the effect): skip a conversation that is mid-POST or within a
5s post-failure cooldown. Also re-queue a failed head ahead of its own
successors instead of at the tail, preserving per-conversation FIFO.
Tests: cooldown blocks an immediate re-POST of a just-failed conversation;
a failed head lands back in front of its successor.
Co-authored-by: Isaac
* feat(web): add UI font family setting to Appearance
Add a font-family control to Settings → Appearance, beside the font-size
stepper. It's a free-text field (Cursor-style): type any font installed on
this device; leave it blank for the system default. The choice re-fonts the
whole UI chrome, is persisted per-device in localStorage, and is applied
before first paint so a reload doesn't flash the default.
Implementation mirrors the just-merged font-size setting (#2040). It can't
reuse --font-sans: Tailwind v4's @theme inline block inlines the literal
stack into the font-sans utility rather than a var() reference, so a runtime
--font-sans override is a no-op. Instead the html rule reads
font-family: var(--ui-font-family, var(--font-sans)), and the preference
module sets --ui-font-family on documentElement — unset falls back to the
existing system stack. The theme picker and font-size stepper are unchanged.
The two .font-heading elements (dialog/card titles) resolve font-family:
var(--font-sans) directly, so they keep the system stack rather than the
custom family — acceptable for this UI-chrome-only change.
Co-authored-by: Isaac
* fix(web): keep font-family input inline; ruff-format e2e test
- The Font family row's longer description pushed the input onto its own
line under flex-wrap. Give the text column min-w-0 flex-1 and the control
shrink-0 so the input stays flush-right on the same row as the label,
matching the font-size stepper above it.
- Apply ruff format to the new e2e test (one-line test signature) so the
Pre-commit CI check passes.
Co-authored-by: Isaac
* fix(web): right-align font-family input with the font-size stepper
Move the Reset button to the left of the input so the input is the
rightmost element in its group; its right edge now lines up flush with
the font-size stepper above it (both at the row's right edge). Reset
stays `invisible` (not removed) at the default so the row doesn't shift.
Co-authored-by: Isaac
* fix(web): keep code surfaces on the mono font, immune to the UI font setting
The UI font-family setting is UI chrome only. Pin the Monaco editor and
xterm terminal roots (.monaco-editor, .xterm) to var(--font-mono) so the
--ui-font-family override can't leak into code surfaces through an unpinned
descendant. Editor/terminal code fonts are intended for a separate, future
code-font setting.
Both surfaces already pin their own font (xterm via its JS fontFamily
option, Monaco via its inline default), so this is a defensive guard;
verified live that with a UI font override active, .xterm/.xterm-screen and
the Shiki code viewer all stay on the mono stack.
Co-authored-by: Isaac
* fix(web): fall back to the default sans for unknown/partial font names
Applying a bare `--ui-font-family: <name>` meant that a font that isn't
installed — or a partial name while the user is still typing — left the
browser with an unresolvable family and no fallback, so the UI dropped to
the browser's default serif (Times) instead of the app's sans.
Append the system stack to the applied value (`<name>, var(--font-sans)`)
so an unusable name degrades to the default sans. The CSS-level
`var(--ui-font-family, …)` fallback only fires when the property is unset,
not when it holds an unusable value, so the fallback must live in the value
too. localStorage still stores just the raw name (the input shows it
verbatim). Verified live: partial/uninstalled names now render as the
default sans, not serif.
Co-authored-by: Isaac
* test(e2e): assert font-family starts with the chosen name
The applied --ui-font-family now leads the chosen family and appends the
system stack as a fallback, so getComputedStyle resolves the custom
property to the full stack (e.g. "Georgia, ui-sans-serif, ..."). Assert the
resolved value startswith the typed name rather than equals it. The
reset/empty assertions are unchanged (property removed → empty).
Co-authored-by: Isaac
* feat(web-ui): global command palette (⌘K)
Add a cross-platform command palette opened with ⌘K (Ctrl+K on
Windows/Linux), with two groups:
- Actions: New chat, Go to Inbox/Settings, toggle the conversations and
workspace sidebars, and open the keyboard-shortcuts dialog. Filtered
client-side against the query.
- Sessions: fuzzy session switching from the same server-search source the
sidebar uses (useConversations → GET /v1/sessions?search_query=),
debounced, so the palette finds sessions beyond the first page rather than
client-filtering one page. Archived excluded, matching the sidebar default.
The hotkey is bound once in AppShell and bails when focus is inside an xterm
terminal or the Monaco editor (both own ⌘K), and is disabled in embedded
mode where ⌘K belongs to the host page. The desktop (Electron) app loads the
same SPA and binds only ⌘N/⌘F natively, so ⌘K reaches the renderer unchanged.
Adds an 'Open command palette · ⌘K' row to the keyboard-shortcuts dialog, a
ResizeObserver test polyfill cmdk needs under jsdom, colocated Vitest
coverage, and a Playwright e2e (tests/e2e_ui/sessions/test_command_palette.py).
Signed-off-by: Dimitar Dimitrov <dimitardimitrov9205@gmail.com>
* feat(web-ui): reuse UI icons in command palette, drop shortcuts action
Give each palette Action the same icon as its equivalent button
elsewhere in the UI (new chat, inbox, settings, sidebar toggles) so the
palette reads as a shortcut to those surfaces. Icons inherit the item's
foreground color rather than the muted tone, matching the label text.
Remove the "Keyboard shortcuts" action — the palette is for imperative
commands, not opening an informational dialog. Widen the palette so the
two columns of longer session labels aren't cramped.
Co-authored-by: Isaac
---------
Signed-off-by: Dimitar Dimitrov <dimitardimitrov9205@gmail.com>
Co-authored-by: Dimitar Dimitrov <dimitardimitrov9205@gmail.com>
Co-authored-by: Daniel Lok <daniel.lok@databricks.com>
The kimi-native forwarder only mirrored `content.part` of type `text`, so
Kimi's reasoning (the `think` block shown in the TUI) never reached the web
conversation — the forwarder's own docstring acknowledged it as "skipped for
v1". The reasoning text lives in `part["think"]`, not `part["text"]`.
Mirror a `think` part as a one-shot transient `external_output_reasoning_delta`
(`started: true`) so the web UI paints a reasoning block — the kimi analogue of
the codex-native fix in #1254, where the project settled this as a required
native-harness capability. `tool.call` / `tool.result` mirroring is left as a
separate follow-up.
Update the existing `_row_to_item` test that asserted think parts are skipped to
assert they now produce a reasoning item.
Closes#1676
Signed-off-by: tomsen-ai <230283659+tomsen-ai@users.noreply.github.com>
Co-authored-by: tomsen-ai <230283659+tomsen-ai@users.noreply.github.com>
The idle reaper snapshots its stale list under the registry lock, then
releases each entry outside it; a single teardown can hold the pass
open for seconds (graceful-SIGTERM wait). A turn that starts on a
later-listed conversation during that window refreshes last_used_at
and marks itself in flight — but release() tore the entry down without
re-checking, SIGTERMing the subprocess mid-turn. Users saw a turn on a
long-idle session die seconds after it started with a harness stream
connection error.
release() now takes only_if_idle_cutoff (passed only by the reaper):
under the registry lock, atomically with the unregister, it skips
entries that were touched after the pass cutoff or have a turn in
flight — they are reclaimed by a later pass once genuinely idle.
Mirrors the pane reaper's busy re-check immediately before teardown.
Signed-off-by: tomsen-ai <230283659+tomsen-ai@users.noreply.github.com>
Co-authored-by: tomsen-ai <230283659+tomsen-ai@users.noreply.github.com>
Scheduled weekday sweep (github-script, modeled on stale.yml +
auto-assign-reviewer) that escalates open PRs/issues an assigned
maintainer has sat on for >5 working days with no reply:
- PRs: re-ping the requested reviewer + add a second reviewer
(lowest-load owner of the touched area(s) in .github/areas.json,
mirrored as an assignee).
- Issues: re-ping the assignee + add a second assignee from the owners
of the area(s) whose comp:* label the issue carries.
- Escalate-once, guarded by BOTH a one-shot `review-sla-escalated` label
and a hidden marker in the comment, so even a failed label write can't
cause daily re-nudging. The second reviewer is added first (best-effort),
so the comment only claims a reviewer that actually attached.
- Cap escalations at 30 per sweep so an existing stale backlog drains
gradually instead of firing all at once, and count each second reviewer
against the in-sweep load so picks rotate across maintainers instead of
concentrating on the current lowest-load one.
Ownership is read from .github/areas.json -- the single source of truth
shared with auto-assign-reviewer.js and issue triage. Runs from the
trusted default branch (reads no PR code). Offline unit test
(review-sla.test.js, 47 assertions, ownership pinned to a fixture) drives
both paths through a mocked client; review-sla-test.yml runs it in CI.
Co-authored-by: Isaac
The Appearance font-size box bound directly to the clamped, committed value
and clamped on every keystroke, so backspacing "13" to "1" snapped straight
to the 12px minimum — you couldn't clear the field or type toward a target.
Decouple the box's displayed text (a free-form draft) from the committed
value: typing shows whatever you enter, applies live only once the draft is a
valid in-range whole number, and clamps + re-syncs on blur/Enter (an empty or
below-min entry settles to the committed size or the minimum). The steppers
still commit and keep the text in sync.
Co-authored-by: Isaac
The design doc had drifted from what actually shipped, and the seam doc
carried a superseded streaming rule. Bring both current:
designs/harness-capabilities-bench-seam.md
- Correct the group-B streaming rule: False → UNSUPPORTED, not PARTIAL.
PARTIAL is a probe observation (coalesced single delta), never declared.
Add the "declare False only from a live 0-delta observation" rule (a static
forwarder grep is insufficient — pi-native disproved it).
docs/harness-bench-design.md
- Add a Status banner up top and a "Current state (shipped)" section: three
transport drivers (sdk-inproc / full-server / native-tui), the six P0
probes, capability-derived matrix, native auto-derivation — and what is not
yet wired.
- Replace the stale "Phasing" (which framed native/full-server as future P1;
both shipped) and refresh "Transport drivers" for the semantic-method driver
design that exists now.
- Note that entry-point plugin discovery now exists (updates the "no discovery
mechanism" constraint), so the bench side of option B is realized.
- Fix the streaming section: only kiro/cursor/qwen are declared non-streaming
(all live-verified 0 deltas), not the earlier blanket seven.
- New "Plugin seamlessness" section: the bench is plugin-ready, but the
server's native-agent seeding is a hardcoded list (the real remaining seam);
the registry-driven-seeding fix closes it.
- New "self-enforcing table in practice" section: kiro/pi/cursor/qwen drift
case studies as worked examples of detect → diagnose → correct-the-source.
- Refresh Open items (drop resolved ones; add the seeding refactor, native-tui
tool/policy, and the per-harness provisioning gaps the bench surfaced).
Docs only; no code change.
* fix(policies): register legacy nessie handler paths in registry
Deployed bundles referencing omnigent.inner.nessie.policies.* were
rejected at session creation because the registry no longer listed
those handler paths after BUILTIN_POLICY_MODULES dropped the shim.
Add the shim back to BUILTIN_POLICY_MODULES with its own POLICY_REGISTRY
that advertises the legacy paths, so old bundles pass validation while
the canonical paths remain under omnigent.policies.builtins.orchestration.
* fix(policies): hide legacy nessie paths from UI with internal_only=True
Explicit /compact on a claude-sdk agent with a pinned bare Anthropic
model (e.g. claude-haiku-4-5-20251001) returned a 500 from the
summarization endpoint. Compaction's Layer-2 summarizer uses the generic
runtime LLM client, whose parse_model_string defaults any prefix-less
model id to OpenAI -- so the Anthropic model id was sent to
api.openai.com, which rejects it, and explicit /compact
(fail_on_summary_error=True) surfaces that as INTERNAL_ERROR (500).
_route_databricks_model_for_compaction already normalized bare
databricks-* ids for this exact reason. Generalize it to
_route_bare_model_for_compaction, which also prefixes bare claude-* with
anthropic/. Already-prefixed ids and bare gpt-* are left untouched.
Co-authored-by: Isaac
* feat(web): add UI font size setting to Appearance
Add a font-size control to Settings → Appearance that scales the whole
interface. The web UI is Tailwind v4 (typography and spacing in rem), so
scaling the root font-size reflows everything uniformly — the same lever
the mobile bump already uses.
The choice is stored as an absolute px value (default 16, range 12–20) and
applied as a --ui-font-scale multiplier on the document root, so it composes
with the mobile @media bump instead of overriding it. Applied before first
paint to avoid a flash, and persisted per-device in localStorage.
The control is a segmented pill ([ − | value | + ]) styled after Cursor's
appearance settings. The theme picker is unchanged.
Co-authored-by: Isaac
* test(e2e): cover UI font size setting
Add a Playwright test mirroring test_theme_toggle.py for the new
Appearance font-size stepper: stepping the value updates the applied
--ui-font-scale on <html> and persists the px choice across a reload,
and the −/+ buttons disable at the 12/20 bounds.
Co-authored-by: Isaac
* feat(models): add Fable 5 and Sonnet 5 to Claude subscription model list
Adds claude-fable-5 and claude-sonnet-5 to the curated subscription
model catalog alongside the existing claude-sonnet-4-6 (kept since
Sonnet 4.6 remains the only option in some regions/workspaces).
* fix(tests): update sys_list_models CI assertion for Fable 5 / Sonnet 5
test_sys_list_models_dispatches_locally_with_static_provider asserted
the old 3-model curated list; missed when claude-fable-5 and
claude-sonnet-5 were added to _SUBSCRIPTION_STATIC_MODELS.
* feat(claude-native): surface Sonnet 4.6 as a distinct /model picker option
Claude Code's /model picker has one fixed alias per family (fable/opus/
sonnet/haiku) plus exactly one extra custom slot
(ANTHROPIC_CUSTOM_MODEL_OPTION). With both claude-sonnet-4-6 and
claude-sonnet-5 in active use, pin the newest Sonnet to the "sonnet"
family alias and the older one to the custom slot so both stay
independently selectable, instead of one silently shadowing the other.
- claude_native.py: a new "sonnet_4_6" key in ucode's claude_models
sets ANTHROPIC_CUSTOM_MODEL_OPTION(_NAME) alongside the existing
per-tier ANTHROPIC_DEFAULT_*_MODEL pins.
- claude_native_forwarder.py: _model_alias_for now special-cases
sonnet-4-6 ids to the "sonnet_4_6" alias before the generic
"sonnet" substring match (a 4.6 id also contains "sonnet").
- claudeNativeModels.ts: adds a "Sonnet 4.6" row; isModelImplicitlySelected
gets the same 4.6-vs-generic-sonnet disambiguation as the backend.
* feat(claude-native): re-enable Fable picker row, label Sonnet rows by version
Fable access is restored, so the withheld row returns. The generic
"Sonnet" row is relabelled "Sonnet 5" so the two Sonnet options read
unambiguously side by side; the id stays the version-agnostic "sonnet"
alias.
* test(e2e-ui): cover the claude-native picker's Fable + dual-Sonnet rows
Asserts the five picker rows and labels, that a bound
databricks-claude-sonnet-4-6 model highlights the Sonnet 4.6 row rather
than the generic Sonnet row, and that picking Sonnet 4.6 PATCHes
model_override and updates the trigger label.
* fix(claude-native): keep Sonnet 4.6 default; add Sonnet 5 as opt-in
#1981 relabelled the primary "sonnet" alias to "Sonnet 5" and put Sonnet
4.6 on Claude Code's one custom /model slot — which presents the newest
Sonnet as the default. Flip it so the default is left alone:
- The "sonnet" alias stays bound to the workspace's existing default
Sonnet (4.6); it's only relabelled "Sonnet 4.6" so it reads clearly
next to the new row. Its model binding is unchanged.
- Sonnet 5 rides the single custom slot (ANTHROPIC_CUSTOM_MODEL_OPTION,
tier "sonnet_5") as an explicit opt-in, not a repointed default.
- Disambiguation (forwarder _model_alias_for + web isModelImplicitlySelected)
routes concrete sonnet-5 ids to the opt-in row; sonnet-4-6 collapses to
the default "sonnet" alias.
- Flip the corresponding unit + e2e assertions.
Builds on #1981 by @dgokeeffe. Fable row + catalog additions unchanged.
Co-authored-by: Isaac
---------
Co-authored-by: Dhruv Gupta <dhruv.gupta@databricks.com>
Resumed claude-native transcripts write a compact_boundary head marker
without a compactMetadata object. Claude Code scans every compact_boundary
on each compaction and destructures compactMetadata, so a missing object
crashes both manual /compact and auto-compaction on resume with:
Error during compaction: Cannot destructure property
'cumulativeDroppedTokens' from null or undefined value
Every subsequent compaction rescans the same transcript and fails the same
way, wedging the session once context fills.
Emit compactMetadata (trigger + postTokens from the item's token_count).
Claude reads every sub-field via ??, so a minimal object is sufficient.
Closes#1955
Signed-off-by: Krzysztof Zarzycki <4157788+kzarzycki@users.noreply.github.com>
Co-authored-by: Krzysztof Zarzycki <4157788+kzarzycki@users.noreply.github.com>
The macOS desktop app raised OS notifications when a session needed
attention (a turn finishing, the agent asking for input, a runner
disconnecting) but never played a sound, unlike the iOS app. Add an
opt-in notification sound driven entirely from the desktop shell, and
stop step-by-step agents from sounding on every milestone.
Desktop shell (web/electron/src/main.js):
- New macOS "Notifications" menu: a "Play Notification Sound" toggle
(OFF by default — the user opts in) and a picker of the system sounds
in /System/Library/Sounds (default Glass); selecting one previews it.
Persisted in settings.json, read live so a change applies to the next
notification.
- The notify handler plays the chosen sound via `afplay` in both the
foreground and background — macOS mutes the frontmost app's own
notification sound, so we mute the toast and play it ourselves, audible
either way and never doubled. A per-session throttle guards a burst.
Notification timing + focus (web/src/hooks/useIdleNotifications.ts):
- Defer a turn-end notification by a 10s settle and cancel it if the
session resumes to running, so a multi-step agent that streams
milestones notifies once at the end instead of once per step. A new
elicitation ("needs response") still fires immediately.
- A session is suppressed while the user is actively viewing it (window
focused AND it's the open conversation). Window focus is read from the
authoritative focus/blur events (and any pointer/key interaction) rather
than a polled document.hasFocus(), which the Electron shell could
misreport.
- Skip notifications for a session whose runner is offline: when nothing
is actively running, the only thing that flips a session terminal is the
server reconciling a dead-runner session (a stale `running` dropping to
`failed`/`idle`), not a real completion — so it must not beep. Stops the
phantom beep after the app sits idle with only stale sessions left.
- Beep a session's turn-end at most once until the user views it: a
session that finishes again while its notification is still outstanding
does not ring again. This also collapses the multiple turn-ends a single
async task produces (launching subagents, then reporting back) into one
beep. The mark clears when the user views the session.
Docs: web/electron/README.md (notification, foreground-cue, and menu
bullets) and the README desktop blurb.
Tests: useIdleNotifications.test.tsx covers the settle, the
focus-from-events fix, the offline-runner filter, and the re-notification
dedup. tests/e2e_ui/sessions/test_idle_notifications.py adds a Playwright
test asserting the turn-end settle deferral end to end — a backgrounded
turn-end stays silent through the settle window, then lands exactly once.
Co-authored-by: Isaac
Signed-off-by: Yuri Chamarelli <yuri.chamarelli@databricks.com>
Co-authored-by: Yuri Chamarelli <yuri.chamarelli@databricks.com>
Submitting the Codex goal dialog rendered a spinner as an extra child
next to the label, widening the button and shifting its neighbours. The
shared Button had no loading state, so every caller inlined its own
spinner beside the text.
Add a `loading` prop to Button that overlays a centered spinner and
hides the label in place (`display: contents` + `invisible`), preserving
the button's width and the flex gap, and forces disabled + aria-busy.
The four Codex goal dialog actions now pass `loading` instead of
inlining a spinner.
Co-authored-by: Isaac
* fix(web): persist brain-harness override across sessions
The per-session brain-harness pick (e.g. claude-sdk vs openai-agents for
bundle agents like Polly) was lost on page refresh because it only lived
in a module-scoped variable. Persist it to localStorage keyed by agent id
so returning users land on the harness they last chose.
* style: fix prettier formatting in NewChatDialog
* fix(web): persist harness under correct agent id on submenu switch
Address Polly AI review feedback:
- Pass the target agent id from the picker when switching agents via
the harness submenu, so the preference is stored under the correct
agent instead of the stale effectiveAgentId from the prior render.
- Fix docstring in harnessPreferences.ts that falsely claimed the
consumer validates stored values against the harness vocabulary.
- Update stale comment on pickedHarness state that still said
"cleared on every agent switch" (now seeds from stored preference).
Show the queued-message Steer button on native sessions too, not just SDK.
The runner delivers a steered message uniformly for every native harness
(POST → buffer → drain → hand to app; each native run_turn returns right
after delivering the input), and the app folds it into the running turn:
deterministically for codex-native (turn/steer RPC) and claude-native (the
TUI folds a pane paste), best-effort for the rest.
Removes the isNativeTerminalSession gate on onSteer (and its now-unused
subscription). steerMessage is harness-agnostic — it just POSTs now.
Verified live: claude-native, codex-native. cursor/pi/hermes/opencode-native
(and the others) get the button too — the mechanism is uniform — but their
mid-response behavior is not yet verified live (tracked as a TODO in
docs/QUEUE_STEER_DESIGN.md; opencode notably has no steer endpoint and queues
as a new prompt).
Co-authored-by: Isaac
The server accepts both native-opencode and opencode-native (harness
aliases), but the web HARNESS_ALIASES map omitted native-opencode, so
nativeCodingAgentForHarness("native-opencode") returned undefined and an
opencode agent forked/switched under that spelling rendered as plain chat
instead of the native terminal wrapper. Add the missing reversed entry.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
* feat(intent-gate): return ASK instead of DENY for off-task tool calls
Switches intent_gate from blocking off-task tool calls outright to
prompting the user for approval, letting them decide whether to proceed.
Also extracts _off_task_reason() to deduplicate the reason string
shared between the cache-hit and fresh-classification paths.
* refactor(intent-gate): rename intent_gate to intent_based_authorization
* refactor(intent-gate): rename display name to Intent Based Authorization
* fix(lint): wrap long log strings in intent_based_authorization
* feat(web): steer a queued message (SDK harnesses)
Adds a per-row steer (send-now) button to the composer's queued strip:
clicking it POSTs that message immediately instead of waiting for the idle
flush. On an SDK harness the server live-injects it into the running turn;
the optimistic bubble promotes on POST. It sends to the agent captured at
enqueue time and can jump ahead of earlier queued messages.
Gated to non-native sessions: native terminals buffer & drain rather than
inject mid-turn, so no steer button is shown there until that path lands
(tracked in docs/QUEUE_STEER_DESIGN.md).
Co-authored-by: Isaac
* fix(web): label steer action and drop the Queued tag
Replace the icon-only steer button with a labeled '↳ Steer' (corner-down-
right arrow + text) and remove the redundant 'Queued' tag — the strip's
position above the composer already signals queued state.
Co-authored-by: Isaac
* test(e2e_ui): steer a queued message sends it mid-turn
Drives the SPA against a spawned server: a first message is acked but
never gets a session.status event, so the session stays busy; a follow-up
queues in the docked strip; clicking Steer POSTs it immediately — which
can only happen via steer, since the session never went idle to trigger
the auto-flush. Asserts the steered message POSTs and leaves the queue.
Co-authored-by: Isaac
* feat(web): edit a queued message from the composer strip
Each queued row gets a pencil button that pulls the message back into the
composer for editing: its text and attachments load into the composer, the
entry is removed from the queue, and the textarea is focused. Any
in-progress draft is preserved (prepended). Re-sending re-queues it (busy)
or sends it (idle).
Stacked on the delete PR.
Co-authored-by: Isaac
* fix(web): edit replaces composer content instead of prepending
Editing a queued message now replaces the composer's text and attachments
with the queued message's, rather than prepending to an in-progress draft
— prepending was surprising when the composer already held content.
Co-authored-by: Isaac
_ensure_default_agents in server/app.py seeded 9 of the 11 native-ui agents
declared in the harness registry (harness_plugins.native_agents) — goose and
hermes were added to the registry but their startup seeders were never wired
in. So `GET /v1/agents` never listed goose-native-ui / hermes-native-ui, and
anything resolving a native agent by that name (the harness bench, and any
head that relies on the built-in row) failed with "not auto-registered".
Add the two missing seeder pairs (_build_*_native_bundle + _ensure_default_*
_agent), mirroring the kiro pattern exactly, and call them from
_ensure_default_agents. goose/hermes have the required _materialize_*_agent_spec
functions already; only the app.py wiring was missing.
Verified: with this change both goose-native and hermes-native get PAST agent
registration in the harness bench (they now reach terminal provisioning, where
each hits a separate downstream issue — hermes a lazy-chat/first-turn gate,
goose a terminal-ensure 500 — tracked separately). test_native_coding_agents
passes; ruff clean.
Note: the per-harness hardcoded seeder list is itself the seam — a native
plugin is invisible until hand-added here. Making _ensure_default_agents
iterate native_agents() from the registry (which already includes plugins) is
the follow-up that would close it.
* feat(web): delete a queued message from the composer strip
Each queued row gets a hover/focus-revealed remove button that drops it
from the client-side queue via a new dequeueMessage(queueId) store action.
Stacked on the client-side message queue foundation.
Co-authored-by: Isaac
* fix(web): make queued-message delete button always visible
The remove button was hover-gated (opacity-0 → group-hover), so the
delete affordance was undiscoverable — users couldn't tell a queued
message could be removed. Show it persistently at reduced opacity;
it brightens on hover/focus.
Co-authored-by: Isaac
* fix(web): use trash icon for queued-message delete
Swap the ✕ for a trash icon so the delete affordance reads as delete,
not dismiss.
Co-authored-by: Isaac
* fix(harness-caps): only declare streaming=False where live-verified (revert #1990 over-reach)
#1990 flipped 7 transcript-mirror natives to streaming=False from a static
"forwarder posts no external_output_text_delta" grep. A live bench run
disproved that for pi-native: it has no delta-posting forwarder yet streams 7
token deltas (its Pi extension emits them by another path), so it drifted
!!✗>✓ (declared UNSUPPORTED, observed SUPPORTED).
The static grep is not a sound basis for asserting a harness does NOT stream.
Revert pi/cursor/goose/qwen/kimi/hermes to streaming=True (their pre-#1990
value, the honest default); keep streaming=False only for kiro-native, which
is live-verified (0 deltas over a full SSE capture). The remaining five are
unverified on this host (own-auth logins the bench can't provision); leaving
them True means the bench will flag a real drift if any turns out not to
stream, rather than asserting an unproven False that drifts the moment the
harness does stream (as pi just showed).
Offline suites: 60 passed / 14 skipped, ruff clean.
* docs(harness-caps): don't claim an unverified emission path for pi-native
The comment asserted pi-native "emits [deltas] by another path" — an inference
that was never traced, the same unverified-assertion habit that caused the
original wrong flip. Soften to the observed fact only: it streams 7 deltas
live, by a path not traced. No behavior change.
* fix(harness-bench): support lazy-chat natives (cursor); mark cursor/qwen non-streaming
Two findings from an all-native bench run:
1. cursor-native could not provision — "native forwarder did not wire up within
90s (no external_session_id)". Root cause: cursor creates its chat id
(external_session_id) lazily, only after the FIRST message lands
(cursor_native_forwarder.py), but the driver hard-gated provisioning on that
id BEFORE posting any turn — a deadlock. claude/codex stamp it at TUI launch,
so the gate worked for them. Add a per-vendor `lazy_chat` flag (NativeVendor)
and skip the pre-turn external_session_id gate for those vendors; the first
probe turn triggers the chat and the forwarder discovers it then. cursor is
the only known lazy-chat native today. Live-verified: cursor-native now
provisions and runs (Basic/Model-override/Interrupt SUPPORTED).
2. With cursor now runnable, its Streaming observed 0 deltas — and qwen-native
likewise (0 deltas) in the same run. Both were declaring streaming=True and
drifting !!✓>✗. Set streaming=False for cursor-native and qwen-native, joining
kiro-native — all three now LIVE-VERIFIED non-streaming (0 deltas observed),
consistent with the "only declare False where observed" rule.
Offline: 60 passed / 14 skipped, ruff clean.
* feat(web): render .ipynb notebooks as read-only previews in the file viewer
Notebooks currently open as raw JSON in Monaco, which is unusable for
reviewing notebook-heavy work. Add a NotebookPreview that renders cells
in order — markdown through the existing react-markdown/GFM pipeline,
code through the shared Shiki CodeBlockContent with execution counts,
and outputs from each cell's mime bundle — with zero new dependencies.
Output handling is safety-first: text/html is never injected into the
DOM (rich outputs like pandas DataFrames fall back to their text/plain
repr with a note), only raster image mimes render as inert data-URIs
(SVG excluded), and stream/error outputs go through the same
ansi-to-react the terminal uses, so colored tracebacks render properly.
Notebooks join markdown/html as previewable: preview is the default
view, with the raw-JSON Monaco source view kept as the escape hatch.
Invalid or truncated notebook JSON shows a parse-error state pointing
at the source view.
* fix(web): make notebook preview robust to real-world .ipynb quirks
The NotebookPreview handled clean, spec-perfect notebooks but broke on
files exported by real kernels:
- Recover from raw C0 control chars (unescaped ANSI in tracebacks/output)
that strict JSON.parse rejects with "Bad control character in string
literal" — retry once after escaping stray control chars inside string
literals.
- Strip all whitespace (not just \n) from base64 image payloads; a
data-URI containing CRLF or spaces is rejected by the browser as a
broken image.
- Validate base64 before building the data-URI (charset + length % 4);
on a corrupt payload show a "could not be decoded" note and fall back
to the text/plain repr instead of an ERR_INVALID_URL broken image.
- Let long unbreakable traceback runs (separator rules, paths) scroll
within the cell (overflow-x-auto + overflow-wrap:anywhere) instead of
widening the whole preview.
Adds regression tests for each case.
Co-authored-by: Isaac
---------
Co-authored-by: Serena Ruan <serena.rxy@gmail.com>
* feat(web): client-side message queue with auto-flush on idle
Follow-ups typed while the agent is busy are now held in a client-side
queue shown in a docked strip above the composer, instead of being POSTed
immediately. The queue head flushes FIFO (one per turn) when the session
goes idle.
The flush is level-triggered — a store action (maybeFlushQueuedHead)
re-evaluated on every status/queue change and on enqueue — so a message
queued just after a turn ends, or after an SSE reconnect that carries no
fresh idle transition, still sends instead of stranding.
In-memory only (no persistence); a hard reload clears the queue.
Per-message actions (delete / edit / steer / reorder) land in follow-ups.
Co-authored-by: Isaac
* fix(web): address queue review — per-conversation flush + edge cases
Fixes from the PR review of the client-side message queue:
- Blocking: flush the first message OF THE BOUND CONVERSATION, not the
global array head. The queue is one flat array across conversations, so
an undrained message from another conversation sat at index 0 and
permanently blocked the bound conversation's messages (the same
never-sends stranding the feature set out to fix). Regression test
covers a foreign head in front of a local entry.
- Pin the agent at enqueue time so a message flushes to the agent it was
composed for even if the binding changed (e.g. a /model switch).
- Hold the flush while the session is unreachable so it doesn't POST into
a void, bypassing the reconnect dialog; drains once reachable again.
- Clear a conversation's queue when it is deleted so entries bound to a
dead session can't linger in memory.
Each fix has a regression test verified to fail without the fix.
Co-authored-by: Isaac
* test(e2e_ui): rewrite cross-session routing test for client-side queue
The client-side message queue changes the routing model the old test
encoded: a follow-up typed while a session is busy is now held in that
session's client-side queue instead of being POSTed on the module-level
send chain. The old repro (hold msg1's POST → msg2 queues on the chain →
switch sessions → chain unblocks → msg2 POSTs to origin) no longer
applies, so the test timed out waiting for a msg2 POST that never fires.
Rewritten to assert the same no-leak guarantee under the new model: a
message queued in B (busy) is held client-side, and switching to idle
session A must never flush it into A. The positive FIFO-flush-on-idle
path is covered by the chatStore unit tests.
Also fixes a real gap the rewrite surfaced: the flush effect now depends
on boundAgentId, so a queue drains correctly when a conversation binds
after navigation (the binding lands after the status settles).
Ran locally against a built web UI: 1 passed.
Co-authored-by: Isaac
The openai-agents harness only handled response.output_text.delta, so a
flagship harness forwarded no reasoning while claude/codex/antigravity all
emit ReasoningChunk. Surface the Responses-API reasoning deltas
(response.reasoning_summary_text.delta and response.reasoning_text.delta)
as ReasoningChunk(event_type="reasoning_text") when non-empty, mirroring
codex. The reasoning_item ghost stays in _NON_OUTPUT_ITEM_TYPES; only the
streaming deltas are mirrored.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
In header/single-user mode the backend already skips admin enforcement,
but the frontend was still waiting on an identity probe that never
resolves an is_admin flag, leaving the page stuck on "Loading..." or
showing the "no permission" message. Mirror the MembersPage pattern:
derive isSingleUser from useServerInfo and bypass the admin gate
entirely when true. Also adds unit tests for the single-user path.
A markdown file containing a blockquote whose only content is a lone
inline image (`> `) or an empty blockquote (`>`) crashed the
markdown editor's panel.
@tiptap/markdown (beta) parses those into a blockquote holding an inline
`image` (or nothing), which violates the blockquote's `block+` content
model. ProseMirror builds the initial document via `nodeFromJSON`, which
does not validate content, so the invalid doc loads silently — then the
first edit transaction that touches the blockquote calls `contentMatchAt`
on it and throws ("Called contentMatchAt on a node with invalid
content"). The viewer's React panel boundary caught the throw and
rendered a crash instead of the file.
Normalize GitHubAlertBlockquote's parsed children to valid `block+`
content (wrap loose inline runs in a paragraph; guarantee at least one
block), so the parsed document is always schema-valid. Round-trip stays
byte-faithful (`> ` re-serialises from the wrapping paragraph).
Co-authored-by: Isaac
The codex-parity sidecar source is frozen (one commit ever) with
rev-pinned deps, yet every CI run recompiled all 73 crates (~3 min)
because the old cache stored the target dir, which restored as a hit
but still forced a full rebuild.
Cache the built binary keyed on sidecar/** + rustc version instead,
and skip `cargo build` on a hit. Warm runs drop from ~4 min to ~15s;
the key self-invalidates when the source, Cargo.lock, or toolchain
changes.
Signed-off-by: Pat Sukprasert <pattara.sk127@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
* fix(members): show friendly message in single-user/header mode instead of auth error
In plain header mode (no accounts, no OIDC), the /auth/users endpoint
does not exist, causing the Members page to show a misleading error.
Add an early return after all hooks when accounts_enabled is false and
login_url is null, rendering a "not available in single-user mode" message.
* fix(members): skip fetch and show not-available message in single-user mode
- Derive isSingleUser from server_version (non-null on a live server,
null on the _OFF probe-failure sentinel) to distinguish real
single-user header mode from a transient /v1/info failure.
- Gate the useEffect on isSingleUser so the identity probe and
/auth/users fetch are skipped entirely in that mode.
- Add a test case asserting the message renders and listUsers is
never called; update mock to expose login_url + server_version
so OIDC and single-user cases are distinguishable.
* fix(harness-bench): classify token-provisioning failures as infra skips
A full-server run over the SDK harnesses exposed a false-drift: codex and pi
fail basic_turn on that transport with a provider/gateway token-provisioning
error ("provider auth command `sh` produced an empty token"; "could not fetch
a gateway token"), which infra_failure_reason did not recognize — so the turn
read as UNSUPPORTED and drifted (!!✓>✗) against the SUPPORTED declaration.
That is an environment/auth gap in the full-server driver's spawn path, not a
capability the harness lacks. Add the token-provisioning phrasings to the infra
markers (with a dedicated skip reason), so such a failure is reported SKIPPED —
matching how a 403 / connectivity error is already handled — instead of a false
capability drift. claude-sdk on full-server is unaffected: it completes the
full matrix (Tool calling + Policy DENY both SUPPORTED and enforced).
Extends the infra-classification test with the codex/pi token-provisioning
messages. Offline 50 passed / 14 skipped, ruff clean.
* docs(harness-bench): document which transport exercises Tool calling / Policy DENY
A default `--profile oss` run shows `·` for Tool calling and Policy DENY, which
reads as "untested" but is really a transport limitation: those two dimensions
only get a real verdict on `full-server` (sdk-inproc harnesses dispatch tools
internally; native-tui isn't wired for them yet). Add a transport-vs-dimension
coverage table, the `--transport full-server` recipe, and the live-verified
result (claude-sdk: Tool calling ✓, Policy DENY ✓ enforced). Record the codex/pi
full-server gateway-auth gap and the native-tui tool/policy gap as open items.
* fix(harness-bench): accurate skip message for a native harness on full-server
Under --transport full-server, a native profile was rejected with "transport
'native-tui' not supported by the 'sdk-inproc' driver" — misleading, since it
is the full-server driver rejecting it and the fix is to use native-tui.
FullServerDriver.unavailable now rejects native profiles itself with an
accurate message ("... is a native-tui harness; ... use --transport
native-tui") and only borrows the SDK driver's CLI gate, not its
sdk-inproc-specific transport check.
Add a test asserting the message names native-tui and never sdk-inproc.
Context: verified on the oss profile that all four SDK harnesses (claude-sdk,
codex, pi, openai-agents) complete the full matrix on full-server with Tool
calling and Policy DENY both SUPPORTED and enforced. The codex "timeout" seen
earlier was a transient cold-start flake under sequential load (codex completes
a basic turn in ~15s solo), not a hang and not an auth failure once the local
Databricks profile was re-authed — no code change needed for it.
Offline 52 passed / 14 skipped, ruff clean.
* fix(server): signal SSE streams to exit on shutdown, reduce graceful timeout
Ctrl-C would hang for up to 30 s because open SSE session streams waited
for their next heartbeat (15 s cadence) before discovering the server was
going away. After the timeout, uvicorn force-cancelled them, producing
spurious "Exception in ASGI application / CancelledError: timeout graceful
shutdown exceeded" tracebacks.
Fix by broadcasting the end-of-stream sentinel to every subscriber queue
in the lifespan shutdown handler (session_stream.shutdown_all()), so SSE
generators return cleanly without waiting for a heartbeat tick. The
graceful-shutdown window is also reduced from 30 s to 5 s: SSE connections
now drain on their own; the remaining window is sized for WebSocket tunnel
teardown, which is fast.
* fix(ci): drop labeled/unlabeled from e2e.yml to prevent automerge label from canceling running E2E
label events share the PR-number concurrency key, so applying automerge
mid-run triggered a new workflow run that immediately canceled the
in-progress suite (cancel-in-progress: true), leaving no E2E result.
e2e-ui.yml and integration.yml already removed these trigger types for the
same reason. Remove labeled/unlabeled from e2e.yml and drop the now-
unnecessary gate `if: github.event.label.name != 'automerge'` condition.
* Revert "fix(ci): drop labeled/unlabeled from e2e.yml to prevent automerge label from canceling running E2E"
This reverts commit f198528373.
* fix(server): move shutdown_all() into Server.shutdown override before graceful wait
The lifespan finally block runs AFTER uvicorn's graceful-shutdown timer
has already expired and force-cancelled in-flight tasks, so calling
shutdown_all() there was a no-op.
Move the call into a uvicorn.Server subclass (_ShutdownSignalingServer)
that overrides shutdown(): the sentinel is broadcast to all SSE subscriber
queues before asyncio.wait_for(_wait_tasks_to_complete(), ...) starts, so
generators exit cleanly within the graceful window instead of being
force-cancelled.
Also clean up session_stream.shutdown_all(): remove the contextlib.suppress
guard (queues are unbounded asyncio.Queue(), so QueueFull is unreachable).
* fix(ci): drop labeled/unlabeled from e2e.yml to stop automerge label canceling running E2E
Applying the automerge label mid-run triggered a new workflow run sharing
the same PR-number concurrency key. With cancel-in-progress: true, that
killed the running suite, leaving no E2E result on the PR.
e2e-ui.yml and integration.yml already removed labeled/unlabeled for the
same reason. Remove them from e2e.yml and drop the now-dead gate condition
`if: github.event.label.name != 'automerge'`.
* fix(server): yield event-loop turn after shutdown_all() before closing transports
Without this pause, generators receive _DONE but cannot run until
super().shutdown() calls connection.shutdown()/transport.close() — at
which point they try to flush "data: [DONE]\n\n" to an already-closing
transport. Writing to a closing transport leaves connections open past
the graceful window, which prevents clear_local_server_record() from
running and leaves the port bound.
One asyncio.sleep(0) turn lets generators consume _DONE, flush their
final chunk, and exit before the transports are torn down.
* fix(server): catch KeyboardInterrupt, use SO_REUSEADDR in port probe
Two issues introduced by the faster shutdown:
1. KeyboardInterrupt now propagates from Server.run() to Click (since we
dropped the uvicorn.run() wrapper that swallowed it), printing
"Aborted!" and exiting non-zero. Add except KeyboardInterrupt: pass
to match uvicorn.run()'s original behaviour.
2. pick_local_port() probed with a plain socket (no SO_REUSEADDR), which
fails on macOS/BSD when recently closed connections are still in
TIME_WAIT with local address 127.0.0.1:6767. The server's listening
socket is already gone, and uvicorn would bind fine (it uses
SO_REUSEADDR), so the probe socket must match.
* revert unrelated e2e.yml change from branch history
* test(cli): update server tests to mock uvicorn.server.Server.run instead of uvicorn.run
The server command now uses uvicorn.Config + _ShutdownSignalingServer(config).run()
rather than uvicorn.run(), so the four tests that monkeypatched uvicorn.run to skip
the blocking server loop were no longer intercepting anything — the real Server.run()
was called, binding to the test port and hanging.
Switch to patching uvicorn.server.Server.run (which _ShutdownSignalingServer inherits)
and capture the same kwarg fields via self.config attributes.
Staged omnigent-site doc PRs all target the per-minor X.Y-docs branch and
carried only the automated-docs label, so maintainers couldn't filter them
by the release they'll ship in. Derive vX.Y.Z from omnigent/version.py in
the existing "Resolve docs branch" step and apply it as a label on both the
create and update paths (backfilling PRs opened before the label existed).
Also add the resolved reviewer as an assignee alongside the review request,
so the PR is filterable by assignee from the site's PR list. The two calls
are independent and best-effort — GitHub rejects non-collaborators with 422,
which stays tolerated as before.
Co-authored-by: Isaac
Label events share the same PR-number concurrency key as code-push events.
With cancel-in-progress: true, applying automerge mid-run fired a new
workflow run that immediately killed the in-progress E2E suite.
Two-part fix:
- Append the label name to the concurrency key for label events (other
events get the suffix '-run'), so each label gets its own isolated slot
and can never preempt a synchronize/push run.
- Add an if: on the gate job to short-circuit for label events that are not
skip-security-scan (e.g. automerge): those runs exit immediately in their
isolated slot rather than spinning up the full suite.
labeled/unlabeled stay in the trigger: they are the fallback recovery path
for skip-security-scan (rerun-security-gate-run.yml calls this out on line 105).
* fix(triage): prioritise load over LLM rank when assigning issues and PR reviewers
LLM rank was the primary sort key, so the first owner listed in areas.json
always won even when their open-issue/review load was far higher than other
eligible owners. Swap to (load, rank, login) so load is the primary signal
and LLM rank only breaks ties within the same load bucket.
* test(triage): update cases 17-19 and stale comment for load-primary sort order
Cases 17-19 previously asserted rank-primary / load-secondary behaviour.
Update them (and their descriptions) to reflect the new load-primary ordering.
Also fix a stale block comment in issue-triage.yml that still said
"rank primary, load secondary".
* ci: re-trigger E2E (previous run canceled by automerge label event)
Injecting a web-UI message while Claude Code is mid-turn grows the footer
with running-state rows (a ○ Explore subagent line, extra spinners) that
push the ❯ input glyph to the 6th non-empty line from the bottom — one
past the readiness gate's 5-line scan window. The gate then times out and
the web UI renders a spurious "did not become ready" runtime-error card,
even though the terminal is healthy and the prompt is on screen.
Widening the window alone would resurrect the scrollback false positive
(an echoed ❯ sits at the same depth). Distinguish them structurally: the
live input box always renders a ──── box rule directly below ❯, which a
scrollback echo never has. Keep the 5-line fast path, and additionally
trust a glyph in a wider 8-line window only when a box rule sits below it.
Co-authored-by: Isaac
Design for a client-side message queue (edit / delete / steer / reorder)
before POST, with auto-flush-on-idle and per-harness steer semantics for
both SDK and native harnesses.
Co-authored-by: Isaac
The policy name "Block Dangerous Shell Commands force-push, rm -rf" read
like an incomplete sentence. Trimmed to "Block Dangerous Shell Commands"
— the description already lists the specific examples.
* refactor(policies): move nessie policies to builtins/orchestration
Move all policy factory functions (blast_radius, spawn_bounds,
headless_subagent_purpose_guard, worktree_guard, read_only_os) and
POLICY_REGISTRY from omnigent.inner.nessie.policies into the proper
omnigent.policies.builtins.orchestration module.
Leave omnigent/inner/nessie/policies.py as a thin re-export shim so
deployed configs that reference handler paths by the old module string
continue to work without any changes. Update BUILTIN_POLICY_MODULES and
all in-repo YAML configs to point at the new canonical path.
* fix(policies): remove redundant F401 noqa on wildcard import in nessie shim
* docs(policies): remove dangling designs/NESSIE.md references
* revert(configs): keep example configs on legacy nessie policy paths
The new orchestration module paths are only safe once all runners have
been updated. The shim at omnigent.inner.nessie.policies handles old
configs indefinitely, so in-repo examples don't need to change.
* fix(policies): add MultiEdit to worktree_guard write-tool set
* fix(harness-bench): streaming=False declares UNSUPPORTED, not PARTIAL
#1990 corrected the transcript-mirror natives to streaming=False, but the
manifest mapped False → PARTIAL while the streaming probe reports a
zero-delta harness as UNSUPPORTED — so kiro-native still drifted (!!~>✗:
declared PARTIAL, observed UNSUPPORTED).
streaming is a binary capability: True → SUPPORTED, False → UNSUPPORTED.
PARTIAL is a probe *observation* (the ambiguous coalesced-single-delta retry
case against a SUPPORTED declaration), never a declared value. Map False →
UNSUPPORTED so a non-streaming harness's declaration matches what the probe
observes. Live-verified: kiro-native now renders a clean ✗ with no drift
(exit 0).
- Add a regression test locking the binary mapping (True→SUPPORTED,
False→UNSUPPORTED, never PARTIAL declared).
- Document in the design doc: how to run/read the bench (a subset suffices;
own-auth natives skip cleanly; read DRIFT + unexpected ✗/· only), and that
streaming is a binary declared capability.
Offline 51 passed / 14 skipped, ruff clean.
* docs(harness-bench): tighten streaming-verdict comments
The binary-streaming rule was explained at length in both the manifest and the
test. Keep the canonical 4-line "why" in the manifest; reduce the test comment
to a one-line pointer. No behavior change.
The harness capability bench flagged a real drift on kiro-native: it declares
streaming=True but emits zero token-level deltas. Root cause is architectural,
not a bench bug: kiro (and the same-shaped goose/qwen/hermes/cursor/kimi/pi
natives) delivers output by mirroring each COMPLETE assistant message
(external_conversation_item) from the vendor's transcript, never posting
incremental external_output_text_delta. So the web UI sees the reply
complete-only, not streamed.
Set streaming=False for those 7 to match reality. kiro-native is live-verified
(0 deltas across a full SSE capture, whole reply arrives as one
response.output_item.done); the other 6 share the identical forwarder shape
(grep-confirmed: 0 external_output_text_delta posts in each). Left as True:
claude-native, codex-native, antigravity-native (forwarders DO post deltas),
and opencode-native (native-server, not benched here).
This is the capability model catching up to the forwarders; no forwarder or
executor behavior changes. tests/test_harness_capabilities.py only asserts the
4 SDK harnesses stream, so it is unaffected.
* test(harness-bench): auto-derive native-tui harnesses from capabilities
Any harness the capability model marks NATIVE_TUI is now probeable by name
with no bench edit -- including a community-plugin native, since
harness_capabilities() already discovers plugins via entry points. This
replaces the hardcoded 2-entry _VENDORS table and wires the 9 remaining
in-repo native harnesses for free.
- native_vendor(harness) derives the driver's per-vendor facts (UI agent name
<harness>-ui, terminal name, own_auth from AuthModel) from the capability
model instead of a static dict. native-server harnesses (opencode-native)
return None -- different transport.
- The manifest registers every NATIVE_TUI harness. Registration is separate
from runnability: OMNIGENT_CREDENTIAL natives (claude, codex) route through
the run's Databricks profile and run unattended; own-auth / session-scoped
natives are registered (visible, honest declared matrix) but skip-gate when
their vendor login is absent.
- Provisioning is now uniform: the native-terminal ensure + external_session_id
readiness gate is the shared protocol every native uses, so claude and codex
no longer need a per-vendor flag. Verified claude-native + codex-native still
pass live with no regression through the unified path.
- cli_binary is not always "<harness> minus -native" (cursor -> cursor-agent,
kiro -> kiro-cli); added an explicit override map for those.
- A provisioning failure is now caught and reported as a per-harness skip
rather than aborting the whole run, so a multi-harness run survives one
unrunnable harness (verified: claude-native + cursor-native -> claude green,
cursor clean-skipped, matrix still rendered).
Offline 49 passed / 14 skipped, ruff clean.
* test(harness-bench): tear down on provisioning failure; address review
Fixes the blocking issue from the Polly review: the provisioning-failure skip
branch returned without tearing down the server + daemon that __aenter__ had
already spawned, so every skipped own-auth native leaked an orphaned server +
daemon process — undermining the multi-harness resilience this path is for.
- Construct the driver context manager outside the try, and in the
__aenter__-failure branch call __aexit__ (suppressing any teardown error) so
a half-provisioned driver is cleaned up. _teardown already null-checks
_client/_proc/_daemon, so it is safe after a partial provision.
- Log the traceback in that branch (warning): it also catches genuine driver
bugs (e.g. an AssertionError), which must not vanish silently behind a
green-looking skip.
- Note the agent_name/terminal_name convention in native_vendor(): it holds
for every in-repo native; a plugin whose names diverge would need an
override map like the manifest's _NATIVE_CLI_BINARY.
- Add a regression test: a driver raising in __aenter__ yields a skip AND is
torn down.
Offline 50 passed / 14 skipped, ruff clean.
* test(harness-bench): drop double-import in provisioning-failure test
Addresses the review nit: the new test imported tests.harness_bench.bench both
via the top-level `from ... import run_harness` and an inner `import ... as
bench_mod`. Patch resolve_driver_class via monkeypatch's string target instead,
and drop the redundant inner Verdict import (already imported at top). No
behavior change.
## Related issue
N/A
## Summary
- The control-mode web-terminal bridge sent one WebSocket frame per tmux
`%output` line. tmux firehoses output as many small per-line writes
(~1 KB each, ~8 MB/s, no throttling), so a heavy burst became thousands
of tiny frames — and when the browser send lags the producer (any real
network), that backlog was flushed one tiny frame at a time.
- Reuse the PTY bridge's queue-driven coalescing forwarder
(`_forward_pty_to_ws`) in `control_bridge.py`: split the old
read-and-send loop into a reader that parses the control stream and
queues decoded `%output` payloads, and the forwarder that drains
everything already queued into one bounded `send_bytes`. A backlog now
collapses into a few large frames; a lone keystroke echo (nothing else
queued) still flushes immediately.
- The reader uses raw `stdout.read()` + its own line buffer instead of
`readline()`, so one wakeup can pull many `%output` lines (giving the
forwarder something to merge) and an oversized line can't raise
`LimitOverrunError`. Reader-finished remains the "session ended" signal
the detach-vs-gone close-code logic keys on.
- Drain-on-exit: because the reader and forwarder are now separate tasks
and shutdown keys on the reader, a burst-then-exit program (dump then
`%exit`) could otherwise have its still-queued tail cancelled mid-drain.
On the reader-ended path the forwarder is awaited (bounded by
`_FORWARD_DRAIN_TIMEOUT_S`) so the sentinel-terminated backlog fully
flushes before teardown — the inline-send loop's ordering guarantee,
restored.
- Reuse `_coalesce_limit_after_input` so the frame right after a keystroke
stays small (xterm's synchronous echo paint path). No browser-facing
wire-protocol change; seed, cursor-restore, scrollback, resize, hex
input, and detach paths are untouched.
## Test Plan
- Before/after with an identical harness (real tmux, 3 MB burst, 1 ms/frame
send): frames dropped from 2,055 (avg 1,459 B) to 162 (avg 18,518 B) for
byte-identical output — ~12.7x fewer WS frames.
- Interactive echo unaffected: a lone keystroke still echoes as 1 frame,
1 byte, ~0.5 ms (coalescing only merges an existing backlog).
- `test_control_bridge_coalesces_burst_when_send_lags`: 500 KB burst behind
a slow send, asserts full delivery AND <100 frames (proves merging).
- `test_control_bridge_burst_then_exit_delivers_full_tail`: 2 MB burst then
immediate exit behind a 5 ms/frame send — asserts the full payload
arrives. Verified this fails without the drain (1.25 MB of 2 MB delivered)
and passes with it (2 MB) — a true regression guard.
- `pytest tests/terminals/test_control_bridge.py` — all 11 pass (seed /
staircase / cursor-restore / scrollback / alt-screen / detach preserved).
Pre-commit clean.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [x] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Coalescing and the drain-on-exit fix are both covered by real-tmux
integration tests that drive bursts behind a slow fake WebSocket and assert
merged frame count / full-tail delivery; the drain test was confirmed to
fail without the fix and pass with it. Manual verification: ran the
before/after measurement harness confirming the ~12.7x frame reduction and
that a lone keystroke echo still flushes as a single immediate 1-byte frame
(no interactive-latency regression). No browser E2E — the WebSocket
TestClient can't drive the streaming receive loop — so the browser-layer
effect stays manual, but the server-side frame-count and no-tail-drop
behavior are pinned by tests.
## Related issue
N/A
## Summary
- Add `omnigent/terminals/control_bridge.py`: a `tmux -C` control-mode
bridge that streams per-pane `%output` into the browser xterm, so the
browser owns scrollback and text selection natively (fixing the
scroll/copy pains of the PTY `tmux attach` transport, which let tmux
own the viewport and capture the mouse).
- Select the transport per attach via `resolve_terminal_transport()`
(`omnigent/inner/terminal.py`): per-attach `?transport=` query ›
per-terminal `TerminalEnvSpec.terminal_transport` › global default.
Control mode is the default; set `terminal.transport: pty` in
`~/.omnigent/config.yaml` to opt the whole install back to the legacy
PTY path. The config is read at attach time (honoring
`OMNIGENT_CONFIG_HOME`), so an edit takes effect on the next attach
without a restart. The PTY bridge is untouched, so the modes run side
by side and revert is a config edit.
- Wire both attach call sites (server fallback `terminal_attach.py`,
runner `runner/app.py`) to pick the bridge; forward `?transport=` over
the runner WS tunnel; stamp `terminal.transport` on telemetry.
- Surface the resolved transport per terminal in resource metadata
(`session_resources.py`) so the web UI (`TerminalView`/`useTerminals`)
switches mouse/selection behavior and drops the hint bar in control
mode, and dedupes redundant resize frames (`TerminalSession`).
- Seed-on-attach fidelity: a control client only receives `%output`
after it attaches, so the bridge seeds the current screen via
`capture-pane -e`. Normalize bare-LF row separators to CRLF (fixes the
staircase), strip the trailing separator (fixes the full-height
off-by-one scroll), restore cursor position + visibility, and capture
`-S -` scrollback only on the primary screen (alt-screen `-S -` would
leak stale primary history).
## Test Plan
- `pytest tests/terminals/test_control_bridge.py` — 8 tests against a
real private tmux server: octal un-escape, `send-keys -H` chunking,
seed streaming + detach close code, CRLF/no-staircase, cursor restore,
full-height no-scroll (verified via a pyte VT emulator), primary
scrollback recovery, and alt-screen no-history-leak.
- `pytest tests/inner/test_terminal.py::test_resolve_terminal_transport_precedence`
— transport selection precedence, reading `terminal.transport` from a
scratch `~/.omnigent/config.yaml` via `OMNIGENT_CONFIG_HOME`; plus the
runner route-dispatch test for `?transport=` bridge routing.
- `vitest` for `TerminalView` / `TerminalSession` / `useTerminals` —
transport plumbing, native-selection + hint-bar gating, resize dedupe.
- Manual: drove the polly claude-sdk REPL and a claude/codex full-screen
session through the web UI, toggling transcript/chat and back, to
confirm no staircase, no off-by-one line, correct cursor, and
recovered scrollback. Reproduced each seed bug against real tmux
before fixing.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [x] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
The control bridge and transport selection are covered by real-tmux
integration tests (seed rendering asserted through a pyte VT emulator)
and frontend unit tests; the config-file default resolution is covered
by writing a scratch config.yaml under OMNIGENT_CONFIG_HOME. Manual
verification covered the parts no automated test exercises: a live
browser reconnect against the polly REPL (primary screen) and
claude/codex (alternate screen), confirming the seed renders without
staircase, extra line, cursor drift, or leaked history. No full browser
E2E was added; the WebSocket TestClient can't drive the streaming
receive loop, so that path stays manual for now.
`chat_stream_to_response_events` only extracted reasoning from typed blocks
nested inside `delta.content` (the Kimi shape). xAI Grok and DeepSeek instead
emit chain-of-thought as a sibling `delta.reasoning_content` string while
`delta.content` is null during the thinking phase, so Grok reasoning was
silently dropped and never reached the REPL/UI.
Surface a non-empty `delta.reasoning_content` as
`ResponseReasoningStartedEvent` + `ResponseReasoningTextDeltaEvent`, reusing the
existing `reasoning_started` sentinel so it interleaves correctly with answer
text and stays out of the final message output.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
* test(harness-bench): wire codex-native native-tui observation
codex-native turns now surface on the bench's shared observe path (basic ✓,
streaming ✓, model override ✓, interrupt ✓ — live-verified on oss, no drift),
so it ships as an official native-tui profile alongside claude-native.
#1880 deferred codex-native on the belief its app-server RPC delivery was
unobservable on the session stream. That was wrong: codex has a runner-side
forwarder that translates app-server RPC into the SAME
response.output_text.delta + response.output_item.done + persisted assistant
item claude-native produces. The gap was provisioning, not observability. A
codex turn needs three things before its forwarder wires up:
1. Provider auth via omnigent config, NOT DATABRICKS_CONFIG_PROFILE.
resolve_native_codex_launch reads the provider from ~/.omnigent/config.yaml
(auth block) / omnigent setup, honoring $OMNIGENT_CONFIG_HOME. Without it
codex falls back to ambient detection, hits the vendor login screen, and
never starts an app-server thread. The driver writes a bench-owned config
home routing codex through the same Databricks profile.
2. Explicit runner launch + bind before the terminal ensure (an unbound
session 503s runner_unavailable).
3. Native terminal ensure + a wait for the forwarder to stamp the session's
external_session_id (the codex thread id) before the first turn.
Gated behind a per-vendor needs_terminal_ensure flag on NativeVendor, so
claude-native is unchanged (its forwarder auto-starts on bind). Once the
forwarder is live, turns drive on the existing shared path unchanged.
Offline 25 passed / 6 skipped, ruff clean. Live: codex-native and
claude-native both pass all wired dimensions with no drift.
* test(harness-bench): trim redundant codex-native comments
The codex-native delivery model was explained in full in four places (module
docstring, NativeVendor.needs_terminal_ensure doc, the _VENDORS comment, and
the manifest comment) plus long inline blocks. Keep the one canonical
explanation (module docstring + the param doc) and cut the duplicates to a
single load-bearing line each. No behavior change.
* feat(tools): add Keenable backend to web_search
Adds a Keenable search backend to the web_search built-in tool, alongside
the existing google / perplexity / nimble / tavily backends, giving
non-OpenAI models another grounded-search option.
Unlike the other backends, Keenable is keyless by default: with no api_key
it calls the public endpoint (/v1/search/public), so it works out of the
box. Supplying an api_key switches to the authenticated endpoint
(/v1/search, X-API-Key header) and lifts rate limits.
- New web_search_keenable.py, mirroring the Tavily/Nimble backends:
optional api_key, max_results clamped 1-20, X-Keenable-Title: Omnigent
attribution header, error-as-string contract, OMNIGENT_KEENABLE_BASE_URL
test override.
- web_search.py gains a _run_keenable dispatch branch (no required key)
plus updated help text and module/_search docstrings.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(web_search): drive backends from a single registry
The selectable search_provider engines were hardcoded in ~5 places
(module + class + _search docstrings, the if/elif dispatch, and two error
strings), so adding a backend meant editing prose in each spot and the
lists had already drifted. Add a `_BACKENDS` registry as the single source
of truth: the dispatch and the error hint both derive from it, and adding
an engine is now a `_run_*` plus one row.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Pat Sukprasert <pattara.sk127@gmail.com>
The "Working folder" header doubled as a collapse toggle (chevron +
aria-expanded), but the file list is the panel's only content — collapsing
it leaves an empty panel with nothing to reveal. Make the header a static
label everywhere; the content is always visible. The drawer keeps its X
close button.
Drops the now-unused `collapsed` preference field and the collapse-specific
unit and e2e coverage, replacing the e2e header test with a guard that the
header is a static label (not a toggle button).
Co-authored-by: Isaac
A reconnecting runner opens a fresh tunnel that supersedes the old one
(newest-wins in TunnelRegistry.register). The new tunnel's
_on_runner_connect recovers the session (clears a stale
runner_disconnected failure to idle), but the superseded tunnel's
teardown then fires _on_runner_disconnect, which re-marks every session
bound to that runner_id failed via a by-runner store lookup - clobbering
the recovery even though the runner is live again.
Guard _on_runner_disconnect: if a live tunnel is still registered for the
runner_id, a newer connection superseded the closing one, so the runner
is not offline - skip the offline-marking. Mirrors the registry's own
generation-guarded deregister(runner_id, session). Genuine offline
runners are unaffected: the WS handler deregisters before invoking the
hook, so no live tunnel is present for a truly-gone runner.
This surfaced as a flaky failure in
test_on_runner_connect_clears_disconnect_failure_on_idle_reconnect
(assert 'failed' != 'failed') under CI load; the recovery path landed in
PR #1593.
main always carries the next unreleased version (X.Y.Z.dev0), so the docs
generated from merged PRs describe a release that isn't out yet. Targeting
omnigent-site `main` deployed those in-progress docs live on merge.
Stage them on a per-minor branch `X.Y-docs` (derived from omnigent/version.py)
instead: doc-sync and sync-openapi-to-site create it off site `main` on the
first doc PR of the cycle and base their PRs on it, so merges accumulate there
without going live. At release, publish-changelog opens a `X.Y-docs -> main` PR
that a human merges to publish the whole batch at once.
The branch name tracks main's version automatically, so there's nothing to
create or retarget by hand across release cycles.
Co-authored-by: Isaac
* test(harness-bench): native-tui transport driver (claude-native skeleton)
Adds NativeTuiDriver, registered as the 'native-tui' transport. A native-tui
turn rides the same HTTP surface as full-server (POST events, GET stream SSE
deltas, item polling), so the driver reuses that machinery (extracted
spawn_omnigent_server as a shared module helper). Three things diverge and
are handled here:
- Provisioning: spawn a host daemon under the real $HOME (vendor login is
inherited, not relocatable), wait for the host online, and create the
session as {agent_id, host_id, workspace} against the auto-registered
<harness>-native-ui agent — not an agent tarball.
- Interrupt: native cancellation surfaces as a session.interrupted SSE
event (no 'interrupted' user-message marker), so run_interrupt_turn keys
off that.
- Per-vendor facts live in NativeVendor records; claude-native is the wired
skeleton, so adding a harness is a config entry (+ a host login), not a
new driver.
Scope / honesty: this is a structurally-complete, offline-tested walking
skeleton. It was NOT live-verified in the authoring environment (native-tui
needs an interactive vendor login the sandbox lacks: 'claude' is aliased to
isaac). The tool/policy dimension is intentionally left unmeasured (returns
a capability-neutral skip) pending native permission-decision observation.
The gated live test runs it where a login exists.
Offline 19 passed / 4 skipped, ruff + pre-commit clean.
* test(harness-bench): add claude-native + codex-native profiles to the suite
The native-tui driver (#1879) added the transport but no selectable profile,
so --harness claude-native KeyError'd before reaching the driver. Ship the
two OMNIGENT_CREDENTIAL native harnesses as official profiles so they are
selectable and appear in the declared matrix:
- _native_profile builds a native-tui BenchProfile with columns + verdicts
derived from the capability model (reusing the #1865 helpers); transport
is native-tui and the driver skip-gates on the vendor CLI binary.
- Only claude-native + codex-native (OMNIGENT_CREDENTIAL) ship as official —
the bench can mint their gateway credential. OWN_AUTH natives stay opt-in.
- model_override now also derives from is_native_harness(): native harnesses
take the model as a launch --model argv (per model_override.py), so the
declaration is truthful rather than absent.
- codex-native added to the driver's _VENDORS (both hit only the shared
session HTTP surface; RPC-vs-tmux delivery is runner-side).
Offline 25 passed / 6 skipped; the declared matrix now renders both native
rows. Still not live-verified (needs a host with the vendor CLI logged in).
* test(harness-bench): fix native-tui streaming subscribe-after-post race
Live smoke of claude-native surfaced a false streaming DRIFT (declared
deltas, observed none). Root cause: _drive_turn subscribed to the session
SSE stream AFTER posting the message, so deltas that fired before the
subscription opened were missed (the stream is not replayed). Basic turn
worked because it reads via item-polling, not deltas.
Fix mirrors the full-server streaming probe: open the SSE subscription on a
background thread and wait until it is connected (ready event) BEFORE
posting the turn, so no deltas are lost. This is the bench catching a real
driver bug via its own drift signal — exactly the intent.
* test(harness-bench): drive native turns from the SSE stream, not stale item polling
The real root cause behind the false streaming DRIFT (a live SSE dump
confirmed 5 response.output_text.delta events DO arrive for claude-native).
The bug was not the event flow: _drive_turn ended the delta read as soon as
_poll_assistant_text found *an* assistant item — but the driver reuses one
session across probes, so it matched a PRIOR turn's stale item and stopped
counting before the current turn's deltas arrived. My earlier
subscribe-before-post fix didn't help because the stale-item read still
ended the turn early.
Fix: drive each turn entirely from the stream. Subscribe first, post, then
read to this turn's response.completed — counting deltas and accumulating
delta text inline, so delta count, text, and terminal state are all scoped
to THIS turn. Interrupt turn gets the same subscribe-first treatment (so it
sees the first delta to trigger on and the terminal session.interrupted).
Event names confirmed live. Removes the stale item-poll helper.
Offline 25 passed / 6 skipped, ruff + pre-commit clean. Awaiting a re-run
to confirm streaming ✓ and interrupt live.
* test(harness-bench): native turn = item-poll text + stream delta count, baseline-scoped
Combine the two observation sources by what each reliably gives, instead of
forcing one to do both (the prior two attempts each broke the other half):
- text from item polling (proven to work for basic turn), but scoped to a
NEW assistant item: record the assistant-item count BEFORE posting and
wait for one beyond that baseline, so the reused session can't return a
prior turn's stale reply.
- delta count from the SSE stream (subscribe-first background thread; the
live dump confirmed 5 response.output_text.delta arrive). A short reply
can complete with zero deltas as a single output_item.done, so
delta-only text was empty for basic turn (the regression the last run
showed) — item text is authoritative.
Offline 25 passed / 6 skipped, ruff + pre-commit clean. Awaiting re-run.
* test(harness-bench): fix native-tui streaming/interrupt (completed fires early)
A per-event SSE diagnostic against real claude-native showed the actual
cause of the streaming DRIFT and skipped interrupt: on native-tui,
response.completed fires ~7s BEFORE the assistant's text deltas -- it marks
the turn being accepted, not the reply finishing. The real end-of-output is
response.output_item.done, right after the last delta.
The reader treated response.completed as terminal, so it exited at t~0.4s
with zero deltas counted (Streaming reported UNSUPPORTED, a false DRIFT), and
the interrupt reader returned before any text streamed (interrupt never
exercised, SKIPPED).
Fixes:
- Reader stops on response.output_item.done, not response.completed
(_READER_TERMINAL drops the early completed event).
- Interrupt timing moves to the main thread: wait for response.in_progress,
hold briefly, then interrupt -- native deltas burst at the very end of the
turn, so firing on the first delta lands too late to interrupt mid-turn.
Live (oss profile, real claude): Basic ✓, Streaming ✓ (9 deltas), Model
override ✓, Interrupt ✓ (cancelled). No drift. Offline 25 passed / 6 skipped,
ruff clean.
* test(harness-bench): ship claude-native only; defer codex-native to follow-up
A live smoke of codex-native showed the shared native-tui observe path
cannot see its turns: codex-native delivers output via app-server RPC, not
tmux paste, so a turn runs (in_progress -> completed) without emitting text
deltas or persisting an assistant item on the session stream the driver
reads. claude-native (tmux-paste) surfaces normally and is live-verified.
Drop codex-native from the shipped OFFICIAL_PROFILES so nothing ships that
the driver cannot drive. Its vendor entry stays in the driver's _VENDORS so
`--harness codex-native --transport native-tui` still resolves and
skip-gates cleanly; wiring RPC-delivery observation earns it an official
profile in a follow-up. Corrected the _VENDORS comment (it wrongly claimed
both vendors drive identically over the shared surface) and the module
docstring scope/verification note.
Offline 22 passed / 5 skipped (the 3 auto-parametrized codex-native cases
drop with the profile), ruff clean.
* feat(web): make the numeric pinned-session jump work in the browser (#7)
usePinnedSessionHotkeys was Electron-only: a browser tab reserves plain
Cmd/Ctrl+digit for native tab-switching, so the hook bailed out outside the
desktop shell. Add a browser-safe chord — Cmd/Ctrl+Alt+digit — that frees a
binding the page can own; the Electron shell keeps the plain Cmd/Ctrl+digit it
can safely claim. With Alt held, macOS rewrites e.key to a composed glyph
(⌥1 → "¡"), so the browser path matches on e.code (physical key) while the
native path keeps matching e.key.
The Keyboard Shortcuts dialog now lists "Jump to pinned session (1–10)" in both
shells, with the matching chord glyphs (Cmd/Ctrl+digit desktop, +Alt in browser).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: jkfnc <56741357+jkfnc@users.noreply.github.com>
* fix(hotkeys): guard getModifierState so a keydown can't throw (#7)
Not every environment (or synthetic event) implements
KeyboardEvent.getModifierState; calling it unguarded would throw on every
keydown and break the sidebar-toggle hotkeys entirely. Guard that it's a
function before the AltGraph check.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: jkfnc <56741357+jkfnc@users.noreply.github.com>
* test(e2e-ui): sidebar keyboard chords — pinned jump + toggle (#7)
Covers both hook changes with real browser keydowns: Ctrl+Alt+1 navigates to
the first pinned session (pin seeded in localStorage; waits for the rendered
Pinned section so the hook's input list is populated), and Ctrl+Alt+[
collapses/expands the left sidebar (asserted via the search input's rendered
width — the rail collapses to icons rather than unmounting). Satisfies the
e2e-ui coverage gate for the web/ changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: jkfnc <56741357+jkfnc@users.noreply.github.com>
* fix(hotkeys): guard AltGraph in the pinned-jump browser chord (#7)
Review finding (Polly, blocking): AltGr reports as Ctrl+Alt on Windows/Linux
intl layouts, so typing AltGr+digit (a composed character) matched the
browser path's Ctrl/Cmd+Alt+code chord and yanked the user to a pinned
session, preventDefault-ing the composition. Bail when
getModifierState("AltGraph") is true - the identical guard (and the same
typeof feature-detect) the sibling useSidebarToggleHotkeys already has.
Adds the companion negative test: an AltGr chord neither navigates nor
prevents default, mirroring the sibling hook's AltGraph test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: jkfnc <56741357+jkfnc@users.noreply.github.com>
---------
Signed-off-by: jkfnc <56741357+jkfnc@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(web-ui): rendered Markdown preview pane for .md files (#970)
Markdown files now open in a read-only rendered Preview by default in the
file viewer — the same affordance HTML already has — with the rich-text
Editor and raw Source one toolbar tap away. Works on the desktop and the
responsive/mobile layout (same FileViewer). Previously .md opened straight
into the editable rich-text editor; the read-only MarkdownPreview existed
and was tested at the CodeViewer level but was unreachable through the UI.
- FileViewer gives markdown a Preview / Edit / Source segmented toolbar;
previewableViewMode defaults to "preview".
- The preview renders headings, lists, tables, fenced code, blockquotes and
task lists via remark-gfm; remark-emoji renders GitHub-style :shortcode:
emoji as glyphs so docs read the same here as on GitHub.
- Schema-versioned preferences (v2) so the new default reaches returning
users whose old build auto-persisted "editor" (diff prefs preserved; a
deliberate future editor choice is still honored).
- HTML's preview<->source toggle now writes the absolute target keyed off
the resolved view, so a single click always flips the surface even when
the shared preference is "editor".
- ?comment= deep links to a .md file open in the editor (the surface that
highlights the comment anchor), since the read-only preview can't.
* fix(web-ui): render raw HTML in markdown preview; collapse view modes into a dropdown
Address review feedback on the markdown preview pane:
- Raw HTML embedded in .md files (<details>, <sub>/<sup>, <kbd>, <br>,
<div align>, inline <img>) rendered as escaped literal text because
react-markdown drops raw HTML by default. Add rehype-raw to parse it and
rehype-sanitize to strip anything unsafe (<script>, event handlers,
javascript: URLs), so the preview matches GitHub while staying safe to
render inline (markdown content is untrusted).
- Collapse the three markdown view-mode buttons (Preview / Edit / Source)
into a single "View mode" dropdown so the toolbar isn't overcrowded:
a picker button inline, a submenu when the toolbar overflows.
- Explain why the deep-link editor bias is a separate override rather than a
seeded previewableViewMode (global persistence + reactivity).
- Update the five markdown-editor e2e tests for the preview-by-default flow
and the new view-mode dropdown, via a shared switch_markdown_view_mode
conftest helper.
Co-authored-by: Isaac
* fix(web-ui): GitHub-style alerts and honored <img> dimensions in markdown preview
Bring the rendered markdown preview closer to GitHub's own rendering:
- GitHub alerts: `> [!NOTE]` / `[!TIP]` / `[!IMPORTANT]` / `[!WARNING]` /
`[!CAUTION]` rendered as plain blockquotes with the literal marker text,
because remark-gfm doesn't implement them. Add rehype-github-alerts so they
become GitHub's typed callouts, and style them GitHub-exact (per-type border
+ octicon + hue, light and dark) reusing the same icons/colors as the
rich-text editor. The plugin's inline <svg> octicon is dropped in sanitize
and redrawn via a CSS mask, keeping the sanitized surface a fixed set of
markdown-alert* classes rather than arbitrary SVG.
- <img width>/<img height>: the attributes survived sanitization but Tailwind
Preflight's `img { height: auto }` overrode them (presentational hints lose
to author CSS), so explicitly-sized images rendered square. A custom img
renderer forwards integer width/height to an inline style, which wins the
cascade — matching GitHub, and how the editor already handles it.
Sanitize stays strict: <script>, event handlers, javascript: URLs, and
non-alert classes are still stripped (markdown content is untrusted).
Co-authored-by: Isaac
* fix(web-ui): honor <img> width/height in the markdown editor too
The rich-text editor had the same image-sizing gap the preview did: its
image node view set width/height as HTML attributes, which Tailwind
Preflight's `img { height: auto }` overrides, so an explicitly-sized image
(e.g. width="200" height="100") rendered square. Forward integer pixel
dimensions to the inline style instead — which wins the cascade — in both
the node view's create and update paths, and clear the style when a
dimension attr is removed. Markdown serialisation is untouched (it reads
node.attrs, not the DOM), so sized images still round-trip to HTML.
Co-authored-by: Isaac
* feat(web-ui): keep markdown opening in the editor by default
Restore the rich-text editor as the default view mode for markdown files.
The rendered preview stays a first-class mode — reachable (with raw source)
from the "View mode" dropdown — but markdown opens in the editor as it did
before, matching how people actually work in these files.
- Revert the previewableViewMode default editor→preview, dropping the
schema-version migration that existed only to force returning users onto
preview. HTML still defaults to its rendered preview.
- The ?comment= deep-link editor bias now only fires when the user's sticky
preference is Preview (otherwise the editor default already lands on a
highlightable surface); its tests seed Preview so they exercise the bias.
- e2e: markdown opens in the editor again, so the initial switch-to-Edit
steps are removed; the mid-test Source/Edit toggles still go through the
dropdown helper (the standalone toolbar buttons are gone).
Co-authored-by: Isaac
* fix(web-ui): always open comment deep links in the markdown editor
A ?comment= deep link now forces the rich-text editor regardless of the
user's sticky view-mode preference, not only when that preference is
Preview. Following a comment link should always land on a surface that
shows the comment's anchor highlight; the read-only preview can't render
it, so a Preview-preferring user would otherwise arrive where the comment
they came to see isn't visible. Drop the `previewableViewMode === "preview"`
guard on the deep-link bias and cover the preview + source preferences.
Co-authored-by: Isaac
* test(e2e-ui): scope comment Edit clicks to exclude the view-mode dropdown
comment_actions.md now opens in the editor by default, so the markdown
toolbar renders a "View mode: Edit" dropdown trigger. get_by_role with a
substring name match then matched both that trigger and the comment card's
"Edit" button, failing under Playwright strict mode. Add exact=True to the
two comment Edit clicks (mirroring the existing exact=True on "Save") so
they target only the comment card affordance.
Co-authored-by: Isaac
---------
Co-authored-by: Daniel Lok <daniel.lok@databricks.com>
* fix(subagents): show "Disconnected" pill for runner disconnect, not red "Failed"
A session/sub-agent whose runner merely DISCONNECTED (tunnel drop) or
EXITED was shown with a red "Failed" badge in the Subagents panel,
indistinguishable from a genuine task failure.
Option B: introduce an explicit, end-to-end "Disconnected" state that is
visually and semantically separate from "Failed".
Backend (omnigent/server/routes/sessions.py):
- On relay tunnel drop, persist the ``runner_disconnected`` cause as
durable ``last_task_error`` labels (alongside the existing clean SSE
``session.status: failed`` terminal event from #1114). Previously the
relay-fed cache only carried a generic ``failed`` and the cause was
dropped from child-session summaries. The snapshot builder already
carries ``runner_failed_to_start`` for runner exits. Genuine failures
keep their own distinct codes, so the cause is preserved end to end and
cleared on the next ``running`` edge like other failure labels.
Frontend (ap-web SubagentsPanel):
- Add a ``disconnected`` variant to the AgentActivity union with an amber
(non-destructive) DOT_TONE entry and a dedicated status pill.
- In ``childStatus()`` and ``sessionStatus()``, branch to ``disconnected``
when the error code is ``runner_disconnected`` / ``runner_failed_to_start``
BEFORE the generic failed branch. Any other failure cause still renders
the red "Failed" pill.
Tests:
- Backend: assert the relay persists the code-preserving
``runner_disconnected`` labels on tunnel close.
- Frontend: assert child + main rows read "Disconnected" (amber, not red)
for the disconnect codes, and still "Failed" for a genuine failure.
Co-authored-by: omnigent <noreply@omnigent.ai>
* style(subagents): recolor disconnected dot blue and hide its inline word
The "Disconnected" pill read amber (--warning) with an inline word. Amber
is shared with the "Needs response" badge, and the word made a benign
liveness loss read louder than the quiet idle/done states.
- Add a dedicated --disconnected blue token (light #2f7fd4, dark #5ca4f5)
wired through the Tailwind @theme block as bg-disconnected; the shared
amber --warning is untouched so "Needs response" stays amber.
- Point the disconnected dot at --disconnected and flip QUIET_STATE so it
renders dot-only (no inline "Disconnected" word), like idle/done. The
hover tooltip / aria-label still carries the error's first line.
- Branch mapping (RUNNER_DISCONNECT_CODES, disconnected-before-failed) is
unchanged for both the main and child rows; genuine failures stay red.
Co-authored-by: Isaac
* test(subagents): harden disconnected-dot coverage from cross-review
Test-only hardening; no visual/routing/condition changes.
- Parametrize the MAIN-row quiet-blue-dot test over BOTH runner-disconnect
codes (runner_disconnected + runner_failed_to_start), mirroring the
child-row it.each so neither code can regress on the main row.
- Add a positive quiet-dot guarantee on both rows: the disconnected pill
routes through the generic quiet-dot path (wrapper keeps the standard
text-muted-foreground, same as idle/done) and the blue bg-disconnected
dot is the only color hook — no warning/destructive bleed on the wrapper
or the dot. No inherited text-color bug found, so no styling change.
Co-authored-by: Isaac
* ui(subagents): swap grey<->blue across pill states (disconnected stays grey)
Reassign which existing token each Subagents-panel pill state uses, scoped
to this panel only — the global --muted-foreground (grey) and --disconnected
(blue) values are unchanged.
- launching: bg-muted-foreground/70 -> bg-disconnected/70 (+ word text-disconnected)
- idle: bg-muted-foreground/55 -> bg-disconnected/55
- done: bg-muted-foreground/55 -> bg-disconnected/55
- disconnected: bg-disconnected -> bg-muted-foreground (quiet dot, no word)
- other (verbatim status fallthrough): stays bg-muted-foreground/55 (exception)
Word visibility, tooltips/aria-labels, running/failed/needs-response, the
runner-disconnect branch ordering, and the global tokens are all unchanged.
Co-authored-by: Isaac
* refactor(subagents): rename --disconnected color token to --session-active
The token was named --disconnected but held the BLUE hue used for the
session-alive-but-not-working states (launching/idle/done). The actual
disconnected state uses grey --muted-foreground. Rename the token (and its
Tailwind --color-* mapping and bg-/text- utilities) to --session-active so the
name matches its meaning. Pure name rename: all hex values, colors, and logic
are unchanged.
Co-authored-by: Isaac
* style(subagents): apply prettier formatting to disconnected details
Collapse the ``details`` ternary in ``childStatus`` onto one line so the
web-prettier hook (and the npm test format:check) pass — CI flagged it as
the sole formatting drift.
Co-authored-by: Isaac
* test(e2e-ui): regenerate chat visual baseline for session-active dot
The subagent quiet-state palette change repointed the done/idle dot to the
new blue --session-active token, so the committed chat snapshot no longer
matched. Adopt the CI-rendered baseline from the pinned Playwright image
(byte-identical to the gate) so the visual check passes; only the dot color
differs.
Co-authored-by: Isaac
* fix(sessions): clear persisted disconnect labels on runner recovery
A disconnect persists durable last_task_error labels (runner_disconnected)
so an ongoing disconnect still projects a "Disconnected" pill after reload.
But runner recovery flips the cached failed status back to idle without a
running edge, so nothing cleared those labels — a healthy reconnected-to-idle
session kept reporting runner_disconnected and the Subagents panel kept the
grey "Disconnected" dot until the next message.
Make _publish_runner_recovered_status async and clear the persisted labels
inside its recovery guard (single source of truth), threading
conversation_store through the two recovery call sites. The durable
persistence itself is unchanged, so the label still survives reload during
an actual ongoing disconnect.
Co-authored-by: Isaac
* fix(sessions): clear disconnect state on runner reconnect-to-idle
A runner tunnel can drop and reconnect to an idle session with no new
turn (a transient WS blip; the runner process survives). On reconnect,
_on_runner_connect re-posted /v1/sessions and restarted the relay but
never cleared the persisted disconnect state, so the session stayed
status=failed with last_task_error.code=runner_disconnected and the
Subagents panel kept the grey "Disconnected" dot until the next message.
Wire the existing _publish_runner_recovered_status helper into
_on_runner_connect so a reconnect drops the stale disconnect state as
soon as the runner is reachable again.
Narrow the helper's guard so recovery only clears a *disconnect*
failure: it now reads the persisted last_task_error code and returns
unless it is runner_disconnected. A genuine task failure (any other
code) survives the reconnect/rebind with its red "Failed" state intact
instead of being silently flipped to idle. This tightens all three call
sites (reconnect, message-forward, PATCH-rebind) to the helper's
documented disconnect-recovery intent.
Co-authored-by: Isaac
* fix(sessions): scope disconnect-code guard to passive reconnect only
The recovery narrowing that clears a stale ``failed`` status only when
the persisted ``last_task_error.code`` is ``runner_disconnected`` was
applied globally, so explicit rebinds/handshakes stopped clearing
genuine stale-failed sessions and broke the PATCH-rebind path.
Gate the guard behind a new ``require_disconnect_code`` flag on
``_publish_runner_recovered_status`` (default ``False`` = clear any stale
failed, still clearing labels). Only the passive tunnel-reconnect caller
(``_on_runner_connect``) passes ``require_disconnect_code=True`` so a
silent reconnect cannot erase a real task failure; the message-forward
handshake and PATCH-rebind keep their clear-any-stale behavior.
Isolate the two reconnect tests from the module-global
``_session_status_cache`` via a snapshot/clear/restore fixture so they
are deterministic in the full integration suite, not just in isolation.
Co-authored-by: Isaac
* test(e2e-ui): regenerate chat baseline for merged tree
After merging main, the chat baseline must reflect both this branch's
session-active blue dot and main's hover-copy-button layout (#1900).
Neither pre-merge baseline had both, so the visual gate failed. Adopt
the byte-exact render the UI Snapshot gate produced for the merge
commit in the pinned Playwright image.
Co-authored-by: Isaac
---------
Co-authored-by: omnigent <noreply@omnigent.ai>
Co-authored-by: Daniel Lok <daniel.lok@databricks.com>
Rework the "Draft release notes" summarizer so the generated highlights
stay user-facing. The drafter now excludes security fixes/hardening and
CI/build/tooling/internal churn from the bug-fixes section, and the
"Bug fixes & hardening" heading becomes plain "Bug fixes" (user-facing
bug fixes only — crashes, reliability, correctness).
Breaking changes get their own section rather than being lumped in with
bug fixes, ordered Features -> Breaking changes -> Bug fixes. An empty
Breaking changes section is omitted entirely by the LLM drafter.
Updates the mechanical scaffold (DRAFT_SECTIONS), the drafter agent
prompt, RELEASING.md, and the changelog tests to match.
Co-authored-by: Isaac
The draft GitHub release now uses the `## [<version>]` section of
editors/vscode/CHANGELOG.md as its notes (only that version's block, up to the
next heading), instead of a generic one-liner. Falls back to a generic note if
no matching section exists, and appends the secure-repo publishing footer.
Co-authored-by: Isaac
When package.json is already at the requested version (e.g. a first release
prepared by hand), the bump + CHANGELOG steps stage nothing, so `git commit`
failed with "nothing to commit" and the release branch never got pushed —
leaving vscode-extension-release.yml with no branch to build from.
Now, on a non-dry run with no staged diff, push release/vscode-v<version> at the
current commit and skip the PR. The build workflow can still build the frozen
.vsix from the branch.
Co-authored-by: Isaac
* fix(editors): use an OpenAI-surface model for the CHANGELOG drafter
databricks-claude-opus-4-8 is only served on the gateway's /anthropic surface,
so POSTing it to /chat/completions 400s (seen in a dry-run of the release-PR
workflow). Switch to databricks-claude-sonnet-4-6 — the id auto-assign-reviewer.yml
already uses on the same endpoint.
Co-authored-by: Isaac
* Apply suggestion from @serena-ruan
Build the .vsix from the frozen release/vscode-v<version> branch instead of
main, so commits landing on main mid-release can't leak into the artifact. The
release PR is merged only after the tag is cut.
- vscode-extension-release.yml: take a `version` input, check out
release/vscode-v<version>, verify the branch's package.json matches, and
target the frozen branch commit.
- Add a `dry_run` input (default true) to both workflows: the release-PR run
shows the bump+CHANGELOG diff without pushing/opening a PR; the release run
builds+checksums without creating the draft release.
- PUBLISHING.md: rewrite "Steps to release" for the freeze-first flow (cut
branch → build from branch → publish draft → merge PR) and document dry_run.
Co-authored-by: Isaac
* feat(web): rename sidebar "Chats" section to "Sessions"
The sidebar's flat session list was headed "Chats" while its create
button reads "New session", so the two disagreed on what a conversation
is called. Rename the visible header to "Sessions" to match.
Only the displayed label changes: the section's persisted collapse-state
key stays "Chats" (as does the drop-zone / hotkey-ordering identity), so
an existing user's collapse preference survives the rename with no
migration. A comment at the call site documents the label/key split.
Co-authored-by: Isaac
* Apply suggestions from code review
Co-authored-by: Daniel Lok <daniel.lok@databricks.com>
vscode-release-pr.yml now drafts the new version's CHANGELOG section from the
PRs merged into editors/vscode since the last release, so the coordinator only
reviews/edits on the PR instead of writing it by hand.
- Harvest merged-PR titles + their `## Changelog` lines since the previous
vscode-v* tag.
- Draft user-facing bullets with a single stdlib urllib POST to the gateway's
OpenAI-compatible /chat/completions (same pattern as auto-assign-reviewer.yml)
— no Omnigent runtime, uv sync, or Claude Code CLI. Fail-open: missing creds,
API error, or empty result keeps the placeholder, so the PR is never blocked.
- Secret-scan the model output for LLM_API_KEY before injecting it.
Also update PUBLISHING.md to use dedicated OMNI_VSCE_TOKEN / OMNI_OVSX_PAT
secrets (separate from databricks-vscode's) so the two teams' release schedules
and revoke-after-release step can't conflict.
Co-authored-by: Isaac
* fix(server): friendly landing for browsers on an API-only (no web UI) server
A server built without the web UI bundle (API-only mode, or an install that
skipped the web UI) served a bare {"detail":"Not Found"} JSON to a browser
opening "/" or a deep link like /c/<conversation_id> — a confusing dead end
for anyone who clicked the conversation URL the CLI advertises.
Serve a short, theme-aware HTML page instead that names the API-only state and
how to install the web UI — but ONLY for a real browser navigation, and ONLY
when no web UI is bundled. Implemented as a 404 exception handler keyed on
Sec-Fetch-Mode: navigate (falling back to Accept: text/html when Sec-Fetch
headers are absent), so:
- programmatic clients (curl, requests, httpx, Go, fetch/XHR — all default to
Accept: */*) keep the exact JSON they got before;
- /api, /v1, /auth always return JSON, even to a browser;
- the "/" metadata is unchanged;
- handler-raised 404s keep their custom detail, and 405s are untouched (a
404-status handler, not a catch-all route, so an unmounted POST route still
404s rather than 405s).
Adds 8 tests covering the browser-navigation, programmatic-client, and
API-namespace paths, including the Sec-Fetch precision case (a browser
fetch() with Accept: text/html still gets JSON).
Co-authored-by: Isaac
* fix(server): API-only landing guidance covers both source and installed
Addresses review feedback (daniellok-db): the landing page only told users
to reinstall, missing the common from-source case. The page can't detect
which situation it's in (it keys solely on whether static/web-ui/index.html
exists), so route by install type instead of assuming one:
- From source: cd ap-web && npm install && npm run build (Vite outDir points
at the dir the server serves), then restart.
- Installed (uv/pip/brew): clear the cache and reinstall. Add the missing
`uv cache clean omnigent` step — `--reinstall` alone can re-serve a cached
UI-less wheel — and call out OMNIGENT_SKIP_WEB_UI as the build-time cause.
Also drop the stale "Node.js 22+" (release CI builds on Node 20) and note
that `npm run dev` runs a separate dev server and won't fix this page.
Co-authored-by: Isaac
* fix(server): correct API-only landing guidance — UI-less is build-time only
A normal install always includes the web UI (the release pipeline gates the
wheel on the bundle being present, and setup.py errors out — rather than
silently skipping — if the npm build fails). So the previous "Installed
(uv/pip/brew) → check OMNIGENT_SKIP_WEB_UI" framing was misleading: a wheel
install ignores that build-time flag and can't land here.
Reframe around the only real causes: a source checkout that hasn't built the
UI, or a build where the UI was deliberately skipped (OMNIGENT_SKIP_WEB_UI),
possibly via a cached UI-less build being reused. Drop the bare
`uv tool install --force --reinstall omnigent` — it can pull an unintended
version (per review) — in favor of clearing the cache and reinstalling the
spec the user originally used.
Co-authored-by: Isaac
* refactor(server): simplify API-only landing — always serve HTML at / (review)
Per review (#908): the browser/Sec-Fetch content-negotiation was convoluted,
and `/` isn't used for anything else. Simplify:
- When no web UI bundle is present, always serve the landing HTML at `/` with a
200 — drop the browser-navigation detection, the JSON-vs-HTML negotiation, and
the 404 exception handler (unmatched paths get the default JSON 404 again).
- Move the HTML out of app.py into omnigent/server/_api_only_landing.py so the
app definition isn't cluttered by a large constant string.
- Rewrite the tests to the new contract (always HTML 200 at /, JSON 404
elsewhere, real routes unaffected).
Co-authored-by: Isaac
* test(server): update root integration test for the HTML landing
The integration test still expected JSON metadata at GET / when no web UI was
present; this PR serves the friendly HTML landing there (200). Update it to
assert the HTML page instead of JSON (it was doing resp.json() and hitting
JSONDecodeError on the HTML body).
Co-authored-by: Isaac
* refactor(server): serve API-only landing from a static .html file
The landing markup is pure static HTML with no interpolation, so a
Python string constant in its own module bought nothing. Move it to
omnigent/server/static/api_only_landing.html and serve it with
FileResponse; ship it in the wheel via package-data. Drops the
_api_only_landing.py module and the HTMLResponse import.
Co-authored-by: Isaac
* fix(server): update landing HTML to reference the renamed web/ folder
The ap-web folder was renamed to web; point the from-source build
instructions at `cd web` to match.
Co-authored-by: Isaac
---------
Co-authored-by: Daniel Lok <daniel.lok@databricks.com>
* fix(web): return to prior conversation from settings back button
The "Back to Omnigent" link in the settings sidebar was hardcoded to
navigate to "/", so leaving settings always dropped the user on the main
landing page instead of the conversation they were viewing. Settings
renders into the shared AppShell outlet under a URL (/settings) that
carries no conversation id, so the link had no context to return to.
Track the last non-settings location (path + search, so ?file= etc. are
preserved) in the Sidebar, which stays mounted across the transition, and
point the back link at it — falling back to "/" when nothing was tracked.
Co-authored-by: Isaac
* test(e2e-ui): cover settings back returning to prior conversation
Drives the real in-app flow — open a conversation, open Settings from the
sidebar, click "Back to Omnigent" — and asserts the URL returns to the
conversation instead of the home landing page. Satisfies the e2e-ui-required
gate for the user-facing navigation fix.
Co-authored-by: Isaac
* feat(web): add hover copy button to user message bubbles
Users could copy assistant responses but had no way to copy their own
messages. Add a Copy action below the user bubble mirroring the assistant
bubble's control: on desktop it's hidden until hover/focus, and on mobile
(no hover) it stays greyed and visible by default.
Co-authored-by: Isaac
* test(e2e-ui): cover user message copy button
Send a message, click Copy under the user bubble, and assert the text
lands on the clipboard and the icon flips to its copied (check) state.
Co-authored-by: Isaac
* test(e2e-ui): regenerate visual baselines
---------
Co-authored-by: omnigent-ci[bot] <294685417+omnigent-ci[bot]@users.noreply.github.com>
## Related issue
N/A
## Summary
- Rename the community harness plugin mechanism from
`omnigent.community.harnesses` to `omnigent.community.harness`.
- Rename the namespace package directory
`omnigent/community/harnesses/` -> `omnigent/community/harness/`.
- Update `COMMUNITY_ENTRY_POINT_GROUP` and `COMMUNITY_MODULE_PREFIX` in
`omnigent/harness_plugins.py` (the entry-point group community plugins
declare and the import-path prefix core validates plugin modules
against), plus the module docstring.
- Update all references in the design doc and plugin tests.
- Note: this is a breaking change for any published community harness
plugin, which must update its entry-point group and module namespace
to `omnigent.community.harness.*` or core will reject it at load time.
## Test Plan
- `uv run pytest tests/test_harness_plugins.py` — all 8 tests pass.
- `uv run python -c "import omnigent.community.harness; import omnigent.harness_plugins as hp; print(hp.COMMUNITY_ENTRY_POINT_GROUP, hp.COMMUNITY_MODULE_PREFIX)"`
confirms the namespace imports and the constants read back as
`omnigent.community.harness` / `omnigent.community.harness.`.
- Repo-wide grep confirms no remaining `community.harnesses` references.
## Type of change
- [ ] Bug fix
- [ ] Feature
- [x] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [x] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
The existing plugin unit tests in `tests/test_harness_plugins.py` were
updated to the new namespace and all pass. Manually verified the renamed
namespace package imports and that the two module constants resolve to
the new group/prefix, and grepped the repo to confirm no stale
`community.harnesses` references remain.
* feat(runner): authenticate managed-sandbox runner HTTP callbacks under accounts/OIDC
Managed runners mint a short-lived owner JWT from POST /v1/runners/{id}/token
(authenticated by the tunnel binding token) and present it on HTTP callbacks,
so require_user-gated routes resolve the owner instead of 401ing. Closes the
HTTP half of #357; builds on the tunnel-owner resolution from #360.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* chore: regenerate openapi.json for POST /v1/runners/{id}/token
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* fix(runner): re-arm managed-mint factory after a transient boot-probe failure
Address Polly review note: the construction probe declined to install the
factory on ANY failure, so a blip at the instant the runner boots left it
unauthenticated until restart. Now it only declines on a definitive no-mint
(HTTP 400 no-auth/header, 404 old server); a transient failure installs the
factory so the next callback re-mints.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* docs: explain intentionally-swallowed exceptions in mint probe and health poll
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* fix(runner): latch managed-mint decline at request time; send bare requests instead of failing closed
The construction probe can lose a boot race (connection refused while
the server is still starting), which installs the managed mint factory.
Every later mint then gets the definitive HTTP 400 of a no-auth server,
the factory returns None, and _RunnerDatabricksAuth fails closed --
bricking every runner->server callback (spec_resolver_failed across the
integration/E2E suites).
Latch the definitive 400/404 decline inside the factory and have
auth_flow send bare requests once declined, matching the no-factory
behavior the construction probe would have chosen.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
---------
Signed-off-by: dbczumar <corey.zumar@databricks.com>
## Related issue
N/A
## Summary
- Add a discreet info (ⓘ) button to the top-trailing corner of the iOS
connect screen — hidden but discoverable, and always reachable since the
connect screen is the app's entry point.
- Tapping it opens a menu with Website, Documentation, and Privacy Policy
links (omnigent.ai, omnigent.ai/docs, omnigent.ai/privacy), satisfying the
need for an in-app privacy policy link.
- Present each link in an in-app Safari sheet via a new `SafariView`
(`SFSafariViewController` wrapper) so users stay inside the app rather than
being kicked out to the system browser.
- Trim the connect screen's server-URL description to a single line.
## Test Plan
- `swift format lint` passes on the changed files.
- `xcodebuild -scheme Omnigent` builds successfully with the new source file
wired into the project.
- Ran the app on the iPhone 17 Pro simulator and confirmed the info icon
renders on the connect screen; verified the menu opens and links present the
in-app Safari sheet.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] Refactor / chore
- [ ] Docs
- [ ] Test / CI
- [ ] Breaking change
## Test coverage
- [ ] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Verified manually: the change is UI-only (a SwiftUI info menu and an
SFSafariViewController wrapper on the connect screen) with no automated UI
test harness in this target. Confirmed via a clean build and running the app
on the simulator that the info icon appears and the menu links open the
in-app Safari sheet.
dbczumar is out of office for a while, so stop routing new issues/PRs to
him. Rather than delete him, move his login from `owners` to a sibling
`owners_paused` array in each of the 18 areas he owned. Every reader (the
reviewer JS, issue-triage, areas.test.js) only consults `owners`, so
`owners_paused` is inert -- reverting when he's back is just moving the
login back into `owners`, no git archaeology.
harness-cursor was [SabhyaC26, dbczumar]; since every area needs 2+ active
owners (enforced by areas.test.js), dhruv0811 takes the active seat there
while dbczumar sits in owners_paused like everywhere else.
Co-authored-by: Isaac
## Related issue
N/A
## Summary
- Add a fastlane `snapshot`-based App Store screenshot pipeline: a new
`screenshots` lane rebuilds the web UI, boots an isolated local Omnigent
server on a non-6767 port (own HOME/data/logs dirs), and drives the
`OmnigentUITests/testLocalServerSnapshot` UI test to capture en-US
screenshots into `fastlane/screenshots`.
- Add DEBUG-only launch hooks so the snapshot run is deterministic: the app
reads its server URL from `--omnigent-server-url` /
`OMNIGENT_SCREENSHOT_APP_URL`, skips auto-opening the saved server, and
suppresses the notification authorization prompt during snapshots.
- Rename the `release` lane to `prod` — prepares the App Store version from an
already-uploaded TestFlight build, reusing metadata + screenshots.
- Add `PrivacyInfo.xcprivacy` privacy manifest, App Store metadata files
(copyright, support URL), accessibility identifiers on the connect form, and
a shared `SnapshotHelper.swift`.
- Drop the iPad-specific `UISupportedInterfaceOrientations~ipad` keys from the
Debug/Release Info.plists.
## Test Plan
- `bundle exec fastlane screenshots` — builds the web UI, starts the isolated
local server, runs the snapshot UI test, and writes screenshots to
`fastlane/screenshots/en-US`.
- `bundle exec fastlane tests` — `OmnigentTests` unit suite still passes with
UI tests skipped.
## Type of change
- [ ] Bug fix
- [x] Feature
- [x] UI / frontend change
- [ ] Refactor / chore
- [ ] Docs
- [x] Test / CI
- [ ] Breaking change
## Test coverage
- [ ] Unit tests added / updated
- [ ] Integration tests added / updated
- [x] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Added the `testLocalServerSnapshot` UI test that drives the connect flow
against a local server and captures screenshots. Verified manually by running
`bundle exec fastlane screenshots` end-to-end and confirming the en-US
screenshots are produced. The DEBUG-only launch hooks are exercised by that
test path and gated out of Release builds.
The expanded shell terminal card cleared the 56px chat header with pt-16
(64px) while the workspace rail uses mt-14 (56px), leaving the terminal
card top 8px lower than the rail. Use pt-14 to match the header height so
the two panel tops line up.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
Copilot's ``assistant.usage`` event reports cache-creation tokens under
``cacheWriteTokens``, but ``_accumulate_usage`` only mapped input/output/
cacheRead, so cache-write tokens were dropped from ``TurnComplete.usage``.
The server cost path (``_accumulate_session_usage`` -> ``compute_llm_cost``)
prices ``cache_creation_input_tokens`` at the cache-write rate, so dropping
them under-counted cost and left the cache breakdown incomplete in telemetry
and the web UI.
Map ``cacheWriteTokens`` -> ``cache_creation_input_tokens`` (the
Omnigent-standard key, matching the cursor harness). Verified live against a
real Copilot turn: a first turn reported ``cacheWriteTokens=14144`` that was
previously discarded.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Co-authored-by: Corey Zumar <39497902+dbczumar@users.noreply.github.com>
* feat(cli): "!" shell passthrough — run a command, fold output into the next turn
A REPL line starting with "!" runs the rest in the user's shell, shows the
output, and folds it into the next agent turn so the assistant can reason about
what ran. "!!" sends a literal leading "!"; a bare "!" prints a usage hint.
- Cross-platform: `$SHELL -c` on POSIX, `%COMSPEC% /c` on Windows.
- Non-interactive (stdin=/dev/null) and timeout-bounded; stdout/stderr captured
separately; ANSI preserved on screen, stripped for the model.
- Buffer model: a bare "!cmd" costs no model turn — output is folded into the
next message's llm_text (ANSI-stripped, capped).
- Lightweight cwd persistence: a standalone "!cd <dir>" changes the directory
later "!" commands run in (a compound "cd x && …" does not persist).
- Huge output spills to a temp file (referenced in the block) instead of being
dropped, so the agent can read it in full.
- Env knobs: OMNIGENT_BANG_TIMEOUT_S (120) / _DISPLAY_MAX (30k) / _CONTEXT_MAX (16k).
Tests (tests/repl/test_bang_command.py): clip; the model-facing context builder
(exit, fences, no-output, ANSI strip, capping, overflow note); cross-platform
shell selection (POSIX + Windows); cd resolution; temp-file overflow; and the
async runner against real commands (echo, non-zero exit, stderr, cwd,
timeout-kills). POSIX-shell tests marked posix_only.
Co-authored-by: Isaac
* test(cli): e2e coverage for "!" passthrough; green composer + echo highlight
- tests/e2e/omnigent/test_repl_bang_e2e.py: drive the real REPL under pexpect —
render + fold-into-next-turn, bare-! hint (no turn), and !! escape.
- Highlight "!" shell input in the omnigent-logo green (#26a079): a composer
lexer while typing, and the echoed command line once it runs.
- Unit tests for the lexer + echo color in tests/repl/test_bang_command.py.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* fix(cli): address Polly review — drop "!" buffer on new conversation
- Clear _pending_bang_blocks on /clear and /new so buffered shell output can't
leak into a fresh conversation's first turn (with e2e coverage).
- _write_bang_overflow: measure the model-facing (ANSI-stripped) size for the
spill trigger, matching the context builder; document the temp-file lifecycle.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
---------
Signed-off-by: dbczumar <corey.zumar@databricks.com>
Co-authored-by: dbczumar <corey.zumar@databricks.com>
* fix(codex-native): picker readiness mirrors the launch resolver, not auth.json
The web picker showed "needs Codex authentication on <HOST> — run `codex
login`" for a Databricks-gateway setup even though codex ran fine.
`_codex_auth_unavailable_reason` only inspected `~/.codex/auth.json`, but
`resolve_native_codex_launch` routes a gateway/provider setup through a
Databricks profile or a `model_provider` override and mints its bearer at
run time (`databricks auth token`) — it never reads auth.json. So auth.json
is legitimately empty and gating on it is a false negative.
Make readiness ask the same question the launch resolver already answers:
available when the launch routes through a provider (profile set, or a
non-`openai` model_provider); fall back to the auth.json check only on the
bare-`codex login` path where auth.json actually is the credential. Reuses
two functions already imported in the module — no new imports, no network
probe. Mirrors the fail-open the claude-sdk / openai-agents gateway
harnesses already rely on.
Co-authored-by: Isaac
* style: ruff format codex_native.py
Co-authored-by: Isaac
* test(harness-bench): --transport wiring + semantic driver protocol
Make the bench's probes run through a selectable transport. Introduces a
Driver protocol (transport.py) with four semantic per-dimension methods —
run_basic_turn, run_streaming_turn, run_tool_turn(deny), run_interrupt_turn
— that both drivers implement. The driver owns the mechanism (request-level
tool + verdict-post deny on the wrap path; builtin tool + spec-baked deny
policy + SSE subscribe on full-server); the probe owns interpretation.
- transport.py: Driver protocol, driver_registry(), resolve_driver_class()
where a --transport override wins over the profile's declared transport.
- SdkInprocDriver + FullServerDriver both implement the four methods;
full-server bridges its sync provisioning/turns to async via
asyncio.to_thread.
- All six probes refactored to call the semantic methods (no more
wrap-specific run_turn kwargs / per-probe tool specs); base.run() typed
against the Driver protocol.
- bench.run_harness/run_bench + the CLI take a transport override
(--transport). Unknown transport fails loud.
- interrupt probe: check result.cancelled BEFORE the delta-count guard, so
a transport that confirms cancellation via a marker (full-server) rather
than a delta count is not falsely SKIPPED.
Verified live on oss: sdk-inproc matrix unchanged; --transport full-server
runs all six probes and fills Tool calling + Policy DENY (·->✓) via real
server dispatch + enforcement, no unexpected DRIFT.
* test(harness-bench): address #1870 review (transport.py stubs, CLI transport guard, shim test)
From the Polly + code-quality review on #1870:
- transport.py Driver protocol: drop the redundant '...' after each
docstring (code-quality 'statement has no effect' x7) — a docstring-only
body is the Protocol stub form. Also drop @runtime_checkable (nothing does
isinstance; it wouldn't cover the data/static members anyway) and document
why.
- CLI: validate --transport against the registry up front, returning a clean
exit-2 error instead of a raw KeyError traceback out of asyncio.run.
- interrupt probe: document the full-server measurement gap (a harness that
IGNORES an interrupt surfaces only via timed_out, else SKIPPED) at the
guard.
- Add an offline test that the FullServerDriver async shims
(__aenter__/__aexit__ + the four run_* to_thread bridges) delegate to the
sync methods, so a regression in the async binding is caught without a
live server.
Offline 18 passed / 4 skipped, ruff + pre-commit clean.
* fix(setup): show the actual install command for optional SDK extras
The setup flow and executor error messages hardcoded `pip install
"omnigent[X]"` regardless of how omnigent was installed. When uv was
available it silently ran `uv pip install` instead, and for `uv tool`
installs neither command could reach the isolated tool venv.
Extract a shared `extra_install` helper that detects the install method
(uv tool / uv / pip) and returns the matching command. All UI surfaces
now display the command that actually runs.
* fix(tests): update install-command tests for shared extra_install helper
Update test mocks to target `extra_install.shutil`/`extra_install.sys`
instead of the removed `*_auth.shutil`/`*_auth.sys` imports. Replace
hardcoded `pip install "omnigent[X]"` assertions with dynamic checks.
Add `uv tool` install path tests for all three harnesses.
* style: fix formatting in install-command tests
* fix(review): add UV_TOOL_DIR caveat and direct _is_uv_tool_install tests
Address Polly review feedback:
- Add docstring note about UV_TOOL_DIR/XDG_DATA_HOME false negatives
(mirrors accepted pipx heuristic gap).
- Add direct parametrized tests for _is_uv_tool_install() covering
Linux, Windows, venv, system, and pipx prefixes.
* style: fix formatting in test_extra_install.py
* fix(setup): keep git-source uv tool installs on their source when adding extras
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* refactor(setup): bind executor install hints to the harness extra constants + guard against pyproject drift
Signed-off-by: dbczumar <corey.zumar@databricks.com>
---------
Signed-off-by: dbczumar <corey.zumar@databricks.com>
Co-authored-by: dbczumar <corey.zumar@databricks.com>
* feat(claude-native): render live tool-call cards in the web chat UI
Native Claude Code sessions already mirror their tool calls (Read/Bash/
Grep) into the web chat, but the cards rendered static (no spinner, no
elapsed timer) so the only live activity signal was a generic "Working…".
The cause: the frontend's live-tool styling only activates when a bubble's
lifecycle is "streaming", which requires a streaming activeResponse whose
responseId matches the bubble. Native "running" status is PTY-activity-
derived and carried no response_id, so the UI never entered that lifecycle.
Feed the existing streaming machinery the id native Claude already knows:
- forwarder: _post_external_session_status gains a response_id param; emit
running+response_id once at turn start (deduped on _ForwardDedupeState so
it survives the delta-hold early-return), and stamp the same id on the
Stop->idle / StopFailure->failed edges. PTY badge edges unchanged.
- server: _publish_status tracks the in-flight id in
_session_active_response_cache (set on running/waiting, cleared on
idle/failed); _build_session_response projects it as active_response_id.
- mid-turn reconnect: SessionResponse.active_response_id -> Session
.activeResponseId -> reconnectStatusPatch reopens the streaming
activeResponse from the snapshot (the SSE stream is snapshot + live
tail, no replay).
No new event types or UI components; reuses the session.status channel.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Regenerate openapi.json for active_response_id
The PR added active_response_id to the SessionResponse schema but did not
regenerate the checked-in openapi.json, so test_openapi_drift failed
(server-rest). Regenerate it via scripts/dump_openapi.py — a purely
additive SessionResponse.active_response_id property.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* test(e2e_ui): cover live tool-card render on mid-turn connect
Add a Playwright e2e_ui test for the PR's user-facing behavior: a session
whose snapshot carries active_response_id reopens the streaming lifecycle on
a fresh connect, so a forwarded (output-less) tool call renders as a LIVE card
(running spinner) rather than a static one. Seeds the exact
external_session_status(running, response_id) + external_conversation_item
(function_call) a native forwarder emits, asserts the snapshot projects
active_response_id, then asserts the transcript shows the running spinner on
both initial load and reload. Extends the existing working-indicator-reload
suite and its _publish_status helper.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* fix(e2e_ui): add required agent field to seeded function_call
The live-tool-card e2e test seeded a function_call external_conversation_item
without the required FunctionCallData.agent field, so the events POST 400'd
(E2E UI Tests shard 0/3) before the DOM assertion ran. Add
agent="claude-native-ui" to match the payload shape native forwarders emit.
Verified against a live local server: the status(running,response_id) and
function_call POSTs both return 202, the snapshot projects
active_response_id, and the item persists with the matching response_id.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* fix(claude-native): drop bridge_dir from turn-start warning log
CodeQL (py/clear-text-logging-sensitive-data, high) flagged the bridge_dir
expression in the new turn-start running-status warning as clear-text logging
of sensitive data. The session_id and response_id already identify the failing
forward, and bridge_dir is derivable from the session, so drop it from the log
to clear the new high-severity alert. Same false positive main already carries
on an analogous transcript-item error log, left untouched.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* fix(e2e): restore mock tool-call config in repl refusal test
test_repl_tool_call_refusal_blocks_tool sends "testing456" and waits for the
"approval required" banner, but the tool-call the banner depends on stopped
being scripted: #1839 rewrote the test for the new abort-on-decline behavior
and, along with the now-obsolete follow-up assertions, dropped the
_configure_mock_tool_then_text call. With no route for "testing456" the shared
mock returns no tool call, so no ASK fires and the expect times out at 45s —
passing only when another test on the same xdist worker happens to leave a
tool-call response in the mock's queue (the ordering flake this hit under -n
sharding; the conftest docstring notes -n 8 has ordering flakes -n 4 avoids).
Restore the echo tool-call config (match="testing456") so the ASK fires
deterministically. Verified: fails in isolation before (pexpect TIMEOUT on
'approval required'), passes 3/3 in isolation after.
Signed-off-by: dbczumar <corey.zumar@databricks.com>
---------
Signed-off-by: dbczumar <corey.zumar@databricks.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(codex): CodexExecutor honors os_env.sandbox.env_passthrough
CodexExecutor builds the codex subprocess env from the hardcoded _clean_codex_env()
allowlist and never consulted the agent's declared os_env.sandbox.env_passthrough — so
a codex-harness agent's shell tools could not see secrets the spec explicitly allows
(e.g. an MCP/REST API token), while the claude-sdk os_env path honors the same field.
Adds an extra_allow param to _clean_codex_env() and a guarded _declared_passthrough()
helper that reads os_env.sandbox.env_passthrough. The _CODEX_ENV_DENY_EXACT rule
(strips OPENAI_API_KEY for subscription auth) still wins — a denied var is never
re-admitted even when declared. Opt-in and targeted: only declared names pass, not the
full host env.
Refs #1022 (the env-allowlist-drops-needed-vars discussion; this is the codex-executor
counterpart to the daemon/runner allowlist case).
* fix(codex): satisfy ruff format and restore allowlist comments
Signed-off-by: dbczumar <corey.zumar@databricks.com>
* style(codex): ruff format test file
Signed-off-by: dbczumar <corey.zumar@databricks.com>
---------
Signed-off-by: dbczumar <corey.zumar@databricks.com>
Co-authored-by: dbczumar <corey.zumar@databricks.com>
The bench hand-maintained a second copy of 'what each harness supports'
(manifest._P0_ALL_SUPPORTED verdicts + _STATIC auth/implementation). Make
it derive from the canonical harness_capabilities() (PR #1847) so there is
one source of truth, and the bench's job sharpens to 'does the harness do
what it publicly claims?'.
- Group A (descriptive columns): implementation from integration_mode, auth
from auth, via small enum->prose maps.
- Group B (capability-backed verdicts): streaming from capabilities.streaming
(True->SUPPORTED deltas, False->PARTIAL complete-only), interrupt from
capabilities.interrupt, model_override from model_env_keys() membership.
- Group C (probe-only, kept explicit): basic_turn, tool_calling, policy_deny.
policy_deny is enforcement, NOT the elicitation ASK surface — deliberately
not derived from the elicitation axis.
- Deleted _P0_ALL_SUPPORTED and the derivable _STATIC dict.
- Tolerates sparse capabilities (community plugins): a harness with no
declared capabilities gets only the probe-only dims, no KeyError.
- reconcile() phrasing now reads DRIFT as 'declared capability vs observed
behavior' — the capability table is self-enforcing.
Reads the STATIC harness_capabilities(), not the runtime Executor.supports_*
methods (different layers). Verified live on oss: openai-agents (SDK) and
codex (CLI-subprocess) reconcile with no unexpected DRIFT on
streaming/interrupt/model_override; offline 17 passed, ruff+pre-commit clean.
The #1839 rewrite of this test dropped the _configure_mock_tool_then_text
setup that scripts the mock LLM to emit the echo function_call. Without it,
sending "testing456" produces no tool call, the TOOL_CALL ASK never fires,
and child.expect("approval required") times out after 45s on every run.
This is a deterministic failure, not a flake: the test's final pre-merge E2E
run was skipped by the merge queue, so the config-less version never ran green
before landing, and it has failed the scheduled main run since.
Re-add the tool-call scripting before spawn. The follow-up text is never
reached (the turn aborts on decline before any second LLM call), so only the
function_call scripting is needed; the rest of the post-#1839 body is unchanged.
Verified locally: 3/3 green.
streaming_probe_turn subscribes to GET /v1/sessions/{id}/stream on a
background thread and counts response.output_text.delta events while the
main thread posts the turn; >1 delta means token-level streaming. Gated
live test asserts it. Verified on oss (~10s, 50+ deltas).
interrupt_probe_turn starts a long turn, posts an interrupt once it is
running (after a short hold so text streams first), and confirms the
server's synthetic 'interrupted' cancellation marker appears. Gated live
test asserts the turn is cancelled. Verified on oss (~9s).
* feat(polly): add cursor and hermes coding sub-agents
Adds `cursor` (cursor-native) and `hermes` (hermes-native) to the polly
orchestrator, taking the roster to six: claude_code, codex, opencode, cursor,
hermes, pi. Both are native terminal harnesses (openable / take-over-able in the
Subagents panel), widening cross-vendor review.
- examples/polly/agents/{cursor,hermes}/config.yaml (new): standard implement /
review / explore contract and blast_radius(gate_pushes=false), matching the
peers.
- examples/polly/config.yaml: roster is now six; preflight checks `cursor-agent`
and `hermes`; tools.agents, routing, cancellation notes, and comments updated;
spawn_bounds.max_dispatches_per_turn 5 -> 6 so one fan-out round can launch
every worker.
- examples/polly/skills/{investigate,fanout,cross-review}: cursor and hermes
wired in as full peers (implementer, reviewer rotation, explore lens).
- tests: roster list, per-worker loops, vendor count (4 -> 6), policy count
(7 -> 9), the shipped-bundle declared set, and the brain-override
worker-harness map updated for the two new workers.
The parent-wake plumbing that makes cursor/hermes usable as headless polly
workers lands in the following commit.
* fix(native): wake parent orchestrator when cursor/hermes finish a turn
cursor-native and hermes-native only emitted the PTY watcher's web-spinner
`session.status: idle` edge, which never wakes a parent orchestrator — so as
polly sub-agents they finished silently while claude/codex/opencode/pi woke the
parent via an `external_session_status: idle` POST. Both now post that event
once per completed turn, deduped against a persisted posted-count and
restart-safe.
cursor: the stop hook records a turn-end marker (cursor_native_status); the
forwarder tails it and posts idle. hermes (no stop hook) derives turn-end from
state.db — an assistant row with no tool_calls is the agentic loop's terminal
step. The runner clears the new poster state on terminal recreation so a stale
count can't skip or re-fire the wake.
Ported from the original cursor/hermes/opencode roster work; without it the two
new polly workers added in the previous commit would dispatch and never notify
polly on completion.
* feat(web): give Hermes its own glyph in the Subagents panel
Hermes rendered with the generic omnigent fallback icon because there was no
HermesIcon component and neither icon resolver had a `hermes` case — even though
`iconKind: "hermes"` was already declared on the native-agent spec. Add an
original caduceus glyph (currentColor, matching its sibling icons) and wire it
into AgentCard.getAgentIcon and SubagentsPanel.brandChildIcon so the hermes
polly sub-agent shows its own icon like the other native harnesses.
* style(web): prettier-format HermesIcon path strings
prettier collapses the two split path-string literals onto single lines
(they fit the print width); match it so format:check passes.
* fix(hermes-native): rebase idle posted-count on compaction re-pin
The completed-turn count is keyed per hermes_session_id, but the idle dedup
baseline (posted_count) is per bridge dir. On an in-session compaction the
forwarder re-pins to the forked child (new session_id, count restarts near 0)
without touching posted_count, so the guard completed_turns > posted_count
stayed False until the child exceeded the parent total — suppressing the
child session's early idle posts and hanging a headless polly worker that
compacts mid-task then finishes. Rebase posted_count to the child's current
count on re-pin (where last_id is reset to 0). Adds a regression test that
fails without the rebase, and corrects the clear_hermes_status_state docstring
(count is per hermes_session_id, not per terminal).
Flagged by the Polly AI review on #1844.
* chore(native): drop unused _logger from cursor/hermes status modules
Neither cursor_native_status nor hermes_native_status logs anything; the
_logger = logging.getLogger(__name__) definition and its import logging were
dead (flagged by github-code-quality). Remove both. No behavior change.
* docs(cursor-native): note idle block runs outside the store-gated branch
The cursor idle-post block sits at the poll-loop body level, deliberately
outside the if store_path mirroring branch, so a stop-hook turn-end marker
is picked up even on a poll where the SQLite store is unbound or empty.
Make that placement explicit (per PR review). Comment-only.
* feat(harnesses): declarative capability model on HarnessContribution
Adds the one axis the dynamic harness registry (#1756) does not cover: a
declarative capability model answering "what can this harness do?" across
seven axes (integration_mode, elicitation, resume, effort, model_family,
auth, subagents), aligned with the harness-integration-guide feature matrix.
- omnigent/harness_capabilities.py: import-safe enums + HarnessCapabilities
dataclass, mirroring the harness_install_spec.py pattern so plugins can
declare capabilities during entry-point discovery without import cycles.
- HarnessContribution gains a per-harness `capabilities` dict; the built-in
contribution declares all 23 harnesses. Community plugins can declare their
own the same way, inheriting the registry's built-in-wins + collision guards.
- harness_capabilities() accessor + harness_catalog() now emits a
`capabilities` object per row, surfacing the matrix on GET /v1/harnesses.
Every value is backed by the implementing module; the two derivable axes
(model_family, subagents) are asserted against their source
(model_override family sets; native subagent_wrapper_label) so the table
cannot silently drift.
This supersedes the parallel omnigent/harnesses/ registry explored in the
now-closed #1793/#1795/#1840 stack: rather than a second registry, capabilities
attach directly to #1756's HarnessContribution as the single source of truth.
Co-authored-by: Isaac
* feat(harnesses): add interrupt + streaming capability axes
Extend HarnessCapabilities with two behavior axes the harness bench probes
(interrupt: can a running turn be cancelled mid-stream; streaming: token-level
deltas vs a single blob), so the bench's declared-support matrix can derive
fully from harness_capabilities() rather than a separate hand-maintained table.
The four P0 SDK harnesses (claude-sdk, codex, pi, openai-agents) are declared
interrupt=streaming=True — matching what the bench verifies live today; a test
pins that alignment. The remaining harnesses declare best-effort values that the
bench's interrupt/streaming probes will reconcile as transport coverage expands.
Both axes serialize into the GET /v1/harnesses catalog.
Co-authored-by: Isaac
* docs(harnesses): seam brief for wiring the bench to capabilities
Adds designs/harness-capabilities-bench-seam.md — the handoff contract for the
follow-up that makes tests/harness_bench/manifest.py derive its declared-support
matrix from harness_capabilities() instead of the hand-typed _P0_ALL_SUPPORTED /
_STATIC dicts. Documents the axis mapping (derive descriptive columns + the
interrupt/streaming/model_override verdicts; leave basic_turn/tool_calling/
policy_deny probe-only), the static-vs-runtime capability-layer distinction, the
best-effort confidence caveat for non-P0 harnesses, and the resulting semantic
shift (DRIFT = a harness's published capability claim is false).
Co-authored-by: Isaac
* refactor(harnesses): name the subagents bool in capability entries
The trailing positional bool in each _BUILTIN_CAPABILITIES entry was the
`subagents` flag — the one unlabeled arg (the enum args are self-documenting via
their _EL./_RS./_MF. prefixes, and interrupt/streaming were already named).
Pass it as subagents=... so each entry reads unambiguously. No value changes.
Co-authored-by: Isaac
* fix(harnesses): correct open-responses capabilities; guard capability collisions
Polly review caught the open-responses row contradicting its own executor
(omnigent/inner/open_responses_sdk.py) — the exact anti-drift failure this table
exists to prevent. Verified against the source and corrected:
- interrupt True (interrupt_session closes the active stream, returns True)
- streaming True (supports_streaming returns True)
- effort OPENAI (drives gpt-5.3-codex, forwards reasoning_effort via cfg.extra)
Also close the collision gap flagged in review: add `capabilities` to
_harness_spellings() so a community plugin declaring capabilities for a built-in
harness id is rejected instead of silently overriding it (last-wins in
_merge_dict). Test asserts the rejection.
Co-authored-by: Isaac
* feat(web): support shift-click range selection in multi-session mode
Extract range computation into a pure, tested helper
(computeShiftSelectRange). Sync the visible-IDs ref directly from
orderedConversationIds (synchronous useMemo) instead of populating it
via useEffect in each ProjectFolder — eliminates the stale-ref timing
bug that caused the previous attempt (#1534) to be reverted (#1652).
* fix: prettier formatting for test file and regenerate package-lock.json
* fix(web): use actual rendered project IDs for shift-select ranges
ProjectFolder fetches its own sessions via useProjectSessions, which
can diverge from the global paginated list. Register each folder's
rendered IDs synchronously during render (via useMemo + ref write)
so shift-select ranges match what's on screen. Unlike the previous
useEffect-based approach (reverted in #1652), this avoids stale-ref
timing bugs because the map is populated before the click handler
can read it.
* fix(web): compute shift-select visible order lazily at click time
Address PR review: the previous approach built visibleIdsRef during
ConversationList's parent render, but ProjectFolder children write
their rendered IDs during their own render — which runs after the
parent. This left the project segment one commit behind and stale
when a child re-rendered independently (async query, session re-sort).
Replace the cached string[] ref with a getter function ref that reads
projectRenderedIdsRef lazily when the user actually clicks. The
closure captures sections/collapsed state from the parent render scope
(stable unless the parent re-renders), while projectRenderedIdsRef is
always read fresh because it's a mutable ref.
Add a test proving shift-select within a project folder uses the
folder's own rendered IDs (including sessions not in the global
paginated window).
---------
Co-authored-by: Corey Zumar <39497902+dbczumar@users.noreply.github.com>
* refactor(policies): remove FunctionPolicySpec.action whitelist field
Drop the `action` whitelist from `FunctionPolicySpec` and all
supporting machinery: the `_parse_action_list` parser helper,
the `_action_permitted` validator, and the `_fail_closed`
branching logic that gave classifier-only and approval-gate
policies special substitution behaviour on error.
The engine now unconditionally returns a fail-closed DENY on any
evaluator exception, simplifying the dispatch contract.
* fix: remove stale action field from test and clean up docstrings
- Drop action=[PolicyAction.ALLOW] from test_omnigent_translator.py
(field no longer exists on FunctionPolicySpec)
- Remove unused PolicyAction import in that test
- Remove action from prompt-policy pass-through docstring in omnigent.py
- Remove stale "omit ASK from action list" guidance in ask_timeout
error messages in parser.py
Adds a keyless DuckDuckGo HTML backend (search_provider: duckduckgo) so
web_search can run with no API key. Not the default — with no search_provider
set, _search() fails loud with a helpful message naming the engines (per
review). Includes hardening, a real-response golden fixture + offline tests,
and a nightly live drift canary.
Co-authored-by: Isaac
* feat(editors): add VS Code extension release + publishing workflows
Set up the release path for the omnigent-vscode extension. The extension
publishes under the shared databricks Marketplace publisher, so releases flow
through the security-hardened secure-release repo — this repo only builds a
SHA256-verified .vsix and attaches it to a draft GitHub release.
- vscode-release-pr.yml: manually-dispatched, opens a reviewed version-bump +
CHANGELOG PR (write-or-higher actor check) so the tag can't diverge from
package.json.
- vscode-extension-release.yml: manually-dispatched, builds the .vsix + .sha256
and cuts a draft vscode-v<version> release (namespace kept separate from the
Python v[0-9]* tags).
- docs/vscode-extension-publishing.md: end-to-end release steps + one-time
setup table.
- Set publisher to "databricks"; add the extension CHANGELOG.
Co-authored-by: Isaac
* docs(editors): move publishing guide into editors/vscode
Keep the VS Code extension's publishing guide alongside the extension it
documents. Update the release-PR workflow's reference to the new path.
Co-authored-by: Isaac
* docs(editors): add local .vsix smoke-test step before marketplace publish
Verify the packaged extension installs and activates in a clean VS Code
before it reaches the marketplaces.
Co-authored-by: Isaac
* docs(editors): clarify the local smoke-test expected result
Replace the "frames it" jargon with a plain description of what to see.
Co-authored-by: Isaac
* fix(editors): enforce strict X.Y.Z extension versions
vsce package rejects prerelease-suffixed versions, so accepting them in the
release-PR workflow could land a version bump on main that then fails at
package time. Validate strict major.minor.patch, and drop the now-dead
pre-release detection in the release workflow.
Co-authored-by: Isaac
* fix(inbox): only surface comments from other people in the inbox
The comment side of the inbox was echoing your own comments back at
you. The filter only dropped a comment when authorship was known
(`viewerId` non-null and matching `created_by`), so single-user
deployments — where every comment is stored with `created_by = null` —
kept showing all of them, and a private session you own showed nothing
useful either.
Tighten the rule to what the inbox is actually for: a comment appears
only if an identifiable *other* person wrote it. A comment can only
carry another user's `created_by` if that user had access, so this also
implies "the session is shared with that person" without needing the
grant list. Consequences: an unshared/private session (and single-user
mode) now contributes an empty comment inbox, while a shared session
still surfaces collaborators' comments and hides your own.
Co-authored-by: Isaac
* test(e2e): assert own comments never surface in the inbox
Adds an e2e_ui case covering the inbox author filter: a comment stored
with created_by = null (authored by the local viewer, as in single-user
mode or a private session) must not appear in the inbox even though the
session reports an unseen draft. Complements the existing test where a
collaborator's comment does surface.
Co-authored-by: Isaac
Replaces the gap-only spacing in the AgentInfoContent popover with
divide-y borders so each section has a clear visual boundary. Also
merges session cost and token usage into a single section.
Follow-up to the HTML-preview comment feature, addressing Polly review
findings:
- Blocking: findAnchorInSource's occurrence-0 fast path used a verbatim
indexOf, which disagreed with the whitespace-normalized occurrence count
the in-frame bridge produces. When an earlier rendered copy was
whitespace-wrapped in the source and a later copy was verbatim, selecting
the first copy anchored the comment to the later one. Dropped the fast path;
always walk whitespace-tolerant occurrences.
- Occurrence counting now skips non-rendered source regions (tag markup and
attribute values, HTML comments, <script>/<style>/<title>/<noscript>) so the
parent's Nth source match lines up with the Nth *rendered* match the bridge
counts over body text nodes.
- Unified the whitespace definition: the parent now folds runs of code points
<= U+0020 (matching the in-frame normWs) instead of regex \s, which also
folds U+00A0 and other Unicode spaces and could diverge from the bridge.
- Perf: repaint() builds the normalized whitespace map once per call and shares
it across comments instead of rebuilding it per comment in anchorRanges.
Co-authored-by: Isaac
* feat(policies): abort agent turn on explicit elicitation decline
When a user explicitly clicks "Decline" on an elicitation card, the
agent turn now aborts cleanly instead of receiving a DENY message and
continuing. This matches the expected native behaviour where a human
refusal stops the run.
Changes:
- Add ElicitationDeclinedError to omnigent/errors.py — a new exception
that callers can catch to distinguish explicit user decline from
timeout, cancel, or malformed verdict
- Add _is_explicit_decline() to approval.py — detects action=="decline"
strictly (cancel/timeout/None all return False)
- _await_elicitation now raises ElicitationDeclinedError on decline
instead of returning False; cancel/timeout/malformed still return False
- _hold_native_ask_gate in sessions.py raises on verdict.action=="decline";
both call sites catch it and return abort:True in the policy verdict
- _stable_elicitation_handler in _executor_adapter.py raises on decline
- _executor_adapter.run_turn catches ElicitationDeclinedError, sets
ctx.cancelled (produces response.cancelled, not response.failed), and
returns cleanly — the LLM never sees the denial
Behaviour unchanged for: cancel, timeout, malformed verdict, and the
proxy-MCP path used by native CLI harnesses (Claude Code, Codex).
* fix(tests): catch ElicitationDeclinedError in ask_cycle e2e harness
* fix(review): update docstrings, drop dead store, interrupt session on decline
* fix(policies): use ctx.cancelled for SDK decline abort; drop inert abort field
The SDK invokes the elicitation handler from a separately spawned
control-request task that wraps the callback in try/except Exception,
so raising ElicitationDeclinedError from _stable_elicitation_handler
was swallowed before reaching run_turn's catch block.
Fix: set ctx.cancelled in _stable_elicitation_handler on decline and
return False. The existing run_turn event loop already checks this flag
between events and takes the interrupt+cancel path — no new mechanism
needed for the SDK path.
Keep except ElicitationDeclinedError in run_turn as a fallback for
non-SDK executors that propagate the exception directly.
Also remove the abort:True field from both ElicitationDeclinedError
catch sites in sessions.py — no consumer reads it, so it was inert
and misleading.
* fix(runner): interrupt harness on explicit elicitation decline
When the user explicitly declines an elicitation, the approval event
arrives at the runner with action=='decline'. Previously this just
resolved the pending_approvals Future (unblocking ProxyMcpManager),
which let the deny propagate as a tool error to the LLM — so the agent
continued running.
Fix: after resolving the Future, immediately POST an interrupt event to
the harness before the ProxyMcpManager task resumes (asyncio cooperative
scheduling ensures the interrupt fires first). The interrupt triggers
interrupt_session in the executor, which stops the in-flight LLM turn
before it processes the deny tool result.
* style: ruff format runner/app.py
* fix(sessions): interrupt native harness before returning deny on explicit decline
For native Claude Code, tool-policy ASKs are resolved server-side via
_hold_native_ask_gate. When the user explicitly declines, the server
was returning POLICY_ACTION_DENY to the PreToolUse hook subprocess,
which would let the LLM continue after receiving the tool error.
Fix: await _forward_session_change_to_runner(interrupt) BEFORE
returning the deny response. This sends the Escape key to Claude Code's
tmux pane (via the runner's _handle_claude_native_interrupt) while the
hook deny is still in-flight. By the time the DENY reaches the hook
subprocess, the abort signal is already queued in Claude Code's input,
cancelling the in-flight LLM generation.
* fix(sessions): interrupt codex-native harness on explicit elicitation decline
Same pattern as the claude-native fix: await the interrupt forward to
the runner before returning the decline response to Codex, so the abort
signal arrives before Codex processes the deny and lets the LLM continue.
* fix(sessions): interrupt pi/cursor/hermes/antigravity native on explicit decline
Same pattern as claude-native and codex-native: await interrupt forward
to the runner before returning the decline result so the abort signal
reaches the native harness before it processes the deny.
Covers:
- cursor_permission_request_hook (cursor-native)
- native_permission_request_hook (pi-native, hermes-native)
- antigravity_elicitation_request_hook (antigravity-native)
* fix(repl): send cancel instead of decline on REPL refusal
REPL refusal (typing 'n') should let the LLM continue with the denial
marker rather than aborting the turn. 'decline' triggers the new abort
path; 'cancel' (dismissed without explicit choice) lets the workflow
continue with the DENY tool result so the LLM can adapt.
'decline' is reserved for explicit web-UI Decline button clicks where
abort is the intended behavior.
* fix(test): update repl refusal test for abort behavior; revert repl cancel change
Explicit decline (typing 'n' in REPL or clicking Decline in web UI)
now aborts the turn rather than feeding a denial to the LLM.
Update test_repl_tool_call_refusal_blocks_tool:
- Remove follow_up wait — no second LLM call is made after abort
- Wait for turn to complete (REPL returns to idle)
- Assert raw tool output never appeared in terminal or reached mock LLM
- Drop the 'denied in function_call_output' assertion — turn aborts
before the deny result reaches the LLM
Revert REPL _handle_elicitation change — 'n' keeps sending 'decline'
since it has the same meaning as the web UI Decline button.
* feat(server,web): surface admin + account settings under OIDC/SSO
Under OIDC the SPA rendered no admin or account chrome at all: the
Members/Policies/Account settings sections gated on `accounts_enabled`
and probed admin via the accounts-only `/auth/me`, which 404s under
OIDC. An SSO operator couldn't see who has accounts, manage global
policies, see their own identity, or even sign out.
Root cause was narrow — admin/account chrome keyed on accounts-only
signals. Fix makes them mode-agnostic:
- `GET /v1/me` now returns `is_admin` (shared `users.is_admin` column).
- `PermissionStore.list_users()` (+ SQLAlchemy impl) backs a read-only
`GET /auth/users` on the OIDC router (same shape as accounts).
- Settings nav + pages gate on `/v1/me` (is_admin / login_url), not
`accounts_enabled`. Members runs read-only under OIDC (no password
invite/reset/delete); Policies is fully functional; Account shows
identity + a mode-aware Sign out (OIDC -> GET /auth/logout), with
Change password hidden under OIDC.
Scopes unchanged: session listing stays per-user in every mode; this
adds no new permission level. Per-user session browse and cost
attribution are intentionally out of scope (tracked separately).
Co-authored-by: Isaac
* test(server): OIDC integration coverage for /v1/policies gating
The default-policies routes gate on the mode-agnostic
permission_store.is_admin, so they already worked under OIDC — this
pins it end-to-end via create_app wired with an OIDC provider: an admin
can CRUD global policies, an unauthenticated caller gets 401, and a
non-admin can read but not write/delete (403).
Co-authored-by: Isaac
* fix(server): sync openapi.json + /v1/me test for is_admin field
CI caught two artifacts of adding is_admin to GET /v1/me:
- Regenerate openapi.json (scripts/dump_openapi.py) so the drift check
passes — only the /v1/me description/return docs changed.
- Update test_me_header_mode_behaviors to expect is_admin=False across
the missing / valid / reserved-name header-mode cases.
Co-authored-by: Isaac
* fix(server): align /v1/me is_admin with the auth-route admin check
Polly review flagged that /v1/me computed is_admin from
permission_store.is_admin() alone, while /auth/users and /auth/invite
gate on permission_store.is_admin(caller) OR admin_list.is_admin(caller).
An identity added to the admin-list file but not yet promoted (the DB
flag flips at next login via promote_if_listed) would be authorized by
those routes yet see no admin chrome in the SPA.
Build admin_list once near app creation and consult it in /v1/me too, so
the chrome signal never under-reports relative to server enforcement.
Adds a regression test (admin-list identity, non-admin DB row ->
is_admin true).
Co-authored-by: Isaac
* fix(changelog): detect the draft release with the App token, edit by id
A manual run against a real draft release still skipped "Enrich the release
draft body". Two causes, both about drafts being invisible/unaddressable the
way we probed:
- The guard probed `gh release view <tag>` with the read-only GITHUB_TOKEN,
but GitHub hides DRAFT releases from tokens without push access — so the
probe always came back empty and is_draft was wrongly false.
- Even with a capable token, the get/edit-by-tag REST endpoint 404s on a draft
(its tag isn't "real" until published), so editing by tag would fail too.
Move draft detection to a new "Resolve draft release" step that runs after the
App token is minted (which has push access), matching by tag_name over the
release list (the only way to see a draft), and expose the numeric release_id.
Enrich now PATCHes the release by id instead of by tag. The read-only guard no
longer probes for the draft, and the "Resolve draft release" step emits the
"no draft found" notice itself, replacing the old note-skipped step.
No behavior change on the happy auto-path; this makes the draft-body
enrichment actually fire (incl. for still-untagged drafts and manual dispatch).
* fix(changelog): pass TAG to jq via env, not string interpolation
Polly review flagged jq-program injection: TAG was interpolated into the
--jq filter (`.tag_name == "${TAG}"`), so a tag containing `"` or jq syntax
could alter which release is selected — and this runs after the contents:write
App token is minted. Read it via jq's `env.TAG` instead, which treats the value
as data. (gh api's built-in --jq has no --arg, and --arg is a standalone-jq
flag gh api rejects, so env is the fix that actually works here.)
Verified adversarially: a tag like `v"; .draft` now yields an empty match and
exit 0 instead of a malformed/altered filter.
* feat(file-viewer): comment on rendered HTML files
Reviewers can now highlight text in the rendered HTML preview and attach
review comments — parity with the Markdown (TipTap) and code (Monaco/Shiki)
comment surfaces. Previously HTML opened in a sandboxed preview iframe with no
way to comment.
The preview iframe stays sandboxed without `allow-same-origin`, so the parent
can't read its selection directly. A nonce-guarded bridge script injected into
the iframe relays selections over a private MessageChannel and paints
highlights (CSS Custom Highlight API) inside the frame. Comments store
raw-HTML-source offsets + anchor_content (resolved parent-side), so the agent
and classifyAndRemapComments keep working unchanged. No backend changes — the
comment store/API are already file-type agnostic.
- htmlCommentBridge.ts: injected bridge script, message protocol + validation,
rendered-selection -> source-offset resolution
- HtmlCommentViewer.tsx: iframe owner, channel handshake, floating button
- CodeViewer.tsx: route HTML preview to HtmlCommentViewer
- unit/component + Playwright e2e coverage
Co-authored-by: Isaac
* fix(ap-web): avoid RegExp.exec false positive in security exfil scan
The CI exfil scanner treats `.exec(` as dynamic code execution; use
`String.match` for the whitespace-tolerant anchor lookup instead.
* fix(file-viewer): correct HTML-preview comment highlighting and navigation
Fixes several issues in the rendered-HTML comment surface found while
reviewing the feature:
- Multi-line anchors never highlighted: the in-frame matcher used exact
indexOf on raw text-node data (which preserves source newlines) while
anchor_content has collapsed whitespace. Made it whitespace-tolerant,
mirroring the parent's findAnchorInSource.
- Dragging the right panel over the preview iframe stuck to the cursor:
mousemove/mouseup fell into the sandboxed frame so the parent never saw
the release. Added a transparent drag overlay in the inline-panel and
comments-panel resize hooks.
- Just-saved highlight stayed grey: the leftover native selection painted
over the Custom Highlight. Clear it once a saved comment covers it.
- Clicking a comment didn't scroll the frame to its highlight; now it does
(only when off-screen).
- Repeated anchor text (e.g. a title reused in the body) highlighted every
copy and resolved selections to the first match. Both directions are now
occurrence-aware: the bridge reports which occurrence was selected and the
parent stores/paints only that one.
- Selecting a highlighted range now activates its comment and scrolls the
comments panel to that card (switching tabs when needed).
Adds unit coverage for the resize-overlay, occurrence resolution, and
panel-reveal logic, plus Playwright e2e cases for each behavior.
Co-authored-by: Isaac
---------
Co-authored-by: Yu Gong <yu.gong@databricks.com>
Co-authored-by: dbczumar <corey.zumar@databricks.com>
Co-authored-by: Serena Ruan <serena.rxy@gmail.com>
Delivers the payoff of the full-server transport, live-verified.
Ad-hoc request-level function tools do not round-trip on the full-server
path (the SDK harnesses handle tools internally, so a client-declared
function tool never surfaces as a server-dispatched, policy-gated call and
the turn hangs). Instead the driver drives a read-only builtin (list_files)
that the server actually dispatches and gates at the tool_call phase.
- FullServerDriver registers the agent with tools.builtins=[list_files]
(spec_version bundle, config.yaml member, spec-format executor).
- tool_probe_turn(deny): ALLOW runs against the base session; DENY runs
against a lazily-created second agent/session whose spec bakes a
tool_call deny policy (the REST policy endpoint's handler allowlist
excludes make_fixed_action_callable, so the deny rides in the spec).
Populates tool_calls and tool_call_denied from the session snapshot.
- Gated live test asserts ALLOW dispatches list_files and DENY blocks it.
Verified on oss: ALLOW dispatches the builtin; DENY yields
function_call_output {"error": "Denied by policy: bench-policy-deny"}.
Follow-ups: SSE streaming, interrupt, and the --transport bench wiring.
Two fixes surfaced from a v0.4.0dev0 tag push:
1. github-release.yml failed with HTTP 422 "body is too long (maximum is
125000 characters)": --generate-notes asked GitHub to list every PR since
the previous tag (193 for the v0.3.0→HEAD range), overflowing the release-
body cap. We draft our own curated notes in draft-release-notes.yml, so
--generate-notes is dead weight. Replace it with a short placeholder body
that draft-release-notes.yml overwrites; the 422 failure mode is gone.
2. A manual run for a dev tag (v0.4.0dev0 --base v0.3.0) harvested 6 PRs but
reported "CHANGELOG.md already up to date" — generate.py gated the write on
a strict ^v\d+\.\d+\.\d+$ regex that a .dev0 tag fails, so it silently
skipped the write. Order CHANGELOG.md by PEP 440 (packaging.Version) using
the full tag string as the block header, so dev/rc tags land in their own
correctly-ordered blocks (v0.4.0 > v0.4.0rc1 > v0.4.0.dev0 > v0.3.0) and
coexist with the eventual final rather than collapsing into it. Re-running a
tag still replaces its own block (idempotent).
previous_final_tag stays finals-only (a real v0.4.0 still diffs against v0.3.0,
not an intervening rc). The workflow_run auto-trigger is unchanged and remains
finals-only — dev/rc changelog blocks are reachable only by manual dispatch.
The harvest step installs packaging (it runs bare python3 before uv sync), and
the dry_run input description is trimmed.
87 tests pass; verified end-to-end that v0.4.0dev0 --base v0.3.0 now writes a
correctly-ordered block instead of no-op'ing.
Co-authored-by: Isaac
## Related issue
N/A
## Summary
- Adds a dynamic harness registry backed by the `omnigent.community.harnesses` entry point group, with built-in and community contributions merged through `HarnessContribution`.
- Adds import-safe harness install metadata and community namespace anchors so optional harness packages can contribute modules under `omnigent.community.harnesses.*` without importing onboarding/provider stacks during discovery.
- Wires aliases, native-agent metadata, model override env vars, runtime harness modules, setup/readiness checks, process-manager errors, and runner spawn env builders through the registry.
- Adds a `/v1/harnesses` catalog route and updates the web UI to merge server-provided harness labels into the picker surfaces.
- Documents the plugin interface and adds registry tests for merge behavior, import-path validation, built-in collision rejection, and external namespace imports.
## Test Plan
- `PYTHONPATH=. uv run --with pytest pytest tests/test_harness_plugins.py tests/test_harness_aliases.py tests/test_model_override.py tests/onboarding/test_harness_readiness.py`
## Type of change
- [ ] Bug fix
- [x] Feature
- [x] Refactor / chore
- [x] Docs
- [x] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
The focused pytest suite passed with 187 tests. I also verified this facilities commit has no provider-specific harness extraction references; concrete harness extraction belongs in a later commit.
Drop the monotonic transition constraint from LabelDef and all
associated infrastructure. Label writes are now validated against
the declared values enum only; free transitions between declared
values are permitted.
Removes _monotonic_ok, _merge_monotonic_writes, and the monotonic
branch in _filter_schema_valid from the policy engine. Cleans up
the omnigent adapter's _OMNI_TO_AP_MONOTONIC mapping and the
loader's monotonic aliasing. Updates all YAML fixtures, parser
tests, and integration tests accordingly.
Manual runs of draft-release-notes.yml were unusable for testing: the guard
only proceeded for a final vX.Y.Z tag (a dispatch with tag=ci-test skipped
every step), and generate.py itself requires a version tag to compute the
range and order CHANGELOG.md.
Add a preview path for workflow_dispatch:
- generate.py gains --base <ref> to override the range start (base..tag,
any refs), plus a clear CLI error when --tag isn't a final vX.Y.Z and no
--base is given. Version-only CHANGELOG.md insertion is skipped for a
non-version tag.
- The workflow gains `base` and `dry_run` (auto|true|false) dispatch inputs.
The guard proceeds for a version tag OR a base override; dry_run defaults to
auto → preview for a non-version tag or base override, real run otherwise,
and is force-overridable. Dry-run renders the CHANGELOG section + draft notes
to the run summary and skips the token mint, CHANGELOG PR, and release-body
edit. The workflow_run (real release) path is unchanged.
Also harden changelog_description: bare omit markers (skip / n/a / none / -,
left over from the old template sentinel) now count as an absent section
instead of leaking in as a literal entry — caught while dry-running against
real history (a merged PR still said "skip").
84 tests pass; verified end-to-end with a local --base dry-run over real
repo history.
Co-authored-by: Isaac
* Otto eyes: look at the caret while typing, the mouse while pointing
Otto's pupils on the new-chat landing tracked only the mouse pointer. The
composer sits directly below the mascot, so while the user types their
attention is on the caret, not the mouse.
Otto now looks at whatever the user last moved: the mouse pointer, or — while a
text field (textarea, text input, or contenteditable) is focused — its text
caret. Moving the mouse pulls his gaze to the pointer even while a field is
focused; a genuine caret move (typing, paste/delete, arrow/Home/End navigation,
click-to-reposition) pulls it back. On mount the pupils rest centered; focus
alone (including the composer's autofocus) never moves them — tracking begins
on the first real activity.
Form fields have no native caret-rect API, so the caret is measured with a
hidden mirror div that wraps identically to the field: its font is copied via
the `font` shorthand (copying individual longhands lets an inherited
font-stretch/variation widen the text and wrap it a word early, which made Otto
glance a line too low), and it uses box-sizing:content-box with
width = clientWidth - horizontal padding (getComputedStyle width is the
content-box value, so copying it onto a border-box element shrank the mirror).
A DOM Range over the character before the caret gives its real position on the
correct line at any width. contenteditable uses the collapsed selection rect.
Only the direction to the target matters — the pupil is normalized onto the eye
rim — so sub-pixel differences are invisible; the existing 90ms transform
transition smooths every hand-off.
Adds a colocated Vitest for the last-activity model (centered on mount,
pointer/caret trade-off, focus alone inert) and a Playwright e2e_ui test
driving the real landing hero.
Signed-off-by: OGordon100 <35759308+OGordon100@users.noreply.github.com>
Co-authored-by: Isaac
* harden(otto-eyes): always clean up caret mirror; drop detached field
Wrap the caret-measurement mirror <div> in try/finally so it's always
removed from <body>, even if a Range measurement throws — otherwise a
persistently-throwing frame would leak one hidden div per rAF and kill
tracking. Also drop activeField back to the pointer when it's no longer
connected (React can unmount a focused field without a matching
focusout), so Otto rests centered instead of aiming at (0,0).
Remove the layout-dependent e2e_ui mascot test; the unit suite in
OttoEyes.test.tsx covers the pointer/caret hand-off.
Co-authored-by: Isaac
---------
Signed-off-by: OGordon100 <35759308+OGordon100@users.noreply.github.com>
Co-authored-by: Serena Ruan <serena.rxy@gmail.com>
Pinning a session while the sidebar's Pinned section is collapsed left
the freshly-pinned chat hidden inside the collapsed group, making it look
like the pin never took. Watch pinnedConversationIds for a newly-added id
and drop "Pinned" from the collapsed set (persisted), so the section pops
open and the just-pinned session is immediately visible. Only reacts to
pins being added — unpinning or reordering leaves the collapse preference
untouched.
Co-authored-by: Isaac
_normalize_cursor_usage copied cursor's inputTokens straight into
input_tokens and also mapped cacheReadTokens/cacheWriteTokens into the
cache buckets without subtracting. cursor's inputTokens is inclusive of
cache read + write (documented in cursor_native_usage.py), and
compute_llm_cost requires input_tokens to be the non-cached portion (it
prices the cache buckets additively). The SDK path is priced via
compute_llm_cost and emits no direct cost_usd, so cached tokens were
billed twice: once at the full input rate, once at their cache rate.
Subtract the mapped cache buckets from input_tokens (clamped at 0),
mirroring the qwen and antigravity executors. No existing test locked the
pre-fix value; strengthen the cache test to assert the non-cached input
and add focused subtraction/clamp regression tests.
Closes#1801
Signed-off-by: abhay-codes07 <abhaysingh0293@gmail.com>
* fix(runner): delete native-harness bridge dirs on session delete
Each native session's prepare_bridge_dir creates a per-conversation dir
holding a bridge token + MCP config (secret material). delete_session
closed the pane but never removed this separate dir, so token-bearing
/tmp/omnigent-* dirs accumulated even on a clean delete (#1350).
Resolve the bridge dir for every native harness (claude/codex/cursor/pi)
and rmtree it after the pane is released. Bridge ids can be rotated via a
session label, so resolve those too and fall back to session_id; we don't
know which harness the session used, so delete every candidate dir with
ignore_errors making wrong-harness / already-gone a no-op. Codex's private
CODEX_HOME lives inside the bridge dir, so it goes with it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: CM <chandrameenamohan@gmail.com>
* fix(runner): clean bridge dirs on the real delete path (/resources)
Polly review found the #1350 cleanup was wired only into the bare
DELETE /v1/sessions/{id} runner route, which production never calls —
server delete_session drives DELETE /v1/sessions/{id}/resources
(cleanup_session_resources), so the token-bearing bridge dir still
leaked on real deletes and the original test passed only because it hit
the unused route directly.
Call _delete_native_bridge_dirs from cleanup_session_resources too (the
server-driven path). Deliberately NOT inside resource_registry.cleanup_session,
since the agent-switch reset (reset_session_state) reuses it while the
session and its bridge live on. Add a regression test through
DELETE .../resources that fails before this change and passes after.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: CM <chandrameenamohan@gmail.com>
* fix(runner): clean up bridge dirs for all 11 native harness families (#1350)
_delete_native_bridge_dirs only removed bridge dirs for 5 families
(claude/codex/cursor/opencode/pi). The other 6 native harnesses
(antigravity/goose/hermes/kimi/kiro/qwen) also leave token-bearing bridge
dirs that leak on session delete. Extend cleanup to cover all 11; resolve
antigravity's rotated bridge-id label like claude/codex/opencode. Also log
non-FileNotFound rmtree failures at debug instead of silently swallowing.
Extend the regression test to parametrize over all 11 families via the real
DELETE /v1/sessions/{id}/resources path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(lint): apply ruff format and import ordering fixes
---------
Signed-off-by: CM <chandrameenamohan@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tomu Hirata <tomu.hirata@gmail.com>
Add the `text-destructive` token to the archived session delete button's
trash icon so it reads as a destructive action, consistent with the
Delete button in the confirmation dialog.
Co-authored-by: Isaac
* feat(changelog): free-text entries tagged by Type of change
Rework the PR `## Changelog` section based on review feedback:
- Drop the `Category: description` format. The changelog tag is now derived
from the "Type of change" checkboxes instead (e.g. checking "UI / frontend
change" renders `[UI] <description>`), so authors write a plain user-voice
one-liner and never restate the category.
- Multi-line entries no longer fail the gate — the harvester takes the first
non-blank line as the description.
- The section is optional: authors delete it (or leave the placeholder) when the
change isn't noteworthy, and the PR is simply omitted from the changelog. No
author-grouped "undocumented" bucket — for large ranges it's just noise. The
one hard rule kept: a Breaking change must carry a real description.
- Replace the `skip` sentinel in the template with
`<Add a line to describe the change, else delete this section>` and update the
guidance comment accordingly.
CHANGELOG.md entries render as a flat, PR-sorted list of `- [Tag] description
(#NNNN)`; the release-notes draft buckets Feature/UI into "Major new features"
and Bug fix/Breaking into "Bug fixes & hardening". The shared `_md.py` parser
(now `changelog_description` + `checked_labels` + `type_tag`/`TYPE_TAGS`) backs
both the gate and the harvester so they can't drift. 75 tests pass.
Co-authored-by: Isaac
* style(changelog): use backticks for `Type of change` in preamble
ruff format normalizes the escaped-double-quote seed string to single
quotes; sidestep the version-dependent quote nit by wrapping "Type of
change" in backticks (also more consistent with the surrounding markdown
in that preamble). No behavior change.
Co-authored-by: Isaac
* perf(web): cut UI bundle ~32% by deduping shiki and dropping dead deps
The web bundle shipped three copies of shiki: root shiki@4.2 (chat +
Monaco), and shiki@3.23 pulled transitively via @streamdown/code and
@pierre/diffs. The version gap blocked npm from deduping, so ~300
duplicate language-grammar chunks (cpp, wasm, etc. — some ~620 KB each)
shipped twice.
- Add a `shiki`/`@shikijs/*` overrides block pinning the family to 4.x
so @streamdown/code resolves the single root shiki. Verified the chat
and streamdown highlighter paths still render.
- Delete the unreachable ai-elements island (43 files) + ui/carousel;
only code-block, conversation, message, reasoning, shimmer, and
streamdown-security are reachable.
- Drop dependencies with no live import: @lobehub/ui,
@databricks/sdk-experimental, motion, @xyflow/react,
@rive-app/react-webgl2, media-chrome, embla-carousel-react,
react-jsx-parser. Move the type-only `ai` package to devDependencies.
- Import the lobehub harness icons via their Mono subpath (as KimiIcon
already did) so the barrel's antd-pulling statics stay out of the
bundle.
- Fix two files that relied on a global JSX namespace leaked by a
removed transitive @types/react@18; use ReactElement instead.
Standalone build: 28.01 MB -> 18.92 MB (-32.5%), 712 -> 411 files.
Type-check, lint, and the full vitest suite (3418 tests) pass.
Co-authored-by: Isaac
* chore(oss): regenerate public lockfiles against public PyPI/npm
---------
Co-authored-by: omnigent-ci[bot] <294685417+omnigent-ci[bot]@users.noreply.github.com>
Make the "Steps to reproduce" field mandatory on the bug report form and
turn off blank issues so reporters can't bypass the structured form. This
raises the floor on bug report quality and cuts low-effort/AI-slop reports.
The field description offers an escape hatch for genuinely intermittent bugs.
Co-authored-by: Isaac
Add a VS Code extension under editors/vscode/ that opens the running local
Omnigent server in an editor-beside webview iframe. It is a thin client of the
local server (localhost discovery via ~/.omnigent/local_server.pid + /health),
contributing an activity-bar icon (omnigent.home view + viewsWelcome), an
editor-title icon, and the omnigent.open command.
Scope is intentionally minimal per the issue: iframe render only. Embed/SPA,
sessions, diffs+SSE, send-selection, the /v1 client, token auth, and remote
servers are out of scope for this first donation.
- esbuild bundle -> dist/extension.js; vitest unit tests (55) for the pure
modules (csp, iframeHtml, host, discovery, config, controller)
- 3-directive host CSP (default-src 'none'; style-src 'nonce'; frame-src origin);
no token ever placed in the iframe URL
- CI deferred to a maintainer-owned follow-up per issue Q5; the proposed
path-filtered, security-gated workflow (mirroring ap-web-tests.yml) is in the
PR description so it does not trip the untrusted-PR workflow guard
- Apache-2.0; DCO sign-off
Refs: #1219
Signed-off-by: Tanner Wendland <tanner.wendland@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(native): bound the permission-hook reattach spin-loop (#1782)
`_post_hook_with_reattach` re-POSTs a permission/ask elicitation with a stable
`_omnigent_elicitation_id` so a proxy-severed long-poll re-attaches instead of
prompting the human twice. But its retry deadline was `_PERMISSION_TIMEOUT_S`
(one day) — the same value that (correctly) bounds a single long-poll. So
against a persistently sick or unreachable server the loop re-POSTed every
<=30s for 24h. Each re-POST re-drives the turn and respawns the harness/tool
subprocesses (node/npm/chromium/tmux/python), which — with the host not reaping
orphans (#1782 Bug A) — piled up as zombies overnight. This is the spin that
produced the repeated same-`elicitation_id` log lines and `Omnigent API failed:
request error`.
Bound CONSECUTIVE FAST failures instead of wall-clock:
- A failure that returns before half the read budget means the server did not
hold the poll (sick / unreachable) — it counts toward
`_PERMISSION_MAX_CONSECUTIVE_FAILURES` (default 8, `OMNIGENT_HOOK_MAX_RETRIES`).
- A failure that surfaced only after the poll was held open a long time (a slow
human, the server working as intended) resets the counter — so raising a
legitimate approval prompt and waiting on it is completely unaffected.
The happy path (2xx on first try) and 4xx-is-final behavior are unchanged; a
regression test locks in that success returns without retry.
Pairs with the host orphan-reaper fix; either alone mitigates #1782, both
together close it.
Co-authored-by: Isaac
* fix(native): classify reattach failures by kind, not wall-clock (#1782 review)
Polly AI review caught a real regression in the first spin-loop fix. It bounded
CONSECUTIVE FAST failures where "fast" = returned in under 12h
(_PERMISSION_TIMEOUT_S * 0.5). But this hook exists precisely for deployments
where "proxies sever idle long-polls" — and a proxy severs a legitimately-
PARKED poll (a human thinking) in seconds-to-minutes, always << 12h. So every
such sever was miscounted as a fast failure, and a real human approval behind a
severing proxy was fail-asked after ~8 severs (~8 min) — contradicting the PR's
own "a slow human is never capped" claim.
Root cause: elapsed wall-clock can't tell a 60s proxy-severed *parked* poll from
a 60s connect failure. Fix: classify by HOW the request failed.
- Hard failure (counts toward the cap = the #1782 spin): a 5xx, a connection
that never established (_NEVER_CONNECTED_ERRORS: ConnectError/ConnectTimeout/
PoolTimeout/ProxyError), or an established connection that dropped in under
_PERMISSION_HELD_POLL_FLOOR_S (10s — a flapping/crash-looping server).
- Held-poll sever (resets the counter): an established connection dropped
mid-poll after being held >= the floor. That is the re-park mechanism working
as intended, so a slow human is never capped no matter how often the proxy
severs.
Also: restore an absolute _PERMISSION_TIMEOUT_S (1-day) backstop on total wait,
and harden the env parse (_env_int ignores a malformed OMNIGENT_HOOK_MAX_RETRIES
instead of crashing the hook at import — another review note).
Tests rewritten to drive by exception kind: down-server and 5xx bound at the
cap; an instant establish-drop flap is bounded; and the key regression —
a proxy severing a held poll every ~60s, 3x the cap, never caps and the human's
eventual 2xx returns. Verified before/after: old 12h logic caps at 8 severs
(~8 min); new logic never caps a held-poll sever.
Co-authored-by: Isaac
* test(native): bound + document the held-sever reset path (#1782 review)
Adversarial review flagged a residual in the kind-based classifier: a *sick*
backend behind a proxy/LB that accepts then silently severs a held connection
(>= the 10s floor) raises RemoteProtocolError — transport-indistinguishable
from a proxy severing a genuinely-parked human poll. Both reset the
consecutive-hard-failure counter, so that case is NOT caught by the cap.
This is fundamental, not fixable client-side: the server holds the POST
silently with no "parked" ack, so "server is waiting for a human" and "proxy
dropped a dead backend" look identical after N seconds. Capping it sooner would
necessarily cap a real slow human on the same topology — so the absolute
_PERMISSION_TIMEOUT_S (1-day) deadline is the tightest safe bound. Blast radius
is limited: this loop only re-POSTs over HTTP from one hook process (it does
not itself respawn subprocesses), and the host orphan reaper (Bug A) reclaims
any subprocesses a re-driven turn spawns — so the worst case is one hook
slow-retrying for a day, not the original zombie pileup.
No behavior change. This commit:
- documents the residual honestly in the docstring (stops implying "a sick
server is always capped"), and
- adds test_reattach_never_resolving_severs_are_bounded_by_deadline, which
proves the previously-untested reset-forever path terminates via the
deadline (returns None, finite call count ~= budget/held) rather than
looping forever.
Co-authored-by: Isaac
* feat(native): make the held-poll floor env-tunable (#1782 review)
Polly non-blocking note: _PERMISSION_HELD_POLL_FLOOR_S (the sole flap-vs-held
discriminator) was hardcoded at 10s. Behind an unusually aggressive proxy/LB
whose idle timeout is under 10s, a legitimate slow-human sever would be
classified as a flap (hard failure) and a real approval could be fail-asked
after the cap — the narrow residual human-capping edge. The retry cap is
already env-tunable; the floor was not.
Make it overridable via OMNIGENT_HOOK_HELD_POLL_FLOOR_S (new _env_float helper,
same fault-tolerant fallback as _env_int; floored at 0 so a negative can't
disable flap detection). Default 10s unchanged. Test covers the override and
the malformed-value fallback.
Co-authored-by: Isaac
* fix(native): reject non-finite held-poll-floor override (#1782 review)
Polly non-blocking note: _env_float accepted inf/nan (float("inf"/"nan") does
not raise ValueError). An inf OMNIGENT_HOOK_HELD_POLL_FLOOR_S would classify
every sever as a held poll — silently disabling flap detection — and nan makes
every `held_s < floor` comparison False. Add a math.isfinite guard so both fall
back to the 10s default like any other malformed value. Test covers inf/nan/-inf.
Co-authored-by: Isaac
* fix(host): reap orphaned harness/tool subprocesses to stop zombie pileup (#1782)
When a runner dies, the harness tool subprocesses it spawned detached
(node/npm/chromium/tmux/python — start_new_session=True) are orphaned and
reparented to `omnigent host`, which is PID 1 in a container (or, with this
change, a child subreaper otherwise). The host installed no child reaper and
only wait()s the runners it tracks directly, so every orphan became a
permanent <defunct> zombie. A run blocked overnight on an unanswered approval
elicitation accumulated ~900 zombies / ~2,300 PIDs / ~6 GB RSS and OOM'd the
shared box.
Install PR_SET_CHILD_SUBREAPER at host startup (Linux; harmless no-op when
already PID 1 or non-Linux) and run a periodic sweep that reaps ready orphans
without disturbing tracked-runner exit accounting:
- Linux/POSIX: os.waitid(..., WNOWAIT) peeks at the next reapable child
without consuming it; a tracked runner is left for its Popen reaper
(_watch_runner) so its real exit code still reaches host.runner_exited.
- Platforms without os.waitid (macOS): waitpid(WNOHANG) reaps, and re-injects
a tracked runner's status onto its Popen so exit-code fidelity is preserved.
A blind waitpid(-1) reaper would steal a just-crashed runner's status and make
Popen.poll() report a bogus exit 0 — verified and guarded against by
test_reap_orphans_never_steals_tracked_runner_exit_code.
This is the containment half of #1782 (stops the box from going down); the
spin-loop that drives the fast spawning is addressed separately.
Co-authored-by: Isaac
* fix(host): pause orphan reaper during host-owned git subprocesses (#1782)
Polly AI review caught a real race in the orphan reaper. Its contract was
"any reapable child not in self._runners is an orphan → reap it", but the host
spawns other DIRECT children besides runners: the git commands in
git_worktree._run_git (subprocess.run, no start_new_session), invoked from the
worktree handlers via asyncio.to_thread. Those git children aren't tracked
runners, so they were indistinguishable from orphans to the reaper.
The race: git exits and becomes reapable; before subprocess.run's own wait()
(in the worker thread) collects it, the 2s reaper sweep fires and waitpid()s
it; subprocess.run then hits ECHILD, which CPython swallows and reports as
returncode 0 — so a FAILED `git worktree add/remove/branch -D` is silently
treated as success (create_worktree/remove_worktree branch on returncode != 0).
Fix: a _host_subprocess_op() context manager increments an
_owned_subprocess_ops counter; _reap_orphans_once() is a no-op while it is >0.
The two worktree to_thread calls are wrapped in it. Counter mutation and the
reaper both run on the event loop, so a plain int needs no lock; the decrement
is in finally so a raising git op can't wedge the reaper off. This also covers
the shutdown `finally: _reap_orphans_once()` path if a worktree op is in flight.
Note: spawning git with start_new_session would NOT fix this — setsid changes
the session/group, not parentage, so the child stays reapable by waitpid(-1)/
P_ALL. Pausing the reaper is the correct scope.
Tests: a git-race regression (failed `sh -c 'exit 42'` stand-in keeps its true
exit code while an op is in flight) and a re-entrancy/exception-balance test.
Verified before/after: without the guard the reaper steals the child and the
owner reads returncode 0; with it, 42 survives.
Co-authored-by: Isaac
* feat(android): native Android WebView shell (#1604)
Add a thin native Android shell that loads the server-served web UI, the
third native runtime of the same bundle alongside the iOS WKWebView shell
(web/ios) and the Electron desktop shell. Mirrors the iOS shell's
native<->web contract so the SPA needs no per-feature branching.
Web side (one bundle, multiple runtimes):
- nativeBridge.ts: add "android" to the shell `kind` union AND the
nativeApi() runtime guard (the guard, not just the type, is what makes
the bridge live), plus an isAndroidShell() sibling to isIOSShell().
- index.css: fold Android-measured insets into --omnigent-safe-* via
max(env(...), var(--omnigent-android-safe-area-*, 0px)), universally —
no isAndroidShell() branching; zero effect off the Android shell.
Android module (web/android, Kotlin):
- Web->native bridge via WebViewCompat.addWebMessageListener,
origin-allowlisted to the pinned server + main-frame gated — the
structural equivalent of the iOS isMainFrame/frame-origin check, so a
sandboxed agent-HTML iframe can't reach the native surface.
- OS notifications with tap routing (cold + warm start, consume-once
replay cache), best-effort badge, POST_NOTIFICATIONS runtime request.
- Edge-to-edge insets measured natively and pushed to CSS (Android
WebView can't rely on env(safe-area-inset-*) alone).
- File upload (WebChromeClient.onShowFileChooser) and microphone
(onPermissionRequest, granted to the pinned origin only + RECORD_AUDIO).
- Downloads incl. blob:/data: exports via a fetch->base64->MediaStore
bridge, which closes#969 (the iOS shell drops these).
- Native connect / recent-servers screen; system-back + predictive-back.
Builds clean: gradlew :app:assembleDebug :app:lintDebug = BUILD
SUCCESSFUL, 0 lint errors (JDK 17, Gradle 8.9, compileSdk 35, minSdk 28).
Not yet exercised on a device. Sidebar edge-swipe and the native floating
bars are deliberately deferred to the web in-page fallbacks (see README).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): keep the OIDC redirect chain in the WebView (#1708)
The shell handed any off-origin top-level navigation to the external
browser (a fail-closed choice from the bridge-hardening work). That
kicked the OIDC login redirect (the server bouncing the main frame to
the IdP) out to Chrome, where auth completed and the session cookie
landed — so the in-app WebView never received the session and login
silently failed.
shouldOverrideUrlLoading now lets all http/https navigation, including
the off-origin OIDC redirect chain, load in the WebView — mirroring the
iOS shell. Only top-level non-http(s) schemes (mailto/tel/intent/custom)
are still handed to the system. This is safe because the native bridge
is origin-allowlisted (addWebMessageListener) and the window.omnigentNative
facade is injected only on the pinned origin, so a foreign auth page
loaded top-level can't reach native.
Verified on a Pixel-6 emulator (API 34) against a live OIDC deployment:
before, logcat showed an ACTION_VIEW handoff of auth.joyful.house to
com.android.chrome and Chrome took the foreground; after, the IdP
(Authentik) login page renders inside the app and login completes the
round-trip in the WebView.
Does NOT cover an IdP that federates to Google social login — Google
blocks embedded WebViews (disallowed_useragent), which needs a Custom
Tabs hand-off with a session hand-back. Tracked in #1708.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(android): brand the app icon + Connect screen to match iOS
The app icon was a generic placeholder and the Connect screen was bare
Material chrome — neither matched the iOS shell or the Omnigent brand.
- App icon: replace the placeholder with the Omnigent starfish (converted
from the shared platform-assets brand source — the same favicon/iOS
AppIcon mark) as the adaptive foreground, a starfish-silhouette
monochrome layer for themed icons, on the brand dark-navy background.
- Connect screen: mirror the iOS ConnectView — the omnigents wordmark
(which embeds the starfish) on top, a muted subtitle, a "Server URL"
label, a bordered field, a filled dark primary button, an inline error
line, and bordered recent-server rows.
- Brand colors: port the iOS DesignTokens palette (foreground #11171C,
border #E8ECF0, primary #11171C, muted, error) into colors.xml plus a
values-night/ dark variant. Type uses the system font (Roboto) — the
same native-font choice the web UI and iOS make (--font-sans is a
system stack), so the setup screen reads consistently across platforms.
Built + screenshot-verified on a Pixel-6 emulator: the wordmark, colors,
field, and button render at parity with the iOS setup screen.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(android): authenticate via Chrome Custom Tabs (fixes Google + passkey) (#1708)
Per RFC 8252, native apps must not run OAuth in an embedded WebView — Google
blocks it (disallowed_useragent) and passkeys/WebAuthn don't work there. The
Layer-1 stopgap (load the IdP in the WebView) only worked for IdP-native
username/password. This does it correctly: authenticate in a Chrome Custom Tab.
Flow (reuses the server's existing browser-login endpoints — the same ones the
`omnigent login` CLI uses, no server change):
- OmnigentWebViewClient intercepts the off-origin OIDC redirect (a server
redirect — no user gesture — to the IdP) and triggers native login instead of
ever loading the IdP in the WebView. A gesture'd off-origin nav is treated as
an external link and handed to the system browser.
- OidcLoginManager: POST /auth/cli-login -> {ticket, login_url}; open login_url
in a Custom Tab (Google/passkey/any IdP all work in a real browser); poll
GET /auth/cli-poll?ticket until it returns the session JWT.
- The Custom Tab and the WebView have isolated cookie stores, so the session is
bridged explicitly: the polled JWT is exactly the session-cookie value (the
server validates the same HS256 JWT as cookie or Bearer), so MainActivity
injects it as the __Host-ap_session cookie via CookieManager and reloads
authenticated, then brings itself back over the Custom Tab.
Verified against the live OIDC server on an emulator: connect -> the shell
intercepts the redirect, POSTs cli-login, opens the Custom Tab to the login URL,
and polls cli-poll (202 pending) — the IdP never loads in the WebView. The login
round-trip (token -> cookie -> authenticated reload) needs a real device with a
set-up browser to complete; pending on-device confirmation.
Adds androidx.browser (Custom Tabs). Follow-up #1708. The `cli-` endpoint naming
is now a misnomer for shared CLI+mobile use — proposed to maintainers to alias,
deferred for blast radius.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): use the system browser for login + return-to-app bridge (#1708)
Verified on-device: the in-app Custom Tab rendered the IdP (Authentik) flow
page blank, while the full system browser works. Switch the login hand-off from
a Custom Tab to a plain ACTION_VIEW browser intent — still RFC 8252 compliant
(the system browser is the canonical external user-agent; Google, passkeys, and
password managers all work). Drops the androidx.browser dependency.
Return-to-app: the poll completes while the browser is foreground, and Android's
background-activity-launch rules block us from foregrounding ourselves, so we
both attempt a reorder-to-front (works within the grace period) and post a
"Signed in — tap to return" notification as the reliable path back.
End-to-end verified against the live OIDC server: login -> session JWT polled ->
injected as __Host-ap_session -> WebView reload is authenticated (server: GET /
304, WebSocket /v1/sessions/updates accepted, /v1/sessions 200), and the app
returns to the foreground. Fully seamless auto-return (browser auto-closing on a
custom-scheme redirect) needs a small server change — tracked in #1708.
Auth-flow logging redacts URLs (OAuth state/PKCE/ticket) — logs origins only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): apply the safe-area insets so mobile chrome isn't under the status bar
The header's top-left sidebar toggle (and the sidebar/panels) were untappable on
Android: the WebView is edge-to-edge and the OS status bar (128px on the test
device) overlaps `.chat-header` (which is `absolute top-0`), so the system
swallows the tap. Root cause: every safe-area rule in index.css was gated on
`[data-ios-native]`, and several used raw `env(safe-area-inset-top)` — which is 0
in Android WebView. The native side already injects the real inset via
`--omnigent-android-safe-area-*`; the web side just never consumed it on Android.
- AppShell sets `data-android-native` for the Android shell (alongside the
existing iOS/Electron markers).
- index.css extends the safe-area rules to `[data-android-native]` — the header
offset, conversation/terminal top padding, sidebar + panel padding, composer
bottom padding, and the drawer slide — and sources them from `--omnigent-safe-*`
(which folds env() on iOS and the injected var on Android) instead of raw env().
The iOS-only floating Liquid-Glass bar rules stay `[data-ios-native]`.
Verified on the emulator: the header drops below the status bar, the toggle is
tappable, the sidebar opens with its header/footer clearing the system bars.
Android: gate WebView remote debugging behind BuildConfig.DEBUG (enable
buildConfig); drop the inset diagnostic logging.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): themed (monochrome) icon shows the starfish eyes, not a blob
The monochrome layer was just the solid body path, so the Android 13+ themed
icon rendered as an eyeless silhouette. A monochrome icon is single-tint, so the
eyes have to be transparent holes: build it from the body + baby starfish with
the eye circles and smile punched out via fillType="evenOdd" (filled body, holes
where the eyes/mouth are). Scaled to match the full-color foreground.
(Validated by build/aapt; the themed-icon appearance needs a launcher with
themed icons enabled — the test emulator's launcher doesn't apply them.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): harden the OIDC login flow (review round 1)
Adversarial review (Codex + Opus) of the browser-login flow:
- Use-after-destroy (HIGH): the poll runs up to 5 min on a background thread, so
it can complete after onDestroy and post onSessionToken into a destroyed
WebView (webView.loadUrl after webView.destroy()). Guard onSessionToken (and
the async setCookie callback) on isDestroyed/isFinishing/::webView.isInitialized,
and hold the session callback in a field that shutdown() nulls.
- Activity leak (MED): the in-flight poll pinned the Activity (via the bound
callback) for up to 5 min. shutdown() now uses shutdownNow() to interrupt the
poll's sleep so the task exits promptly and releases the host.
- Login-loop guard (MED): cap browser-login relaunches at MAX_LOGIN_ATTEMPTS so a
rejected cookie / expired token can't loop the browser forever; the counter
resets in onPageReady once a pinned-origin page actually loads.
- POST /auth/cli-login (LOW): set Content-Length: 0 on the bodyless POST (strict
servers/WAFs can 411 otherwise).
- Logging (LOW): route the auth-flow traces through authLog() (Logging.kt), which
only emits in debug builds — no auth event traces in release logcat.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): guard login routing on scheme + validate token shape (review round 2)
Two robustness fixes surfaced by the Gemini adversarial pass (the must-fixes
all three models converged on landed in the prior commit):
- OmnigentWebViewClient.onPageStarted: only treat a real http(s) off-origin
landing as an OIDC bounce. A null / about:blank / chrome-error:// URL is a
failed or transitional load of the pinned server (e.g. it's offline), not an
IdP redirect — the old check popped the system browser for it. Mirrors the
http(s) gate shouldOverrideUrlLoading already had. Facade injection is now
explicitly gated on the pinned origin (a non-http off-origin URL falls
through the first gate instead of returning).
- MainActivity.onSessionToken: reject a token that isn't JWT-shaped before
building the cookie string. Defense-in-depth — the token is interpolated into
the cookie value, so a ';'/whitespace-bearing value could smuggle attributes
(e.g. Domain=, defeating __Host-). A real HS256 JWT always passes.
Also folds in a behavior-preserving simplifier pass: name the repeated 10s HTTP
timeout (HTTP_TIMEOUT_MS), hoist duplicated originOf() lookups into locals, and
correct stale "Custom Tab" comments to "system browser".
Build + lint green (0 errors); 32/32 web bridge tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): canonicalize origins (default port + case); share http-scheme check
Post-review polish surfaced by the round-2 reviewers (the substantive loop had
already converged — all three models reported no new must-fix):
- originOf now canonicalizes like a WHATWG browser origin: lowercase scheme +
host and omit the default port (443/https, 80/http). The WebView reports an
origin with the default port stripped, so a user who typed `https://host:443`
previously got pinnedOrigin="https://host:443" that never matched the page's
"https://host" — breaking the bridge / looping login. Both the pinned origin
and every page URL flow through originOf, so they canonicalize identically.
(Gemini flagged this as a pre-existing latent edge.)
- Extract the duplicated http/https scheme test into isHttpScheme() and use it
at all three sites (originOf-adjacent normalizeServerUrl + both WebViewClient
nav gates). (Simplifier FYI.)
Build + lint green (0 errors); 32/32 web bridge tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): make isHttpScheme normalize case internally
Round-3 review nit (Codex): isHttpScheme gates a security boundary — which
navigations load in the bridged WebView vs. trigger login / hand off to the
system — but relied on an implicit "callers pass an already-lowercased scheme"
contract. A future caller passing a raw Uri.scheme ("HTTPS") would silently
fail to match. Lowercase internally so the predicate is self-contained; idempotent
and behavior-identical for the 3 current (already-lowercased) call sites.
All 3 round-3 reviewers (Codex/Gemini/Opus) confirmed the loop converged with no
new must-fix; this is the one accepted LOW hardening. Build + lint green (0
errors); 32/32 web bridge tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): server-independent, IME-aware safe-area insets
The shell pins to a server whose web build may predate it, so it can't rely on
the bundle's own inset rules. emitInsets now feeds the app's existing
--omnigent-safe-top/bottom vars (which every build lays out from) alongside
--omnigent-android-safe-area-*, and the bridge injects a <style> that re-asserts
the inset paddings with !important — the server's semantic inset rules otherwise
lose the CSS cascade to the Tailwind utility classes on the same elements, so the
OS inset was dropped (content under the status bar, the chat/terminal switcher
behind the gesture nav). The bottom inset is IME-aware
(max(0, systemBars.bottom - ime.bottom)) so the composer sits flush to the soft
keyboard, not a nav-bar height above it.
Reviewed via a 3-model adversarial loop (Codex/Gemini/Opus) + code-simplifier,
converged clean. Build + lint green; injected bridge JS syntax-validated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(android): system-back dismisses in-page overlays + clears login history
Android system back was leaving the app / doing nothing / landing on stale pages.
Back now first asks the page to dismiss an open in-page overlay:
- Detects an open sidebar drawer, modal dialog, or panel drawer via
data-state="open" + an on-screen (center-in-viewport) test, so the panel
drawers — which stay in the DOM at full size when closed, translated
off-screen — no longer false-match and swallow the press.
- Gated to the <768 drawer width: at md+ the side surfaces dock as persistent
rails that back must not close.
- Closes via the overlay's own Close control, else a single Escape (one per
back, so stacked overlays don't collapse together).
If nothing was open, back navigates WebView history / leaves the app.
clearHistory() drops the pre-auth + login-redirect entries on the first
authenticated load (re-armed on each re-login) so back can't walk into the IdP
redirect or a blank page. The handler is async but races a 600ms timeout
fallback (guarded against a torn-down host) so a back press always acts even if
the renderer is unresponsive.
Reviewed via a 3-model adversarial loop (Codex/Gemini/Opus) + code-simplifier,
converged clean over 2 rounds. Build + lint green; injected bridge JS
syntax-validated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): themed icon eyes — eyeball + pupil + highlight, both starfish
The monochrome (themed) launcher icon rendered the eyes as hollow holes. A
single-tint icon can't reproduce the full-color icon's white-eyeball/dark-pupil,
but it can read as eyes-with-pupils: cut the eyeball as a hole, fill a tinted
pupil dot inside it, and cut a small highlight glint in the pupil — matching the
standard icon's sparkle. The baby starfish gets the same treatment, separated
from the mama by a thin moat so both read as distinct faces.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): don't burn the login retry budget on re-entrant OIDC redirects
A multi-hop OIDC redirect can re-enter startLogin() before the first
browser hand-off settles. start() no-ops via compareAndSet when a login
is already in flight, but loginAttempts++ (and the one-shot history-clear
re-arm) ran unconditionally beforehand — so a 2-3 hop bounce could
exhaust MAX_LOGIN_ATTEMPTS without ever relaunching, suppressing a
legitimate later retry.
Make OidcLoginManager.start() return whether it actually began a flow,
and count / re-arm only on a real launch.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): harden against off-device session leak and unusable download names
- allowBackup=false: the WebView cookie store holds the authenticated
__Host-ap_session cookie, so cloud Auto Backup / adb backup would
otherwise copy a live session off-device. A server URL is trivially
re-entered; a session is not worth exfiltrating.
- BlobSaver.safeFileName: ""/"."/".." now fall back to a timestamped
name — the API 28 File path resolves "."/".." to a directory, which
would fail the write.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(android): drop stale ProGuard keep rule for a non-existent class
The rule kept ai.omnigent.android.NativeBridge with @JavascriptInterface
members, but no such class exists and @JavascriptInterface is used
nowhere — the bridge is OmnigentBridgeListener : WebViewCompat.WebMessageListener,
kept via ordinary R8 reachability plus androidx.webkit's consumer rules.
Replace with an accurate note.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): lift bottom-anchored content above the soft keyboard
Edge-to-edge (setDecorFitsSystemWindows=false) neutralizes the manifest's
adjustResize, so when the IME opens the window doesn't shrink and bottom-
anchored web content (a chat composer, a terminal input) sat BEHIND the
keyboard. The inset listener now resizes the WebView's laid-out HEIGHT by the
IME inset — a bottom margin, not padding: 100vh / the visual viewport that
fixed/sticky content anchors to tracks the view height, not its content box,
so padding alone wouldn't reflow the composer. The status/nav bars stay CSS
safe-areas so content still draws behind them when the keyboard is hidden.
Verified on an API-34 emulator (CDP: window.innerHeight and visualViewport
shrink 915->578 on IME open; a position:fixed;bottom:0 element rises to the
keyboard's top edge) and on a physical Pixel 10 Pro Fold in a real chat
composer and terminal.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(android): satisfy web format/line-ending hooks on shell files
CI's `npm run format:check` and pre-commit hooks flagged files the Android
shell added:
- README.md: Prettier normalizes `*shell*` -> `_shell_` (markdown emphasis).
- .prettierignore: exclude the Android Gradle build output, mirroring the
existing `ios/build/` entry — Gradle writes HTML lint reports that Prettier
would otherwise choke on during a local `--check`.
- ic_launcher_foreground.xml, omnigents_logo.xml: add the trailing newline
end-of-file-fixer requires.
- gradlew.bat: normalize CRLF -> LF for mixed-line-ending (--fix=lf); the repo
enforces LF everywhere and has no CRLF-preserving .gitattributes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(e2e-ui): cover the Android shell's web-side detection + safe-area fold
The Android WebView shell injects window.omnigentNative = {kind:"android"}; the
web layer feature-detects it (isAndroidShell) and tags AppShell with
data-android-native, which gates the [data-android-native] chrome in index.css —
notably the safe-area max() fold that lets the OS inset (injected as
--omnigent-android-safe-area-*) reach --omnigent-safe-*.
Mirror the desktop shell tests (sessions/test_pinned_session_hotkeys.py): inject
the bridge via add_init_script and assert data-android-native plus the resolved
inset fold, with a paired plain-browser negative test proving the gate is
Android-only. Covers the web/** change end-to-end — the chain the nativeBridge
unit tests can't reach.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(android): harden review-flagged edge paths in auth, downloads, and tap routing
Addresses the non-blocking findings from the Polly review pass:
- OidcLoginManager: accept only a rooted relative login_url from
/auth/cli-login (the server always returns "/auth/login?ticket=..."),
so a hostile/malformed absolute or scheme-relative value can't send the
one-time ticket flow off the pinned origin.
- MainActivity.onSessionToken: bail when the cookie injection is rejected
instead of reloading unauthenticated, which re-launched the browser and
burned the capped login retries on a failure retrying can't fix.
- MainActivity.downloadFile: gate on isHttpScheme(Uri.parse(url).scheme)
like the navigation gate — accepts "HTTPS://", rejects "httpfoo:" values
that DownloadManager.Request would throw on.
- MainActivity.flushPendingActivation: keep a notification tap pending when
the WebView is parked off-origin (mid re-login) rather than emitting into
a bridgeless page and dropping the path; the next pinned-origin
onPageReady flushes it.
- BlobSaver.safeFileName: take the basename past backslashes too, so a
Windows-flavored suggestion saves as "bar.txt" instead of "foo_bar.txt".
assembleDebug + lintDebug green; each change adversarially reviewed against
its call sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
_WEB_UI_DIST resolves relative to the installed package's static/web-ui/
by default. Let a deployment override it with the OMNIGENT_WEB_UI_DIST
env var, so a deploy can ship the SPA outside the wheel (e.g. as loose
files in the app source tree, to keep the wheel under a per-file size
cap) and point the server at it without rebuilding or repackaging.
Backwards-compatible: when the env var is unset the value is byte-identical
to before, so `pip install omnigent`, `omnigent serve`, and the published
wheel are unaffected. The static/web-ui package-data glob is unchanged, so
the published wheel still bundles the UI.
Co-authored-by: Isaac
Both issue triage and PR reviewer assignment now decide *who* via an LLM,
routing from one source of truth (.github/areas.json) that replaces the
split .github/reviewers (path->owners) and .github/ISSUE_ASSIGNEES
(owner->domains) files.
Each area carries a prose definition (for the LLM), file-path prefixes (for
matching), a comp:* label, and 2+ owners. Areas cover server/runner/host,
web/desktop-app/mobile-app, one per harness group, setup/onboarding,
policies, etc.
Selection: the LLM RANKS an area's owners by fit, given the definitions +
touched files (PR) or issue text. Trusted code takes the top-ranked owner,
breaking ties by open-work load. Hard constraint: the LLM can ONLY reorder an
area's own owners -- its output is allowlist-filtered against areas.json
before any GitHub call, so a hallucinated or prompt-injected login can never
be assigned.
PR path: a fail-open gateway step (same secrets/gateway as triage, via the
OpenAI-compatible /chat/completions endpoint with a Bearer token) writes a
rank file; the assigner falls back to today's pure load-balancing if it is
absent. Only changed-file PATHS are sent to the model -- never diff contents
or PR prose. All existing reviewer invariants (exactly-1, linked-issue
adoption, reconcile, push-down, fork-only, fail-closed) are preserved.
Issue path: ALLOWED_COMPONENTS is now derived from areas.json (kills the
prior drift between issue-triage.yml and config.yaml); ranked_owners + load
tie-break replaces the issue_number % N round-robin. A maintainer-authored
issue is still assigned to its author first (unchanged).
Tests: areas.test.js guards the areas.json invariants (owners in MAINTAINER,
real comp:* labels, hzub excluded, 2+ owners, path resolution incl. the
web/ ordering and kimi/kiro prefix split). auto-assign-reviewer.test.js
keeps all 16 prior assertions green (fallback = load order) and adds 4 for
rank>load, allowlist enforcement, and adoption-overrides-rank. The live
gateway wire format + ranking quality were verified end-to-end on CI.
Co-authored-by: Isaac
* fix(runner): align ws-tunnel protocol keepalive to the 90s app-level budget (#1116)
The runner<->server tunnel left its WebSocket protocol-level keepalive at the
library/uvicorn default of 20s ping-interval + 20s ping-timeout on both ends
(the runner's websockets.connect set no ping params; the server's uvicorn.run
set no ws_ping_*). That default is 4.5x stricter than the deliberate app-level
liveness budget the server already runs (_ping_loop: 30s x 3 misses = 90s), so
it pre-empts that policy: the moment a healthy runner's event loop stalls for
~20s (a synchronous / CPU-bound dispatch), the peer closes the tunnel with
"1011 keepalive ping timeout", causing reconnect churn and the downstream
"Timed out waiting for runner stream relay to subscribe" failures + 503 storms.
Set ping_interval=30s / ping_timeout=90s on both ends (shared constants in
ws_tunnel/limits.py) so the protocol keepalive is no tighter than the app-level
budget: a loop stall up to 90s (the system's own "is it dead?" line) no longer
drops a live tunnel, while a genuinely dead peer is still detected. The 30s ping
is also the runner's only liveness probe for a silently-dead server (the
app-level _ping_loop only runs server->client). The same uvicorn config covers
both the runner and host tunnel server endpoints.
This is the surgical mitigation; the deeper fix is keeping >Ns blocking work off
the event loop so a tight, responsive keepalive is safe again.
Tests: limits invariant (protocol timeout >= app-level budget, both tunnels) so a
future tightening fails CI; serve wiring (connect passes the aligned params); cli
wiring (uvicorn ws_ping_* set).
Co-authored-by: Isaac
* docs(#1116): document server-global ws_ping_* scope + precise dead-peer bound
Address Polly review on #1727 (non-blocking):
- cli.py: note that uvicorn ws_ping_* is server-global, so the 30s/90s budget
also reaches /v1/sessions/updates + terminal-attach — deliberate (those carry
their own app-level heartbeat traffic; only effect is ~120s vs ~40s half-open
reap, not a correctness change).
- limits.py: state the precise worst-case dead-peer detection bound (~120s =
30s interval + 90s timeout), correcting the earlier ~60-90s figure.
- test_limits.py: scope note that the global reach is intentional and untested
here (uvicorn-internal), pointing at the cli.py rationale.
Co-authored-by: Isaac
* fix(#1116): align host-tunnel client keepalive too (symmetric with runner)
Polly non-blocking note on #1727: the PR frames the fix around 'both tunnels'
and the test_limits.py invariant covers host_tunnel, but the host CLIENT
(host/connect.py websockets.connect) still used the 20s/20s library default —
so the host->server tunnel was only half-aligned (server tolerant, host client
would still drop the server with 1011 the instant the server loop stalls >20s,
the same failure class in the mirror direction).
Set ping_interval/ping_timeout from the shared TUNNEL_KEEPALIVE_* constants,
symmetric with serve.py's runner-side connect(). Now both tunnels are aligned
on both ends.
Co-authored-by: Isaac
* docs/test(#1116): precise idle-socket keepalive reasoning + _ConnectKwargs fields
Address Polly (non-blocking) on the rebased #1727:
- cli.py / test_limits.py: correct the 'carry their own app-level traffic'
caveat — for an IDLE sessions-updates or terminal-attach socket the protocol
PING/PONG is in fact the ONLY half-open detector (the updates heartbeat is a
server->client send; an idle terminal has no traffic). Conclusion is unchanged
(dead idle socket reaped ~120s vs ~40s, bounded, not a leak) but the stated
reason is now accurate; note the terminal-attach proxy holds its runner socket
+ tmux child ~80s longer on a half-open browser.
- test_serve.py: add ping_interval/ping_timeout to the _ConnectKwargs TypedDict
so it fully describes the asserted kwargs.
Co-authored-by: Isaac
## Related issue
N/A
## Summary
- Add `OMNIGENT_`-prefixed aliases for provider credential env vars so hosted sandboxes can keep raw provider variables out of harness processes when needed.
- Resolve prefixed aliases during provider detection, provider config secret expansion, non-interactive provider selection, global API-key auth expansion, and host-to-runner credential forwarding.
- Document the Modal setup for Claude Code API-key auth with `OMNIGENT_ANTHROPIC_API_KEY`, and keep the deployment config/docs aligned with the Modal-backed sandbox setup.
ELI5: operators can store `OMNIGENT_ANTHROPIC_API_KEY` in Modal secrets, and Omnigent translates it for its own config paths without setting raw `ANTHROPIC_API_KEY` in the Claude CLI environment.
```text
Modal secret -> sandbox host env -> Omnigent resolver -> Claude Code apiKeyHelper
`OMNIGENT_ANTHROPIC_API_KEY` no raw `ANTHROPIC_API_KEY`
```
## Test Plan
- `UV_CACHE_DIR=/private/tmp/omnigent-uv-cache PYTHONPYCACHEPREFIX=/private/tmp/omnigent-pycache uv run --extra dev pytest tests/onboarding/test_ambient.py tests/onboarding/test_detected.py tests/onboarding/test_provider_config.py tests/onboarding/test_provider_selection.py tests/test_claude_native.py tests/host/test_connect.py -q`
- `UV_CACHE_DIR=/private/tmp/omnigent-uv-cache PYTHONPYCACHEPREFIX=/private/tmp/omnigent-pycache uv run --extra dev ruff check omnigent/env_credentials.py omnigent/host/connect.py omnigent/onboarding/ambient.py omnigent/onboarding/detected.py omnigent/onboarding/provider_config.py omnigent/onboarding/provider_selection.py omnigent/runtime/workflow.py tests/host/test_connect.py tests/onboarding/test_ambient.py tests/onboarding/test_detected.py tests/onboarding/test_provider_config.py tests/onboarding/test_provider_selection.py tests/test_claude_native.py`
## Demo
N/A - non-visual environment and deployment configuration change.
## Type of change
- [ ] Bug fix
- [x] Feature
- [ ] UI / frontend change
- [ ] Refactor / chore
- [x] Docs
- [x] Test / CI
- [ ] Breaking change
## Test coverage
- [x] Unit tests added / updated
- [ ] Integration tests added / updated
- [ ] E2E tests added / updated
- [x] Manual verification completed
- [ ] Existing tests cover this change
- [ ] Not applicable
## Coverage notes
Added focused tests for prefixed credential detection, provider config resolution, non-interactive provider selection, native Claude `apiKeyHelper` wiring, and host runner env forwarding. Manual verification was the focused pytest suite and targeted ruff check listed above.
Terminals created from the web UI (POST /resources/terminals) land as
"declared" terminals — the requested name is gated against the agent
spec's terminals: block. The runner's declared-terminal branch passed
that spec's cwd straight through, and for the common placeholder
(cwd: ".") create_terminal_instance fell back to Path(".").resolve() —
the runner's process cwd, i.e. the directory `omni host` was launched
in. So new shells opened there instead of the session workspace.
Resolve the placeholder against compute_default_env_root before launch,
reusing the same _materialize_terminal_spec_for_launch /
_synthesize_parent_os_env helpers the sys_terminal_launch tool path
already uses for this. The resolved cwd is baked into the spec (not a
cwd_override, which is gated by allow_cwd_override). The synthesised
branch and the LLM tool path already resolved correctly; only this
declared-terminal REST branch was missing the step.
Fixes OMNI-1007. Also fixes OMNI-977 (managed lakebox): the workspace
comes from compute_default_env_root, which returns OMNIGENT_RUNNER_
WORKSPACE when set.
Co-authored-by: Isaac
* Add Bell-LaPadula "no write-down" to gdrive_policy; fix MCP field/tool gaps
Extend the built-in Google Drive policy (gdrive_policy) with an optional
confidential-file compartment implementing Bell-LaPadula's "no write-down"
rule: once the session reads a file in `confidential_files`, its writes are
confined to that set, so confidential content can't leak into a less-protected
file. Declared explicitly (not inferred from a per-document label), so it works
on any Drive tenant. Off by default — base access behavior is unchanged.
Also fix two gaps found while running the policy against the real Google MCP:
- Recognize `docs_document_edit_section` as a write tool (it was falling
through to the unknown-tool fail-closed branch).
- Match snake_case create-result id fields (`document_id`, `spreadsheet_id`,
`presentation_id`, `file_id`) in addition to camelCase, so files the agent
creates this session are tracked and remain writable.
Clean up the risk_score example so it no longer depends on a proprietary
`label_classification` field: the demo drives its threshold via `tool_points`,
with `sensitive_labels` documented as optional/tenant-dependent.
Adds a runnable example agent (info_flow_agent.yaml), unit tests, and
end-to-end policy-engine scenarios; existing gdrive tests unchanged.
* Address Polly review: confidential_files is containment-only, not a write grant
Revert the write-scope widening that let any file listed in confidential_files
be written/deleted even if the agent never created it and it isn't in
write_files. confidential_files is now purely a containment declaration:
writing to a confidential file still requires it to be created this session or
in write_files, matching the pre-existing write boundary. The demo CUJ is
unaffected (it writes to a doc the agent created this session).
Also document that the read-latch engages only on reads that name a confidential
file by id — content-returning reads that don't target a specific file
(drive_search, listing, exports) can surface confidential text without engaging
containment.
Update tests to the corrected semantics and add a guard that declaring a file
confidential does not by itself grant write access.
The claude-sdk harness stores sys_advise_models tool results as a JSON
content array ([{type:"text", text:"<json>"}]) rather than a raw JSON
string. parseRecommendations was calling JSON.parse on this array and
seeing no `recommendations` key, causing the SmartRoutingCard to render
"· unavailable" even when the router returned valid recommendations.
Unwrap the first text block when the parsed value is an array, then
recurse to parse the actual recommendations object.
* test(harness-bench): full-server transport driver skeleton (phase-2)
Spins up a real Omnigent server + runner OUTSIDE pytest (reusing the
live_server spawn recipe via the shared compat helpers), registers the
harness as an agent, creates a runner-bound session, and drives a basic
turn through the full session path. Live-verified: openai-agents on the
oss profile returns the marker (completed, no error).
This is the lifecycle walking skeleton. Next increments layer on the
probe-facing behaviors so the full-server path can be selected per run:
streaming-delta counting via the session SSE stream, policy DENY via
pre-attached session policy, server-dispatched tools, and interrupt/cancel
- each returning the shared TurnResult so existing probes consume it.
Bearer minting isolates DATABRICKS_TOKEN/DATABRICKS_BEARER (issue #1781).
* wip(harness-bench): full-server run_turn — tools + policy pre-attach (NOT live-verified)
Extends the full-server driver's run_turn to the probe interface
(tools/deny_phases/auto_tool_output/interrupt) and adds:
- tool_call-scoped deny policy pre-attach (POST /v1/sessions/{id}/policies
with make_fixed_action_callable action=deny on_phases=[tool_call]);
- snapshot scan for function_call / function_call_output items to populate
tool_calls and tool_call_denied, and to submit auto_tool_output on an
action_required call;
- approximate interrupt (post on running) with cancel detection.
VERIFIED: lifecycle + basic turn (openai-agents returns marker).
NOT VERIFIED: the tools/policy live path — a live openai-agents tool turn
did not complete and surfaced no function_call in the snapshot, so either
the full server does not dispatch ad-hoc request-level function tools or
the snapshot item shape differs. Needs full-server log inspection (keep the
tmp logs, trace the runner) as the next increment. Committed WIP so the
wiring is not lost; streaming via the SSE subscribe stream still pending.
* test(harness-bench): full-server transport foundation (lifecycle + basic turn)
Adds FullServerDriver: spins up a real Omnigent server + runner outside
pytest (reusing the live_server spawn recipe via the shared compat
helpers), registers the harness as an agent, creates a runner-bound
session, and drives a basic turn through the full session path (post
message, poll the snapshot to terminal, extract assistant text). A gated
live test (test_full_server.py) spins the stack up on --profile and
asserts a basic turn round-trips; it skips without creds.
Foundation for the full-server transport, whose payoff is exercising the
dimensions the wrap path cannot prove. Stacked follow-ups: server-
dispatched tools, tool-call policy enforcement (pre-attached tool_call
deny policy), delta streaming via the SSE subscribe stream, interrupt, and
the --transport selector that runs the probes through this driver.
* Add GenAI semconv attrs to AGENT and TOOL spans, gate content capture
This PR re-authored on top of upstream/main after main moved
omnigent/inner/tracing.py to raw OTel (it now returns plain
opentelemetry.trace.Span instead of mlflow LiveSpan and records I/O
via span.set_attribute(_INPUT_VALUE, ...)). The original branch's
diff was patched against the pre-refactor mlflow-shaped API and no
longer applied; this commit rebuilds the feature against main's
current shape.
What this adds
- 5 OTel GenAI semconv attribute constants in omnigent/inner/tracing.py
(_GEN_AI_OP_NAME, _GEN_AI_AGENT_NAME, _GEN_AI_PROVIDER_NAME,
_GEN_AI_REQUEST_MODEL, _TOOL_NAME) per
https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-agent-spans/
- start_agent_span now sets gen_ai.operation.name=invoke_agent,
gen_ai.agent.name=<name>, and (when model is set)
gen_ai.provider.name + gen_ai.request.model from parse_provider_name
- start_tool_span now sets gen_ai.operation.name=execute_tool and
uses the _TOOL_NAME constant for tool.name (still set unconditionally
as metadata)
- Per-attribute content-capture gate around span.set_attribute(_INPUT_VALUE)
/ _OUTPUT_VALUE on agent + tool + policy spans, controlled by
OMNIGENT_OTEL_CAPTURE_CONTENT (off by default for PII safety)
What this removes
- The dead helpers start_llm_span and end_llm_span. They had zero
production callers; production LLM spans come from inside the
spawned executor subprocess via the SDK's own tracing, not from
omnigent.inner.tracing. Per call-site-audit.md: do not ship
instrumentation on a dead path. Locked with test_dead_llm_helpers_removed.
- The _SPAN_KIND_LLM constant (no longer used).
What this scopes OUT (deferred)
- gen_ai.* attributes on LLM-level spans. Those spans do not exist in
omnigent's main process today (subprocess-side concern). Subprocess-
side instrumentation is a follow-up.
- Cross-process trace correlation (TRACEPARENT etc.) is tracked
separately on PR #1070 design discussion.
Tests
7 new tests in tests/inner/test_tracing_genai_semconv.py exercise
the production TracingContext path through a real OTel TracerProvider
+ InMemorySpanExporter (no mlflow internals, no singleton poking).
Coverage: AGENT span attrs (with and without model, with and without
provider prefix); TOOL span attrs; content-capture off/on (with PII
negative assertion that the off-path drops nothing into any attr key);
dead-helper removal lock.
Real-data verification
The semconv attributes are emitted via OTel SDK primitives, so any
real OTLP collector receives them. To verify against a real collector:
# Terminal 1: local OTel collector with debug exporter
docker run --rm -p 4318:4318 -v $PWD/dev/otel-collector.yaml:/etc/otelcol-contrib/config.yaml \
otel/opentelemetry-collector-contrib
# Terminal 2: run omnigent with the OTel exporter pointed at it
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf \
ANTHROPIC_API_KEY=$KEY \
uv run omnigent server
# Terminal 3: drive a real request
curl -X POST localhost:8000/v1/responses -d @examples/anthropic_tool_request.json
Expected: the collector debug log shows AGENT and TOOL spans with
gen_ai.operation.name, gen_ai.agent.name, gen_ai.provider.name,
gen_ai.request.model, tool.name, plus the OpenInference span-kind
attrs that main already set.
Signed-off-by: debu-sinha <debusinha2009@gmail.com>
* Apply ruff format and lint fixes
Run ruff format and ruff check on every changed file. Move atexit
import to module top (E402). Add noqa: BLE001 to telemetry-emission
swallow blocks where catching the broad Exception is intentional
(telemetry failures must not break the request path). Reorder imports
where needed (I001).
Signed-off-by: debu-sinha <debusinha2009@gmail.com>
* Hoist telemetry imports to module top + clean voice violations
Three cleanups flagged by senior-staff review:
1. omnigent/inner/tracing.py had 8 function-level imports of
should_capture_content and 1 of parse_provider_name in the hot
path (start_agent_span, end_agent_span, start_tool_span,
end_tool_span, start_policy_span). Each ran on every span creation
and was harmless but pointless. Hoist to module-top imports.
2. 2 em dashes in tracing.py comments, 3 em dashes in the test file.
Voice rule bans em dashes in code comments. Replace with periods.
3. 520 box-drawing section separators in the test file (U+2500). Voice
rule bans non-ASCII punctuation. Replace with '# ---'.
9 of 9 tests still pass. Lint clean.
Signed-off-by: debu-sinha <debusinha2009@gmail.com>
---------
Signed-off-by: debu-sinha <debusinha2009@gmail.com>
* feat(polly): add opencode as a fourth coding sub-agent
Adds an `opencode` sub-agent (harness: opencode-native) to the polly
orchestrator alongside claude_code, codex, and pi. OpenCode is a native
terminal harness, so a human can open it in the Subagents panel and take
over, and it gives polly a fourth cross-vendor implement / review / explore
worker.
OpenCode was previously dropped from polly after the version-skew incident
(#1145): older clients that did not recognize opencode-native failed to load
the whole agent. That is now mitigated on the execution path. spec.load(...,
prune_invalid_sub_agents=True) gracefully drops an unknown sub-agent instead
of failing the parent, and opencode-native is a recognized harness on current
clients, so the worst case on an old client is polly running without the
opencode worker rather than a crash.
Changes:
- examples/polly/agents/opencode/config.yaml: new worker with the standard
implement / review / explore contract and blast_radius(gate_pushes=false).
- examples/polly/config.yaml: roster is now four; preflight checks opencode;
tools.agents, routing, cancellation notes, and comments updated.
- examples/polly/skills/{investigate,fanout,cross-review}: opencode wired in
as a full peer (implementer, reviewer rotation, explore lens).
- tests: flip the polly opencode guard to expect the worker (debby stays
opencode-free), update the polly structural test roster and counts, and
update the builtin-bundles declared set.
Config plus example-agent text and tests only; no product Python touched.
* test(polly): include opencode in brain-override worker-harness map
test_materialize_bundle_overrides_brain_harness pins polly's sub-agent
name -> harness map to assert a brain-only override never rewrites
agents/<name>/config.yaml. Add the new opencode worker (opencode-native)
so the map matches the four-worker roster.
* fix(opencode-native): gate the turn path on cold-boot readiness
An opencode-native sub-agent's first (cold) turn could be dispatched before
`opencode serve` finished booting (its readiness wait is up to ~30s). The turn
path (`_stream_message_to_harness`) had no terminal-ensure for opencode, so it
raced the boot: the harness found no ready server / bridge state, produced no
result, and silently hung the parent orchestrator (polly). A warm re-dispatch
worked because boot had completed in the background by then.
Add a readiness gate on the opencode-native turn path: before obtaining the
harness client, ensure the terminal is booted (idempotent, under the same
per-session lock the session-init path uses), so the turn WAITS for the boot
instead of racing it. The events POST budget is ~1 day, so a one-time
cold-boot wait is safe, and the turn actually running means the forwarder posts
the external_session_status: idle wake as usual. A boot failure now surfaces as
a 503 turn failure (routed to the parent inbox) instead of a silent hang.
Scoped to harness_name == "opencode-native"; other harnesses are unchanged.
Set OMNIGENT_OTEL_HTTP_CLIENT_INSTRUMENTATION=false to suppress
internal httpx client spans (server↔runner↔harness API calls) from
appearing in the trace backend alongside agent/tool spans.
Co-authored-by: Isaac
From the PR #1768 automated review:
- Security: policy_deny could false-pass by denying ANY policy phase. The
driver now answers DENY only for phases the probe asks for; policy_deny
scopes its DENY to PHASE_TOOL_CALL and requires both a surfaced tool call
and a PHASE_TOOL_CALL DENY before concluding SUPPORTED. Live-confirmed:
openai-agents (previously a false SUPPORTED) now correctly reports
SKIPPED - its wrap-direct path surfaces no tool-call evaluation, so real
enforcement is a full-server (phase-2) concern.
- SdkInprocDriver.unavailable now returns a clean skip when a profile's
transport != sdk-inproc, instead of force-running a native/community
harness through the in-process driver.
- Offline (--no-live) now renders the DECLARED matrix (labeled 'declared,
not observed') instead of a grid of skips, matching the docs.
- _post records a downward verdict as delivered only on a non-error
response, so a raced/rejected policy_verdict is not counted.
Blocking finding #1 (tool-call event vocabulary) was already fixed in the
merged MVP (response.output_item.done / function_call), so no change here.
PR #1412's core change — isolate agy's config/state via the hidden
`--gemini_dir` flag while keeping the real HOME so macOS keyring auth keeps
working — already landed on main via #1598, which explicitly cherry-picked
#1412's commits. Rebased onto main, the only content this branch still adds
that main lacks is:
- test_seeding_and_mcp_config_never_mutate_real_gemini_dir: a Linux
non-regression proving seed_isolated_agy_home + write_mcp_config leave a
fully-populated real ~/.gemini (including the user's own mcp_config.json)
byte-for-byte untouched, writing only under the per-session isolated dir.
- test_auto_create_antigravity_prepends_gemini_dir_to_generated_flags:
guards that --gemini_dir is prepended ahead of every generated agy flag
(--conversation/--model/…) so the arg order is never corrupted.
- a stale-comment fix in the runner's fallback relay path: it still said
"isolated-HOME mcp_config" though main now uses the isolated --gemini_dir.
Co-authored-by: SabhyaC26 <sabhyachhabria@gmail.com>
* feat(changelog): automated changelog generation and publishing
Introduce an end-to-end changelog pipeline that turns merged PRs into a
granular CHANGELOG.md and a curated, per-version release post on the docs
site, split across the two moments in the release flow.
Authoring signal:
- Add a `## Changelog` section to the PR template; the author (or their
agent) writes one-line `<Category>: description` entries, or `skip`.
- Enforce it in the merge gate (validate.py): entries must parse, and a
Breaking change may not be `skip`. format_body.py scaffolds the section.
- Factor the shared Markdown-section + changelog parser into _md.py so the
gate and the release-time harvester never disagree.
At release cut (draft-release-notes.yml, fires via workflow_run after the
GitHub Release draft is created — runs from main, so no tagged code runs):
- Harvest each merged PR's `## Changelog` section into CHANGELOG.md and open
a PR to main (version-ordered, idempotent).
- Synthesize concise two-section release notes (release-notes-drafter agent,
tools-less claude-sdk, doc-sync security posture) and fill the GitHub
Release draft body, preserving the auto-notes in a collapsed <details>.
Falls back to a deterministic mechanical scaffold if the LLM is absent; a
hard isDraft guard never clobbers human-curated notes.
At release publish (publish-changelog.yml, site-only): mirror the curated
release body to an MDX-safe app/releases/<version> post on omnigent-site via
the omnigent-ci App token.
generate.py computes the range statelessly from git tags. Unit-tested end to
end (prev-tag selection, grouping, skip, sanitize, ordered insertion, draft
rendering, MDX transform); RELEASING.md documents the flow.
Co-authored-by: Isaac
* fix(ci): pass release tag via env in draft-release-notes to avoid injection
CodeQL flagged a critical "Code injection" alert: the "Note draft skipped"
step interpolated ${{ steps.guard.outputs.tag }} directly into the run: shell
script. Since this workflow is workflow_run-triggered, CodeQL treats the tag
(from workflow_run.head_branch) as externally controlled. Route it through a
TAG env var and reference ${TAG} instead, matching every other step in the
file — the canonical remediation, with no behavior change.
Co-authored-by: Isaac
* style(changelog): apply ruff format + lint fixes
Pre-commit ruff surfaced formatting/lint on the changelog scripts once
rebased onto main: drop unused `# noqa: E402` (RUF100), collapse
now-fitting `SCRIPT`/import statements (ruff format), and fix C416
(redundant set comprehension), RET504 (assign-before-return), and RUF005
(list concat → unpacking). No behavior change; 73 tests still pass.
Co-authored-by: Isaac
skills_filter was decoded and stored but never reached the Hermes CLI:
_build_hermes_args never emitted -s/--skills, so a configured skill set was
dropped, while the harness docstring claimed bundle_dir sourced bundled
skills. Thread skills_filter into the args (a list preloads named skills via
-s a,b; "none" maps to --ignore-rules; "all"/None add nothing) and correct
the docstring to note bundle_dir/agent_name are reserved (no hermes chat
flag yet), matching the executor's own wording.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
context_tokens (context-window fill) was only assembled in the
ResultMessage branch at successful completion, so a turn that ends the
stream without a ResultMessage (early CLI stream close, or a turn cut
short before its final usage is reported) yielded TurnComplete(usage=None).
The context-occupancy meter then froze at the previous successful turn's
value, showing a misleadingly low fill exactly when a session is in
trouble.
The latest prompt size is already observed mid-turn from each
message_start event (last_call_usage). When no ResultMessage arrives,
fall back to that observed usage and still emit context_tokens so the
meter keeps refreshing. The ResultMessage path is unchanged and still
wins whenever it runs; output_tokens is reported as 0 on an incomplete
turn rather than guessed.
Related to #1533.
Signed-off-by: abhay-codes07 <abhaysingh0293@gmail.com>
A top-level session bound to a custom agent that declares a native
terminal harness (e.g. a `polly` orchestrator with
`executor.harness: codex-native`) carries no `omnigent.wrapper`
presentation label, so `_is_native_terminal_session` returned False.
The server then persisted the inbound user message (persist-before-forward)
AND the native transcript forwarder mirrored the rendered turn back,
so every web message landed twice.
Recognize a native session by wrapper label OR resolved harness via a
shared `_native_coding_agent_for_session` helper, used by both
`_is_native_terminal_session` and `_native_terminal_runtime`. Such a
session now takes the native single-writer path (the server skips its
persist; the forwarder is the sole writer) while stamping no
presentation label, so it stays chat-first — routing is decoupled from
presentation.
Co-authored-by: Isaac
* test(harness-bench): add capability conformance suite (MVP)
Pluggable bench that probes a harness and reports a verdict per P0
dimension (basic turn, streaming, tool calling, interrupt, policy DENY,
model override), reconciling observed behavior against a self-declared
BenchProfile to surface drift.
- BenchProfile + manifest (official SDK harnesses, built from
tests/e2e/_harness_probes) with name-based resolution for community
harnesses via 'module:attr'.
- SdkInprocDriver drives turns over the harness-wrap SSE endpoint
(same path as test_harness_wrap_e2e), handling policy/tool/interrupt
round-trips.
- Six P0 probes; Verdict vocabulary maps to the support-matrix glyphs
plus SKIPPED and DRIFT.
- CLI (python -m tests.harness_bench) renders Markdown/JSON, non-zero
exit on drift.
- test_bench.py: offline conformance (always) + live layer gated on
--profile and a runnable harness CLI.
Design: docs/harness-bench-design.md. Phase-2 (native transports,
remaining harnesses, P1 dimensions) tracked there.
* test(harness-bench): classify infra/auth failures, short-circuit, progress output
Addresses two issues surfaced running the live bench:
- A gateway 403/auth failure was rendered as capability DRIFT
(basic turn/tool calling/model override ✓->✗). Turn failures whose
error matches infra/auth markers (403/401/Invalid Token/unexpected
status/connection) are now SKIPPED with an actionable reason, never
UNSUPPORTED, so a bad token can't masquerade as drift.
- When the prerequisite basic_turn does not pass, remaining probes are
short-circuited to SKIPPED (prerequisite) instead of running against a
dead turn and emitting misleading UNSUPPORTED/DRIFT (e.g. interrupt
falsely reading ✓ off a failed turn).
- The live run was silent for minutes; the CLI now streams per-harness
and per-probe progress to stderr.
- Interrupt probe no longer claims support off a turn that produced no
text before terminating.
- Live pytest skips (not fails) when basic_turn is an infra SKIP.
Adds a unit test for the infra-failure classifier.
* test(harness-bench): accurate probes + terminal-friendly output
Probe accuracy (from driving the live oss run):
- Tool calls surface as response.output_item.done (function_call item,
status action_required), not response.tool_call; the driver now matches
that and answers with tool_result, so tool-calling completes.
- Interrupts emit response.cancelled; the driver treats it as terminal,
so the interrupt probe reads SUPPORTED instead of UNKNOWN.
- Tool-calling reports SKIPPED (not a false UNSUPPORTED) when a harness
does not dispatch a request-level tool (claude-sdk/pi register tools via
config/MCP, not the wire).
- Policy DENY reports SKIPPED when no policy evaluation is surfaced in the
wrap-direct path (a server-path concern), not UNSUPPORTED.
- Interrupt probe runs last (cancelling a turn leaves the session mid-
processing and contaminated the next probe, e.g. pi 'already processing');
that error is also classified as a transient skip.
Result: the live matrix is clean (all cells ✓ or a justified ·), no false
drift.
Terminal-friendly output:
- Default is now an aligned, ANSI-colored table (color auto-off when piped
or --no-color), plus a Notes section explaining every non-supported cell.
- Markdown grid moved behind --markdown (for docs/PRs); --json unchanged.
* test(harness-bench): harden streaming probe against coalesced-delta flakiness
A streaming-capable harness (e.g. claude-sdk) occasionally coalesces a
short reply into a single delta, which read as complete-only (PARTIAL) and
drifted against the declared SUPPORTED. The probe now retries once when it
sees a single delta and only concludes complete-only if it reproduces, so
'streams sometimes' resolves to SUPPORTED and only 'never streams' stays
PARTIAL. Also uses a longer prompt and classifies infra/timeout on either
attempt as SKIPPED.
* test(harness-bench): skip hint flags stale DATABRICKS_BEARER/TOKEN
A stale DATABRICKS_BEARER (or DATABRICKS_TOKEN) exported in the shell
overrides profile OAuth in the codex gateway auth command, so a 403 keeps
firing even after re-login. The gateway-auth skip reason now points at that
env var, not just 're-login the profile'.
* test(harness-bench): make auth-skip hint provider-neutral
The 401/403 skip hint named DATABRICKS_BEARER/DATABRICKS_TOKEN, but the
symptom (an expired or ambient-env-shadowed credential overriding the
configured auth source) is not Databricks-specific: any harness can hit it
(ANTHROPIC_API_KEY, OPENAI_API_KEY, GITHUB_TOKEN, cached auth files, ...).
Reworded to point at 'the harness auth source (profile, API key, or token
env var)' without naming one provider. Detection was already provider-
neutral (401/403/Invalid Token markers).
The qwen-native forwarder stored posted-event uuids in a `set` and persisted
`list(seen)[-512:]`. Because `set` iteration is hash-ordered, that kept an
arbitrary 512 uuids, not the most recent 512 the docstring promises. After a
qwen TUI relaunch (offset rewinds to 0, file re-read from the top) for a session
with >512 events, recent uuids evicted from the window were re-posted as
duplicate bubbles in the web session.
Back `seen` with an insertion-ordered dict (an ordered set), mirroring the
sibling opencode-native forwarder, so the `[-_DEDUP_WINDOW:]` cap keeps the real
recent tail. `_read_new_events`' membership-only param is typed `Container[str]`.
Closes#1779
Signed-off-by: tomsen-ai <230283659+tomsen-ai@users.noreply.github.com>
Co-authored-by: tomsen-ai <230283659+tomsen-ai@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(copilot): gate native tools through PHASE_TOOL_CALL policy
Copilot's session was created with on_permission_request=approve_all, so
every native tool (bash/edit/view/create) was auto-approved and the
executor never evaluated PHASE_TOOL_CALL for them. Bridged sys_* tools are
gated server-side, but Copilot's built-ins could run shell commands and
edit files with no policy enforcement (cursor evaluates PHASE_TOOL_CALL for
its native tools; Copilot did not).
Install an on_permission_request handler that evaluates PHASE_TOOL_CALL via
the runtime-installed policy evaluator: a DENY rejects the individual call
(the model sees the denial and continues, rather than aborting the turn);
otherwise it approves. When no policy evaluator is wired (single-process /
pre-turn paths) the call defaults to approved, preserving prior behavior.
A small helper maps the non-uniform Copilot PermissionRequest union to a
(name, arguments) policy input, falling back to the variant's kind
discriminator when it carries no tool_name.
Interactive elicitation for native tools (the other half of the documented
limitation) is left as a follow-up; this change covers the security-
critical policy gate.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
* feat(copilot): add elicitation for native tools in on_permission_request
Adds a second stage to _on_permission_request: after a policy hard-deny
short-circuits (unchanged), the new _elicitation_handler is invoked so
users can approve or reject native tool calls from the web-UI approval
card. No handler wired → default approve, preserving prior behavior.
The adapter already installs _elicitation_handler on any executor that
declares the attribute, so no adapter changes are needed.
* fix(copilot): set harness_label to Copilot so elicitation card reads correctly
---------
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Co-authored-by: Tomu Hirata <tomu.hirata@gmail.com>
Implements intent-based permissioning as a zero-config factory in
omnigent.policies.builtins.routing.
Two-phase enforcement:
- request (first message only): records the user's stated goal as the
immutable session intent in session_state.
- tool_call: classifies each tool invocation against the stored intent
via the server-level LLM client. OFF_TASK calls are denied before the
tool runs; results are cached by (intent, tool, args) hash so
identical tool calls pay for only one classifier round-trip.
Fails open (abstains) when: no intent recorded yet, no llm_client, or
the classifier call throws. Adds 12 unit tests; updates the registry
test to cover both entries.
When an ``llm:`` block is present, ``parse`` rebuilds LLMConfig to keep
model/connection in sync with the authoritative executor fields, but the
rebuild omitted ``profile`` — silently dropping a declared credentials
profile from ``spec.llm.profile``.
This is not cosmetic: the policy/guardrail builder resolves a Databricks
workspace connection from ``spec.llm.profile``
(runtime/policies/builder.py::_resolve_server_llm_connection), so the
dropped profile makes the policy/guardrail LLM and web_fetch sub-agent
fall back to env/default auth instead of the declared workspace profile.
Carry ``profile=llm.profile`` through the rebuild. Adds a regression test
that parses llm.model + llm.profile and asserts the profile survives.
Closes#1743
Signed-off-by: abhay-codes07 <abhaysingh0293@gmail.com>
* fix(version): single source of truth for the omnigent version
The host and runner hard-coded version="0.1.0" in their hello frames,
so every host/runner reported a stale placeholder in the server's
version popover regardless of the build actually running. The server
had its own metadata->pyproject->PEP440 fallback to cope with installs
whose package metadata reports a non-PEP-440 "source" placeholder.
Introduce omnigent/version.py holding a single VERSION constant that the
runtime imports directly (no importlib.metadata round-trip), and wire the
host hello frame, runner hello frame, server /api/version, and CLI
--version to it. Importing the constant is correct regardless of how the
package was installed, so the server's fallback dance is deleted.
VERSION mirrors the canonical [project].version in pyproject.toml; a
pre-commit fixer (scripts/sync_version_py.py) rewrites the constant to
match pyproject and aborts the commit for re-staging on drift, so
releases stay a pyproject-only bump (via scripts/update_versions.py).
Co-authored-by: Isaac
* fix(version): teach the release bump path about omnigent/version.py
Polly review on #1772: the automated bump path (scripts/update_versions.py
+ .github/workflows/bump-version.yml) rewrote only the three pyproject.toml
files, never omnigent/version.py, and its `check` verified only the
pyprojects. A bot bump would therefore commit a stale VERSION constant and
trip the new test_version_matches_pyproject backstop — breaking the
"pyproject-only bump" story this change relies on.
Extend set_version() to also stamp the VERSION constant in
omnigent/version.py (anchored on its own `VERSION = "..."` line), and
extend check() to verify the constant equals the resolved [project].version
so a forgotten bump fails in the release tooling rather than on the bot PR.
The workflow's `git add -A` already picks up the extra file, so no YAML
logic change is needed — only the descriptive comment/PR body are updated.
Also soften sync_version_py.py's --check docstring, which implied a CI
wiring that never existed (per the review's non-blocking note).
Co-authored-by: Isaac
* test(version): don't assert /api/version against frozen package metadata
Polly review on #1772: the server version tests re-added
`== importlib.metadata.version("omnigent")` assertions. Since pyproject's
version is static (no dynamic wiring), that metadata is a frozen build-time
snapshot that can legitimately differ from VERSION — a stale editable
install or a "source" placeholder — the exact cases the removed server
fallback handled. Equality only holds right after a clean reinstall, so the
assertions are a latent spurious failure that undercuts the PR's
"authoritative regardless of how the package was installed" contract.
Drop the `_pkg_version` assertions in test_version_returns_source_of_truth_version
and test_info_includes_server_version (keep `== VERSION`), and remove the now
-unused import.
Also address non-blocking note 1: the --version banner (format_help) now reads
VERSION instead of importlib.metadata, for consistency with `--version`. The
upgrade path (cli.py) intentionally keeps reading installed metadata — it must
compare the on-disk install against PyPI.
Co-authored-by: Isaac
A native CLI sub-agent's completion reaches the parent orchestrator's inbox
(waking it) only when an external_session_status: idle POST hits the runner,
which rebuilds delivery via the in-memory work entry. Two gaps broke this:
- The work entry (registered at dispatch) is lost after a runner reconnect /
restart, or never registered for a sys_session_create child (the server
records a parent_session_id but no sub_agent_name). The idle handler then
found no entry and returned a silent 204, dropping the completion. Now the
runner rebuilds the entry from the server snapshot's parent linkage, and
returns 503 (so the forwarder retries) when delivery still can't be confirmed.
- cursor-native never posted the turn-end idle at all: its forwarder mirrors
only conversation items and the PTY-activity watcher is suppressed for it, so
nothing triggered delivery. cursor-agent fires a stop hook once per completed
turn (used for usage); the usage forwarder now also posts
external_session_status: idle on each newly-observed turn, the authoritative
wake edge. Idle delivery is idempotent, so a restart re-posts (server dedupes)
rather than risk skipping a wake.
The external_session_status POST helper is extracted to the shared
_native_post_delivery module so the claude-native and cursor-native forwarders
use one implementation.
Verified live: a polly-launched cursor reviewer now wakes the parent and its
result lands in sys_read_inbox instead of the parent parking idle forever.
Co-authored-by: Isaac
* fix(web): keep settings sidebar put on Members/Policies sub-pages
Clicking Members or Policies from the settings Account page navigated to
the standalone /members and /policies routes, which live OUTSIDE the
settings surface. useSettingsRoute() then reported inSettings:false, so
the sidebar swapped its section nav back to the conversation list and lit
up "New session" — the sidebar appeared to jump back to sessions.
Redesign Members and Policies as settings sub-categories:
- Add `members` / `policies` to SettingsSectionId so /settings/members and
/settings/policies resolve as in-settings sections (inSettings stays true).
- settingsNavGroups() gains an isAdmin flag and emits an admin-only "Admin"
group with Members + Policies nav items; SettingsSidebarBody reads admin
status via a new shared useMe() hook (accounts deploys only).
- SettingsPage renders the (lazy-loaded) MembersPage/PoliciesPage for those
sections and drops the now-redundant Account-section links.
- App.tsx redirects the legacy /members and /policies paths to their new
/settings/* homes so existing bookmarks still work.
Co-authored-by: Isaac
* fix(web): address Polly review notes on settings admin sections
- Fall back from the accounts-only Members/Policies sections when accounts
auth is off. `members`/`policies` are in SECTION_IDS, so useSettingsRoute
previously resolved /settings/members to an in-settings admin section even
on a non-accounts deploy — where the sidebar shows no nav item and the page
renders an empty panel. Gate them on accountsEnabled so they fall back to
the default section (still in-settings) instead of a dead one.
- Correct the useMe() doc comment: it overstated the dedup. MembersPage /
PoliciesPage still probe via a direct getMe() call (their own loading /
login-bounce state predates the hook), so they don't share this cache yet;
note that as a follow-up rather than claim it's done.
Co-authored-by: Isaac
* feat(web): click-to-zoom images in the file viewer
The file viewer rendered image files as a static <img>, while the rest of
the app (chat/session images) already opens images in a shared full-screen
lightbox with wheel/button/double-click zoom and pan. Wire the file viewer's
ImageViewer into that same lightbox via the existing useLightbox() hook so
clicking a previewed image opens it zoomable, matching the rest of the UI.
Kept the existing fit-to-container layout by calling the hook on the current
<img> rather than swapping in ZoomableImage (whose button wrapper has no
height constraint and would break max-h-full).
Co-authored-by: Isaac
* test(e2e-ui): cover file-viewer image click-to-zoom lightbox
Adds a Playwright test to tests/e2e_ui alongside the existing image-render
test: clicking a previewed image opens the shared full-screen zoom lightbox
(dialog + zoom in/out controls, same blob-backed <img>), and Escape closes it.
Satisfies the E2E UI Required gate for this UI behavior change.
Co-authored-by: Isaac
The Doc sync workflow's Plan step queried the commit→PR association index
seconds after merge, hitting GitHub's async-indexing lag and wrongly
concluding "commit has no associated PR (direct push?)" — so the merged PR
was never classified or drafted.
- Retry the commits/{sha}/pulls query with backoff (0/3/6/9s) to ride out
the indexing lag, then fall back to parsing the PR number from the merge/
squash commit subject (index-independent) if it still comes back empty.
- Move the label-driven decision into a shared block so manual
workflow_dispatch runs also honor a pre-existing label: no-doc-update
skips, needs-doc-update drafts directly, unlabeled classifies. This skips
the costly classifier turn whenever a human already labeled the PR.
- Teach the doc-classifier that a built-in policy under
omnigent/policies/builtins/ (add/remove/param change) is always
needs-doc-update — the case that slipped through (detect_task_switch, #1742).
Co-authored-by: Isaac
* chore: drop PR/issue references from code comments
Per the AGENTS.md code-comment guidance, comments should describe the
scenario rather than point at PR/issue numbers a reader must chase. Strip
the internal PR/issue/finding references from inline comments and
docstrings across production code and tests, rewording where needed so
each comment still explains what the code handles and why.
External upstream references (claude-code, coreweave/cwsandbox-client) and
local fix enumerations are left intact.
Co-authored-by: Isaac
* chore: tighten reworded comments after issue-ref removal
Fix two comments that read awkwardly after their issue references were
dropped: remove a now-duplicated parenthetical in the codex sandbox-error
guidance, and make the openai-executor regression-test docstring name the
actual scenario (missing databricks-sdk falling through to the env-var
client) instead of a vague "missing/invalid config".
Co-authored-by: Isaac
* chore: leave the initial-schema migration comment untouched
Revert the comment edit in the initial-schema migration; that file should
not change.
Co-authored-by: Isaac
* feat(routing): use live runner model catalog for intelligent routing
Pass harness→model mapping to the routing judge so it can select both
model and harness, and fetch live availability from the runner rather
than relying solely on the static lookup table.
Changes:
- runner: add GET /v1/sessions/{id}/models endpoint (catalog_for_spec)
- smart_routing: RoutingResult gains harness field; RoutingClient.route
and LLMRoutingClient accept dict[str, list[str]] (harness→models);
judge prompt now shows harness names + descriptions; harness/model
consistency enforced with fallback re-resolution on mismatch
- smart_routing: fetch_runner_models() fetches live catalog from runner;
route_turn() accepts session_id + runner_client, prefers live catalog
over infer_models fallback
- sessions: both route_turn call sites thread runner_client through;
_handle_advise_models_mcp fetches runner catalog once per call and
uses it per-agent, falling back to infer_models static table
- polly prompt: instruct polly to call sys_advise_models before fan-out
- tests: 22 tests covering new harness selection, fetch_runner_models,
runner catalog fallback, and harness/model mismatch re-resolution
* fix(routing): fix chip SSE order and restrict brain routing to self worker
- route_turn: filter runner catalog to "self" worker only; previously
the full catalog (including pi's GPT models) was passed to the judge,
causing it to pick a GPT model for a claude-sdk session
- _forward_event_to_runner: emit routing_decision chip after
_publish_input_consumed so the live SSE stream delivers the user
bubble before the chip, matching the persist order
* fix(routing): emit native chip after terminal forward, not before
Mirrors the SDK path fix: _emit_server_routing_decision now fires after
_forward_native_terminal_message so the user bubble (echoed back by the
CLI) arrives in the SSE stream before the routing chip.
* fix(routing): improve judge prompt GPT naming conventions
The judge was picking gpt-5.5 for simple tasks because the prompt
didn't clarify that -mini/-nano suffixes are cheaper than base models
regardless of version number. Clarify that nano < mini < base is the
tier order, with an explicit example.
Also log available_models before the judge call for debuggability.
* fix(routing): abstract GPT naming convention example from concrete versions
* fix(routing): fix line length in judge prompt
Add a Code comments section to AGENTS.md instructing agents to keep
comments brief (avoid >3 lines) and to describe the scenario rather than
referencing PR/issue/ticket numbers.
Co-authored-by: Isaac
* docs: add harness test bench design
Design for a standardized, pluggable capability conformance suite that
probes a harness and reports a verdict per dimension (model override,
streaming, interrupt, steering, policy DENY, etc.), reconciling observed
behavior against declared Executor flags to detect drift.
* docs: rename unofficial harnesses to community harnesses
The APPLY-mode run auto-dismisses alerts by PATCHing the Dependabot API
with dismissed_comment set to the LLM's reason. The reason was capped at
280 chars, but the "auto-triage: " prefix pushed the field to 293, over
GitHub's 280-char limit -> HTTP 422, so the dismissal silently failed
(the aws-sdk-s3 alert stayed open despite a wont_fix verdict).
Cap the whole comment (prefix included) at 280. Also split failed API
calls (status "ERR...") out of the "Auto-dismissed" headline into a
"Failed" count and emit a ::warning, so a failed dismissal is visible
instead of being counted as a success.
Co-authored-by: Isaac
* fix(ci): broaden demo-check to flag bug-fix/feature PRs and require real media
- Expand trigger from UI-checkbox-only to Bug fix, Feature, and UI /
frontend change — PRs like #1739 (bug fix with behavior change) were
previously missed.
- Replace placeholder-text matching with positive media detection:
hasDemoContent() now requires an actual image/video (markdown image,
HTML img, direct gif/mp4/mov/webm, Loom, YouTube, or GitHub-hosted
attachment). "N/A — reason" and any other non-media text no longer
pass as a valid demo.
- Narrow scan window from 14 days to 1 hour to match the hourly cron
cadence; use ISO 8601 timestamps for sub-day precision.
Co-authored-by: Serena Ruan
* fix(ci): widen demo-check scan window from 1 hour to 24 hours
Ensures PRs opened just before a cron tick aren't missed, and catches
PRs whose authors add a demo within the first day after opening.
The needs-demo label still prevents duplicate comments on re-runs.
Co-authored-by: Serena Ruan
* feat(web): installable PWA (manifest + service worker + update prompt)
Rebase of PR #116 onto upstream/main (c0907f74), relocating ap-web/ -> web/
after the upstream directory rename. Squashes the four original PWA commits
(installable PWA; build/SW hardening; Playwright e2e_ui coverage; native
desktop app icons).
Conflict resolutions:
- omnigent/server/app.py: folded the `.webmanifest` MIME registration into
upstream's new `_register_web_mimetypes()` helper (was a standalone add_type).
- tests/e2e_ui/conftest.py: kept upstream's `_codex_cli_supports_goal_mode`
alongside `_assert_pwa_build`, and pointed `--ui-skip-build` at
`_assert_pwa_build` (it subsumes the index.html existence check).
Verified: web build emits manifest.webmanifest + fingerprinted sw.js +
version.json + icons; oxlint shows no new findings; 14 PWA unit tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(e2e-ui): point PWA build guard at renamed web/ dir
The ap-web/ folder was renamed to web/; update the embed-build guard's
cwd so test_embed_build_ships_no_service_worker runs against the new path.
Co-authored-by: Isaac
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Daniel Lok <daniel.lok@databricks.com>
* feat(ci): hourly scan for contributor PRs missing UI demo
Adds a scheduled GitHub Actions workflow (every hour) that scans open
contributor PRs from the last 14 days and posts a comment + applies a
`needs-demo` label when the "UI / frontend change" checkbox is checked
but the Demo section is empty or contains only a placeholder (N/A, none,
-, tbd, todo). Drafts, maintainer-association authors, and already-flagged
PRs are skipped to avoid noise.
Co-authored-by: Serena Ruan
* fix(ci): strip unclosed HTML comment remnants in demo-check
CodeQL flagged that after removing complete <!-- ... --> blocks, an
unclosed <!-- could still remain, enabling HTML injection in the
extracted demo content. Add a second replace to strip any trailing
unclosed comment fragment.
Co-authored-by: Serena Ruan
* fix(ci): address CodeQL alert and Polly review notes in demo-check
- Fix CodeQL incomplete-sanitization: use a single regex
/<!--[\s\S]*?(?:-->|$)/g to handle both complete and unclosed HTML
comment fragments in one pass, eliminating the intermediate value
that triggered the alert.
- Flip label/comment order: comment first so a transient comment
failure leaves the PR unlabeled and retried next run, rather than
permanently suppressing the reminder.
- Remove dead COMMENT_MARKER constant (was embedded in comment body
but never read back for dedup; label is the sole dedup mechanism).
- Fix inaccurate "Skip bots" code comment to reflect what is actually
skipped (drafts + maintainer association/file).
Co-authored-by: Serena Ruan
export_agent called shutil.rmtree on a fully LLM-controlled absolute
target path, enabling arbitrary directory deletion on the user's
filesystem (contradicting its own "must not already exist" docstring).
It also built `source` with no workspace containment and copied with
copytree's default symlink dereference, so a traversal path or a
symlink inside the source could pull host files/secrets out of the
sandbox.
- Resolve `source` via safe_resolve so traversal paths and escaping
symlinks are rejected (workspace containment).
- Refuse an existing `target` instead of rmtree-ing it; never delete a
path on the user's filesystem.
- Copy with symlinks=True so symlinks in the source are preserved as
links rather than dereferenced into the export.
Extend tests: existing target is refused (no deletion), out-of-workspace
source is rejected, and a source symlink is not dereferenced out.
claude-native bakes the model at spawn time; model_override alone
doesn't change the running terminal. Send a model_change event to
the runner so it types /model <name> into the tmux pane.
Co-authored-by: Isaac
claude-native bakes the model at spawn time; model_override alone
doesn't change the running terminal. Send a model_change event to
the runner so it types /model <name> into the tmux pane.
Co-authored-by: Isaac
2026-07-01 13:57:34 +09:00
1006 changed files with 102261 additions and 28726 deletions
"definition":"The web frontend (web/) shared by all clients: React UI, components, embed. NOT the desktop or mobile app shells (those are separate areas below).",
"paths":[
"web/"
],
"owners":[
"serena-ruan",
"daniellok-db",
"hzub"
]
},
{
"key":"desktop-app",
"label":"comp:web-ui",
"definition":"The desktop app shell (Electron wrapper around the web UI): main process, packaging, native desktop chrome.",
"paths":[
"web/electron/"
],
"owners":[
"fanzeyi",
"serena-ruan",
"daniellok-db"
]
},
{
"key":"mobile-app",
"label":"comp:web-ui",
"definition":"The mobile app shell (iOS wrapper around the web UI): native mobile integration and packaging.",
"paths":[
"web/ios/"
],
"owners":[
"serena-ruan",
"fanzeyi",
"daniellok-db"
]
},
{
"key":"inner",
"label":"comp:harnesses",
"definition":"Core agent runtime and the harness/executor layer shared by all harnesses (loader, executor base, tool bridge, sandboxes). Harness-specific code has its own areas below.",
"paths":[
"omnigent/inner/"
],
"owners":[
"dhruv0811",
"TomeHirata",
"SabhyaC26",
"bbqiu",
"fanzeyi",
"aravind-segu"
],
"owners_paused":[
"dbczumar"
]
},
{
"key":"runner",
"label":"comp:runner",
"definition":"The agent runner: the execution engine that drives a turn.",
"paths":[
"omnigent/runner/"
],
"owners":[
"dhruv0811",
"bbqiu",
"fanzeyi",
"aravind-segu"
],
"owners_paused":[
"dbczumar"
]
},
{
"key":"runtime",
"label":"comp:runner",
"definition":"The agent runtime and execution scaffolding surrounding the runner.",
"_fixture_note":"FROZEN TEST FIXTURE for auto-assign-reviewer.test.js -- do NOT sync with .github/areas.json. Intentionally pinned so reviewer-logic tests don't churn when real ownership changes. Real ownership lives in .github/areas.json (validated by areas.test.js).",
"_readme":[
"Central area / codeowner map. Single source of truth for BOTH issue triage",
"(.github/workflows/issue-triage.yml) and PR reviewer assignment",
"(.github/workflows/auto-assign-reviewer.js). Replaces the old .github/reviewers",
"and .github/ISSUE_ASSIGNEES files.",
"",
"It is .json (not .yaml) on purpose: the github-script sandbox has no YAML parser",
"and the CI runner has no PyYAML, so JSON is read natively by both the JS",
"(JSON.parse) and Python (json.load) with zero dependencies.",
"",
"Each area:",
" key - stable identifier (not user-facing)",
" label - the comp:* GitHub label applied to issues in this area. MUST be",
" one of the 8 labels that already exist in the repo",
"definition":"The web frontend (web/) shared by all clients: React UI, components, embed. NOT the desktop or mobile app shells (those are separate areas below).",
"paths":[
"web/"
],
"owners":[
"SabhyaC26",
"serena-ruan",
"daniellok-db"
]
},
{
"key":"desktop-app",
"label":"comp:web-ui",
"definition":"The desktop app shell (Electron wrapper around the web UI): main process, packaging, native desktop chrome.",
"paths":[
"web/electron/"
],
"owners":[
"SabhyaC26",
"serena-ruan",
"daniellok-db"
]
},
{
"key":"mobile-app",
"label":"comp:web-ui",
"definition":"The mobile app shell (iOS wrapper around the web UI): native mobile integration and packaging.",
"paths":[
"web/ios/"
],
"owners":[
"SabhyaC26",
"serena-ruan",
"daniellok-db"
]
},
{
"key":"inner",
"label":"comp:harnesses",
"definition":"Core agent runtime and the harness/executor layer shared by all harnesses (loader, executor base, tool bridge, sandboxes). Harness-specific code has its own areas below.",
"paths":[
"omnigent/inner/"
],
"owners":[
"SabhyaC26",
"TomeHirata",
"dhruv0811",
"dbczumar"
]
},
{
"key":"runner",
"label":"comp:runner",
"definition":"The agent runner: the execution engine that drives a turn.",
"paths":[
"omnigent/runner/"
],
"owners":[
"SabhyaC26",
"TomeHirata",
"serena-ruan",
"fanzeyi"
]
},
{
"key":"runtime",
"label":"comp:runner",
"definition":"The agent runtime and execution scaffolding surrounding the runner.",
# Bumps the project version across ALL lockstep locations in one PR:
# the three pyproject.toml files (each package's [project].version plus
# its sibling ==pins) and the regenerated uv.lock. Modeled on MLflow's
# its sibling ==pins), the runtime VERSION constant in omnigent/version.py,
# and the regenerated uv.lock. Modeled on MLflow's
# dev/update_mlflow_versions.py (pre-release / post-release), adapted to
# this repo's three-package layout.
#
@@ -121,6 +122,6 @@ jobs:
--title "Bump version to ${resolved}" \
--body "Automated version bump via \`.github/workflows/bump-version.yml\` (mode: \`${MODE}\`, input: \`${NEW_VERSION}\`).
Rewrote \`[project].version\` and sibling \`==\` pins across all three packages (\`pyproject.toml\`, \`sdks/python-client\`, \`sdks/ui\`) and regenerated \`uv.lock\`.
Rewrote \`[project].version\` and sibling \`==\` pins across all three packages (\`pyproject.toml\`, \`sdks/python-client\`, \`sdks/ui\`), the runtime \`VERSION\` constant in \`omnigent/version.py\`, and regenerated \`uv.lock\`.
Generated by \`scripts/update_versions.py\`. CI does not auto-trigger on GITHUB_TOKEN PRs — re-open or push to run it."
`@${author} This PR is a **Bug fix**, **Feature**, or **UI / frontend change** but the **Demo** section is missing or only contains a placeholder.
These change types require a screenshot or screen recording so reviewers can see the new behaviour without checking out the branch. Please update the **Demo** section with:
- A screenshot or screen recording of the change, or
- A link to a hosted video or GIF showing the new behaviour.
_Use \`N/A\` only when the change has no user-visible effect whatsoever (e.g. a pure refactor or test-only change). If that's the case, uncheck the relevant type box and check **Refactor / chore** or **Test / CI** instead._`;
module.exports=async({context,github,core})=>{
const{owner,repo}=context.repo;
try{
// Load maintainers from the API so a PR can't self-grant by editing the
// file (same approach as maintainer-approval.yml).
if [ -n "$(gh pr list --repo "$SOURCE_REPO" --head "$BRANCH" --state open --json number --jq '.[].number')" ]; then
echo "CHANGELOG PR already open for ${BRANCH} — force-push updated it."
exit 0
fi
body="$(printf 'Records **%s** in `CHANGELOG.md`, harvested from the `## Changelog` section of each merged PR. Merge as part of cutting the release so the draft notes '"'"'Full Changelog'"'"' link resolves.\n\nGenerated by `.github/workflows/draft-release-notes.yml`.' "$TAG")"
gh pr create \
--repo "$SOURCE_REPO" \
--base main \
--head "$BRANCH" \
--title "docs(changelog): record ${TAG}" \
--body "$body"
# --- 5) Enrich the GitHub Release DRAFT body (only while still a draft) ---
echo "::warning::${img}:latest-dev not found yet; skipping nightly promotion"
fi
done
reconcile-floating:
# Manual reconcile (workflow_dispatch with reconcile_floating=true): repoint
# :latest and :latest-rc onto the correct EXISTING version images, computed
@@ -393,7 +378,7 @@ jobs:
fi
}
for img in ghcr.io/omnigent-ai/omnigent-server ghcr.io/omnigent-ai/omnigent-host ghcr.io/omnigent-ai/omnigent-server-openshell; do
for img in ghcr.io/omnigent-ai/omnigent-server ghcr.io/omnigent-ai/omnigent-host ghcr.io/omnigent-ai/omnigent-server-openshell ghcr.io/omnigent-ai/omnigent-server-kubernetes; do
if [ -n "$(gh pr list --repo "$SITE_REPO" --head "$RELEASES_BRANCH" --state open --json number --jq '.[].number')" ]; then
echo "Release-post PR already open for ${RELEASES_BRANCH} — force-push updated it."
exit 0
fi
body="$(printf 'Publishes the **%s** release post at `/releases/%s`, mirroring the curated GitHub Release notes.\n\nGenerated by omnigent `.github/workflows/publish-changelog.yml`. Edit the GitHub Release, not this file.' "$TAG" "$VERSION")"
echo "${DOCS_BRANCH} has nothing beyond main — nothing to publish." \
| tee -a "$GITHUB_STEP_SUMMARY"
exit 0
fi
if [ -n "$(gh pr list --repo "$SITE_REPO" --head "$DOCS_BRANCH" --base main --state open --json number --jq '.[].number')" ]; then
echo "docs → main PR for ${DOCS_BRANCH} already open." | tee -a "$GITHUB_STEP_SUMMARY"
exit 0
fi
body="$(printf 'Publishes the staged **%s** documentation to the live site: merges `%s` (%s commit(s) of doc-sync + OpenAPI updates accumulated this cycle) into main.\n\nOpened by omnigent `.github/workflows/publish-changelog.yml` on the **%s** release. Review the batch and merge to go live.' "${VERSION%.*}" "$DOCS_BRANCH" "$ahead" "$TAG")"
gh pr create \
--repo "$SITE_REPO" \
--base main \
--head "$DOCS_BRANCH" \
--title "docs: publish ${VERSION%.*} docs to the live site" \
assert("partial failure: reminder comment still posted",s.comments.length===1,JSON.stringify(s.comments));
assert("partial failure: comment does NOT over-claim a second reviewer",!/second reviewer/.test(s.comments[0].body),s.comments[0]&&s.comments[0].body);
assert("partial failure: no reviewer was actually requested",s.requested.length===0,JSON.stringify(s.requested));
assert("partial failure: still labelled (won't re-nudge next run)",JSON.stringify(s.labels)===JSON.stringify([script.LABEL]),JSON.stringify(s.labels));
assert("partial failure: the reviewer-add error is warned, not fatal",s.warnings.some((w)=>/couldnotaddsecondreviewer/.test(w)),JSON.stringify(s.warnings));
// ---- orchestration: marker fallback -- prior nudge exists but the label didn't --
assert("stale issue: one reminder comment posted",s.comments.length===1&&s.comments[0].issue_number===9,JSON.stringify(s.comments));
assert("stale issue: re-pings the assignee",/@hzub/.test(s.comments[0].body),s.comments[0]&&s.comments[0].body);
assert("stale issue: a second assignee from the label's owners",["webx","weby"].includes((s.assigned[0]||"").toLowerCase()),JSON.stringify(s.assigned));
echo "PR #$existing already open for $SYNC_BRANCH (base $DOCS_BRANCH) — the force-push updated it."
exit 0
fi
# Build the body with printf so YAML block indentation never
# leaks leading spaces into the Markdown.
short="${GITHUB_SHA:0:7}"
body="$(printf 'Automated sync of `public/openapi.json` from [omnigent@`%s`](https://github.com/%s/commit/%s).\n\nGenerated by `.github/workflows/sync-openapi-to-site.yml`. Merging publishes the updated API reference at `/reference`.' "$short" "$GITHUB_REPOSITORY" "$GITHUB_SHA")"
body="$(printf 'Automated sync of `public/openapi.json` from [omnigent@`%s`](https://github.com/%s/commit/%s).\n\nStaged on `%s` (the per-minor docs branch); publishes the updated API reference at `/reference` when that branch merges to main at release.' "$short" "$GITHUB_REPOSITORY" "$GITHUB_SHA" "$DOCS_BRANCH")"
gh pr create \
--base main \
--base "$DOCS_BRANCH" \
--head "$SYNC_BRANCH" \
--title "chore(api): sync OpenAPI reference from omnigent" \
echo "::error::package.json version ($pkg_version) != requested version ($VERSION). Is release/vscode-v$VERSION the branch created by vscode-release-pr.yml?"
# If nothing is staged, `main` is already at this version (e.g. a first
# release where package.json + CHANGELOG were prepared by hand). There
# is no diff to open a PR for, but the release branch must still exist
# so vscode-extension-release.yml can build the frozen `.vsix` from it.
# Push the branch at the current commit and skip the PR.
if git diff --cached --quiet; then
git push --force-with-lease origin "$BRANCH"
echo "No changes to release for v$VERSION — main is already at this version." \
| tee -a "$GITHUB_STEP_SUMMARY"
echo "Pushed branch \`$BRANCH\` at the current commit (no PR). Build from it with the **VS Code Extension Release** workflow." \
| tee -a "$GITHUB_STEP_SUMMARY"
exit 0
fi
git commit -m "Release (vscode): v$VERSION"
git push --force-with-lease origin "$BRANCH"
gh pr create \
--base main \
--head "$BRANCH" \
--title "Release (vscode): v$VERSION" \
--body "Bumps the Omnigent VS Code extension to \`v$VERSION\` and drafts its CHANGELOG section from the PRs merged since the last release. **Review the CHANGELOG entries and edit if needed** before merging. After merge, run the **VS Code Extension Release** workflow to build the \`.vsix\` and cut the draft release. See \`editors/vscode/PUBLISHING.md\`."
@@ -179,7 +179,7 @@ mirrors work out of the box; override with `OMNIGENT_INDEX_URL` if needed.
also launches a local web UI at `http://localhost:6767` that shows the same
session in the browser, or on a phone on your network (step 4). The
[desktop app](https://omnigent.ai/docs/interact/desktop) wraps that same UI
in a native window and adds OS notifications and a dock badge —
in a native window and adds OS notifications (with a configurable sound) and a dock badge —
[download it for macOS](https://omnigent.ai/download/mac).
> [!NOTE]
@@ -451,6 +451,10 @@ Polly at [`examples/polly/`](https://github.com/omnigent-ai/omnigent/tree/main/e
Contributions are welcome. See [CONTRIBUTING.md](https://github.com/omnigent-ai/omnigent/blob/main/CONTRIBUTING.md) for how to set up your environment, run the checks, and open a pull request.
Adding or changing support for a harness (Claude, Codex, Cursor, OpenCode,
Hermes, Pi, ...)? Run the [harness test bench](https://github.com/omnigent-ai/omnigent/tree/main/tests/harness_bench)
to check its capability matrix against observed behavior.
@@ -42,10 +42,11 @@ the generated runner Pod is already restricted-compliant (non-root uid 1000, dro
## Prerequisites
1.**A server image built with the `kubernetes` extra.** The base image omits
it, so `_ensure_sdk()` would fail every launch. Build with
`--build-arg OMNIGENT_EXTRAS=kubernetes` (see `deploy/docker`) and set the
image in `kustomization.yaml` (`images:`→`newName`/`newTag`).
1.**A server image built with the `kubernetes` extra.** The overlay's
`images:` block already points at the official `omnigent-server-kubernetes`
variant, which includes it — nothing to build. If you self-build instead,
keep `kubernetes`in`OMNIGENT_EXTRAS` (see `deploy/docker`) or
`_ensure_sdk()` fails every launch, and point `images:` at your build.
2.**Harness credentials.** The runners read their LLM / git credentials from a
Secret named by `secret_name` (default `omnigent-creds`); you create it out of
band after applying the overlay — see step 2 of **Apply**. It is deliberately
@@ -130,9 +131,9 @@ writing nothing to disk — use HTTPS repository URLs. Details by provider match
| `namespace` | Runner-Pod namespace (defaults to `omnigent-sandboxes`). |
| `secret_name` | Harness-creds Secret projected into every Pod via `envFrom`. |
| `service_account` | ServiceAccount the runner Pods run as (powerless). |
| `image` | Optional runner image override (defaults to the official amd64 host image). |
| `image` | Optional runner image override (defaults to the official multi-arch amd64/arm64 host image). |
| `env` | Optional list of SERVER env-var names to inject as literal Pod env (prefer `secret_name` for credentials). |
| `node_selector` | Optional extra node labels, merged with the mandatory`kubernetes.io/arch: amd64`. |
| `node_selector` | Optional extra node labels, merged with a default`kubernetes.io/arch: amd64` — set that key to `arm64` to schedule runners on arm64 nodes. (arm64 note: the CEL policy module is unavailable there — `cel-expr-python` ships no aarch64 wheel — and degrades gracefully.) |
The forwarded set covers the variables the harnesses themselves
@@ -324,7 +324,7 @@ like [OpenRouter](https://openrouter.ai) and
| Variable | Enables |
|---|---|
| `ANTHROPIC_API_KEY` | Claude models on the Anthropic API (claude-sdk, pi, claude-code harnesses) |
| `OMNIGENT_ANTHROPIC_API_KEY` or `ANTHROPIC_API_KEY` | Claude models on the Anthropic API (claude-sdk, pi, claude-code harnesses). Prefer the `OMNIGENT_` form for Claude Code so the raw `ANTHROPIC_API_KEY` env var is not present in the CLI process. |
| `ANTHROPIC_AUTH_TOKEN`, `ANTHROPIC_BASE_URL` | Anthropic-compatible gateways — point claude-code at a LiteLLM proxy, a Bedrock/Vertex bridge, or a corporate gateway |
| `CLAUDE_CODE_OAUTH_TOKEN` | claude-code with a Claude subscription (no API key) |
| `OPENAI_API_KEY` | OpenAI models on the OpenAI API (codex, openai-agents harnesses) |
@@ -334,7 +334,10 @@ like [OpenRouter](https://openrouter.ai) and
Common setups:
- **Claude with an API key** — put `ANTHROPIC_API_KEY` in the secret.
- **Claude with an API key** — put `OMNIGENT_ANTHROPIC_API_KEY` in the secret.
Omnigent resolves it into Claude Code's `apiKeyHelper`; do not also set
`ANTHROPIC_API_KEY` unless you are okay with Claude Code detecting the raw
custom key env var.
- **Claude with a subscription** — run `claude setup-token` on your own
machine (one-time browser auth) and store the resulting long-lived
The bench's axes are not 1:1 with capabilities: probes measure **behaviors**,
capabilities describe **traits**. Three groups:
### A. Descriptive columns → derive directly from capabilities
Replaces the hand-typed `manifest._STATIC`:
| `manifest._STATIC` column | Capability field | Note |
|---|---|---|
| `implementation` | `integration_mode` | e.g. `SDK_IN_PROCESS` → "SDK in-process". Map enum→prose in one helper. |
| `auth` | `auth` | `OMNIGENT_CREDENTIAL` / `OWN_AUTH` / `SESSION_SCOPED_CONFIG`. The old free-text ("Anthropic key / Databricks gateway") is richer prose; keep a small enum→string map if you want the exact wording, or simplify. |
| *(new columns available for free)* | `model_family`, `effort`, `resume`, `elicitation`, `subagents` | Pure metadata the report can now show without new plumbing. |
### B. Declared verdicts → derive where a capability backs the probe
| `model_override` | `SDK_MODEL_OVERRIDE_HARNESSES` (already in the registry via `model_env_keys()`) or `native` metadata | already derivable from #1756; no new field |
> **Correction (implemented, supersedes the original `False → PARTIAL` idea).**
> `streaming` is **binary**: `False → UNSUPPORTED`, not `PARTIAL`. `PARTIAL`
> is a *probe observation only* — the streaming probe returns it for the
> ambiguous coalesced-single-delta case against a `SUPPORTED` declaration — and
> is **never a declared value**. Declaring a non-streaming harness `PARTIAL`
> drifts against reality, because the probe reports zero deltas as
> `UNSUPPORTED`. This was found live: kiro/cursor/qwen-native observe 0 deltas
> and are declared `False → UNSUPPORTED` (no drift). The rule now: **declare
> `streaming=False` only from a live observation of 0 deltas** — a static
> "the forwarder posts no delta" grep is not sufficient (pi-native has no
> delta-posting forwarder yet streams live).
### C. Probe-only — no capability backing; leave hand-declared
These are behaviors with no single trait to key off. Keep them in the manifest
as-is (or a small explicit table):
-`basic_turn` — every harness is expected to complete a turn; not a
differentiating capability.
-`tool_calling` — not modeled as a capability axis (all P0 harnesses support
it; would need a new axis if that changes).
-`policy_deny` — related to `elicitation` but *not* identical (policy DENY is
enforcement, elicitation is the ASK surface). Do **not** derive `policy_deny`
from `elicitation`; keep it explicit unless you add a dedicated axis.
**Rule of thumb:** derive A and B; leave C. If you find yourself forcing a
probe-only behavior onto a trait, add a new capability axis instead (see below).
---
## Semantic shift after wiring
`verdict.reconcile()` compares declared vs live-probed. Today "declared" is a
typed guess. After this seam, "declared" = the harness's **published capability**.
So a DRIFT now means **"a harness's capability declaration is false"** — which
makes the capability table self-enforcing (you can't lie in `_BUILTIN_CAPABILITIES`
without the bench catching it on the next live run). Say this in the reconcile
output so the signal is legible.
---
## Confidence caveat (important for correctness)
Only the **four P0 SDK harnesses** — `claude-sdk`, `codex`, `pi`,
`openai-agents` — have `interrupt`/`streaming`**verified live** by the bench
today (declared `True/True`; a test in `test_harness_capabilities.py` pins this).
The other 19 harnesses' `interrupt`/`streaming` values are **declared
best-effort by integration mode**, not yet probe-verified. That is fine and
intended — it is exactly the declare-then-reconcile workflow — but the bench
wiring must not treat those 19 as ground truth. As transport drivers land for
phase-2 harnesses, their live verdicts either confirm the declaration or raise
DRIFT (which then corrects the declaration). Do not silently assume the
best-effort values are right.
---
## Adding a new axis (if a probe-only behavior needs backing)
1. Add the field to `HarnessCapabilities` (+ `as_dict()`), in
`omnigent/harness_capabilities.py`.
2. Fill it for all 23 in `_BUILTIN_CAPABILITIES`.
3. If derivable from an existing constant, add a guard test in
`tests/test_harness_capabilities.py` asserting the declaration matches its
source (see `test_model_family_matches_model_override_sets`).
Keep the model small — only add an axis when a real consumer (a probe) needs it.
---
## Suggested sequence
1.#1847 lands (capability model + `interrupt`/`streaming` axes).
2. Follow-up bench PR:
- a `manifest.py` helper `_declared_from_capabilities(harness) -> dict[dimension, Verdict]` for group B, and enum→prose helpers for group A;
- delete `_P0_ALL_SUPPORTED` and the derivable parts of `_STATIC`;
- keep group-C dimensions explicit;
- update `reconcile()` phrasing to "declared capability vs observed".
3. Phase-2 harness rollout then gets its metadata for free (all 23 already
declared) — only transport drivers remain bench-side work.
---
## Gotchas checklist
- [ ] Read the **static**`harness_capabilities()`, not `Executor.supports_*`.
- [ ] Derive groups A + B only; leave `basic_turn` / `tool_calling` /
`policy_deny` explicit.
- [ ] Don't equate `policy_deny` with `elicitation`.
- [ ] Treat non-P0 `interrupt`/`streaming` as best-effort until probed.
- [ ] Community-plugin harnesses flow through `harness_capabilities()` too —
the manifest should tolerate harnesses with no declared capabilities
| **codex-native** | explicit **`turn/steer`** RPC when a turn is active | ✅ deterministic *(verified)* |
| **claude-native** | `send-keys` into the **live pane**; the TUI folds the paste into the response | ✅ verified (best-effort timing) |
| cursor-native / hermes-native | `send-keys` paste into the **live pane** (`supports_enqueue=True`) | ⚠️ app-defined — mechanism confirmed in code, **not yet verified live** |
| pi-native | queued to the **resident extension** (`supports_enqueue=True`) | ⚠️ app-defined — mechanism confirmed in code, not yet verified live |
| opencode-native | HTTP prompt (`supports_enqueue=True`); the native server has **no live-steer endpoint** → admitted as a new prompt, promoted by the server's own queue at turn end | ❌ next turn (code-confirmed) |
| qwen / goose / kimi / kiro / antigravity -native | paste / file / RPC into the app (`supports_enqueue=True`) | ⚠️ app-defined — not yet verified live |
> **TODO (live verification):** every native harness above reports
> `supports_live_message_queue = True` and its delivery mechanism is confirmed
> in code (see the enqueue path per harness), but whether the vendor app folds
> the steered message in **mid-response** vs. at the **next turn** is confirmed
> against a *live* runner only for claude-native + codex-native. Run a live
> steer per harness to upgrade the ⚠️ rows. opencode-native is settled: its app
> server exposes no live-steer endpoint, so the steered message is always
> promoted at the next turn boundary.
**No runner change is required for native steer** — every native `run_turn`
returns right after delivering the input (decoupled from the response), so the
drain fires the next message quickly and it reaches the app while the prior
response is likely still running; the app does its own steering. Frame the UX
honestly: *"send now; the agent folds it into current work if it can"* — which is
exactly how native type-ahead already feels. Do **not** promise deterministic
mid-turn for the unverified natives.
**Steer is not interrupt.** In every case above, steer *does not cancel* the
running turn — the message is folded in at the agent's next natural breakpoint
(after the current tool/step completes), the same feel as steering native Claude
by typing while it works. For SDK, `enqueue_session_message` adds the message to
the running session's queue; the SDK surfaces it at its next turn-boundary — no
teardown. This is distinct from the **Interrupt** button, which really does
cancel the turn (`turn.cancel()`).
### Edges to handle
| Edge | Rule |
|------|------|
| POST fails after promote | revert the bubble to the queue (or error-badge it) |
| Agent goes idle mid-edit | editing pins the message out of auto-flush until re-committed |
| Native mirror-back | consume/mirror still needed as a **reconcile** signal (id-match the optimistic bubble to the real transcript item) so native round-trips don't double-render |
| native | `UserPromptSubmit` hook | `Stop` / `StopFailure` hook (relayed by the transcript forwarder) |
Both surface to the client as the same `sessionStatus` field, seeded from the
snapshot on bind (correct after refresh, across tabs).
### Live-injection gate (SDK steer)
```python
_can_forward=(
not_native# native uses paste / turn-steer, not this path
andnot_awaiting_approval# don't steer a turn parked on a human gate
andconversation_idin_live_response_id# a response is actually streaming
)
```
### Native decoupling (why paste-steer works)
Native `run_turn` returns as soon as `send-keys` finishes pasting (not when the
agent finishes). `_active_turns` clears immediately, so the buffer drains the
next message quickly and it pastes into the still-live pane — the native app then
decides to steer it. `_native_pane_status` is the reliable liveness signal for a
long autonomous native turn (since `_active_turns` clears early).
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.