Commit Graph

366 Commits

Author SHA1 Message Date
Katia Bulatova 49a1504310 docs: cut the release note back to what a user notices 2026-08-06 13:01:12 +00:00
Katia Bulatova 44b4e57003 docs: say the investigation card closes while the panel is open 2026-08-06 12:46:39 +00:00
Katia Bulatova 07b83da609 docs: note the agent's closed investigations, watch email honesty and image sources 2026-08-06 11:39:00 +00:00
Katia Bulatova 05b3472801 Merge branch 'main' into feat/dashboard-agent-flows
Route conflicts were the tab-title work meeting the agent page-context handle: both
sides kept, duplicate meta exports resolved to pageMeta, duplicate imports merged
with unused bindings dropped. Lockfile regenerated from the merged manifests.
2026-08-06 10:27:43 +00:00
Katia Bulatova 337dda1e97 feat(webapp): name of the page in tab titles (#4517)
Adds a shared `pageMeta()` helper and 74 route declarations, so a title
reads `run_abc | Runs | Trigger.dev` — the specific thing first, then
the page. Org pages also carry the organization: `Team | Acme |
Trigger.dev`. Inside a project no scope is added, because the dashboard
switches projects in every tab at once.

Page names are unchanged; what's new is that a page says which one it is
at all. Three wording changes on purpose: the queue page now names the
queue, the model page names the model, and entity pages carry their
section.
2026-08-06 10:34:30 +02:00
Katia Bulatova 50fd10bf1b fix(webapp): replay a watch submission's recorded outcome instead of re-running it
A retried card submit was only repairable while the first attempt's watch was
still active. Once it had fired, expired, or answered in one shot, the retry
re-evaluated the condition and created a second operation.

A watch_submissions ledger, keyed (chat_id, client_request_id), is now written
before the condition is read and carries the outcome once there is one. A retry
looks it up first: a recorded outcome is replayed, a different draft under the
same id conflicts, and only a pending row proceeds - converging on the watch id
reserved up front rather than creating another.
2026-08-06 07:14:40 +00:00
Katia Bulatova 49a953871a fix(webapp): record a watch request before the watch starts, and repair a retried submit 2026-08-06 01:54:44 +00:00
Katia Bulatova d719c9b2d3 fix(webapp): start the wake poll for a browser that has only an active watch
The page load reported unread wakes only, so a fresh browser whose watch was
created elsewhere and hasn't fired yet never started polling: the wake landed
without a toast or a dot until a reload. The loader now returns the active-watch
presence too, in one read per page load.
2026-08-06 01:53:16 +00:00
Katia Bulatova 6c993f213b chore: note the summarised long conversations in the release note 2026-08-06 00:49:02 +00:00
Katia Bulatova 29d9a8e423 chore: note the conversation scoring in the release note 2026-08-06 00:11:26 +00:00
Katia Bulatova ebad471b43 fix(webapp): cap the size of a message to the agent 2026-08-05 23:48:01 +00:00
Katia Bulatova cafd920060 chore: keep the release notes as one 2026-08-05 22:18:03 +00:00
Katia Bulatova e109a243f5 fix(webapp): read the alert channel and the watch target on the primary 2026-08-05 22:17:41 +00:00
Katia Bulatova 6a6fa1f4fd chore: keep the release notes as one 2026-08-05 22:13:07 +00:00
Katia Bulatova a7a29f368f chore: add a server-changes note for the report card fixes 2026-08-05 22:12:43 +00:00
Katia Bulatova 5800e8929d chore: drop the unrendered demo cards and fold the release notes back into one 2026-08-05 21:49:37 +00:00
Katia Bulatova b062490cbf chore: add a server-changes note for the report layout parity 2026-08-05 21:17:34 +00:00
Katia Bulatova 1af895bb47 fix(dashboard-agent): keep a failed turn in the conversation
A turn that ends in an error was only a stream event, so reloading history showed
a turn that just stops. It is now appended to the transcript through the same
id-deduped path a wake uses, in the same wording the live stream shows, and the
panel drops its live callout once that record is the last message.
2026-08-05 21:13:39 +00:00
Eric Allam b20806247f fix(run-store): stop run-create failing on a brief write stall (#4514)
## Summary

On the run-ops store, creating a run could intermittently fail with a
"Transaction already closed" error, and the run would never be created.
Single-write run creates no longer run inside an interactive
transaction, so a brief database write stall can't blow the transaction
budget and drop the run.

## Fix

The dedicated run-ops `createRun` / `createFailedRun` wrapped a single
nested `taskRun.create` in an interactive `$transaction`. Its default 5s
budget is wall-clock from `BEGIN`, so when a write briefly stalls the
transaction expires before the create completes and throws, even though
the statement itself is fast at the database.

A single-write create does not need an interactive transaction: Prisma's
implicit nested create is already atomic and holds no app-side budget,
so it now runs directly. Only the `triggerAndWait` path (run plus its
associated waitpoint, two writes that must commit together) keeps an
interactive transaction, now with headroom over the default.

Verified with a red/green test against the real split topology
(reproduces the exact expiry on the unchanged code, green after) and an
end-to-end run created and completed through the dedicated store.
2026-08-05 17:35:46 +01:00
Eric Allam 58bf4e2833 feat(webapp): per-client database pool and connect timeout overrides (#4515)
## Summary

Follow-on to #4513. The database connect timeout is now honored, but a
single global value has to serve three separate databases at once
(control-plane, legacy run-ops, and run-ops). This adds optional
per-client overrides for the Prisma pool and connect timeouts, one pair
for the writer and one for the read replica of each of the three
databases, each falling back to the shared `DATABASE_POOL_TIMEOUT` /
`DATABASE_CONNECTION_TIMEOUT` when unset.

That lets one database's clients run a fail-fast connect timeout (with a
bounded pool wait) while another keeps more headroom, without a single
knob forcing the same tradeoff everywhere. No behavior change until an
override is set.

It also tags each client's queries with its specific datasource
(`control-plane` / `legacy-run-ops` / `run-ops`, writer or replica) via
the `db.datasource` span attribute, so telemetry can attribute
connection behavior to a specific database instead of just
writer-vs-replica.
2026-08-05 17:26:41 +01:00
Eric Allam 771937adf5 fix(webapp): clamp run priority so a large value can't fail run creation (#4512)
## Summary

Triggering a run with a very large `priority` could fail run creation
outright with an opaque database error. `priority` is multiplied by 1000
and stored in a 32-bit integer column, with nothing bounding it, so a
big enough value overflowed the column and the create failed. The
trigger now caps the value to the highest supported priority instead of
erroring, so the run is still created.

## Fix

`priorityMs` (the stored `priority * 1000`) now goes through a
`clampPriorityMs` helper before the write. It rounds to a whole number
and clamps into the column range at both ends, so only a valid integer
ever reaches the column and an out-of-range priority caps rather than
failing. Single and batch triggers share the write path, so both are
covered.
2026-08-05 16:28:22 +01:00
Eric Allam 3039bc14d6 fix(webapp): honor the configured database connect timeout (#4513)
## Summary

Every Prisma client built its connection URL with a `connection_timeout`
query param, but the Postgres connector's parameter is
`connect_timeout`. The misspelled param is silently ignored, so all
clients fell back to Prisma's 5s default instead of the configured
timeout. When establishing a new connection briefly took longer than 5s
(for example during connection spikes), it failed with `Can't reach
database server` even though the database was healthy.

## Fix

All four client builders now construct their connection URL through one
shared helper (`buildPrismaConnectionUrl`) that sets `connect_timeout`,
so the configured value actually applies, and the parameter name lives
in exactly one place. Covered by a unit test.
2026-08-05 15:52:12 +01:00
Katia Bulatova ef65f468a2 chore: one release note for the dashboard agent, one short changeset 2026-08-05 13:23:36 +00:00
Katia Bulatova 70559be991 fix(webapp): attribute env-authed API calls to the acting user (TRI-11095)
The env JWT exchange now stamps a signed `act` claim (acting user + client
kind), auth surfaces it as `actor`, and the tenant context prefers it over
`orgMember`, which only exists on dev environments. Identity only — the JWT
still authorizes as the environment.
2026-08-04 20:36:10 +00:00
Katia Bulatova 93c4a02f4e feat(webapp): a consented watch investigation conducts itself
The wake seeded the card and said it had started looking, then nothing ran
it — the findings were left to a turn only the user could start. The watcher
now reports a delivered consented wake to the webapp, which mints the same
delegated user-actor token a turn gets and sends a `watch.investigate` action
into the chat; the agent conducts a real investigating turn on that card and
delivers the findings as its own message. Best-effort throughout: nothing here
can retry or invalidate the wake.
2026-08-04 18:44:06 +00:00
Katia Bulatova 14170648f0 fix(webapp): settle investigation cards left in progress between turns
A card opened by a turn that died, or opened for a later turn that never
came (a wake's narration does this), sat in_progress forever — a spinner on
the card and an Investigating marker in History. A new sweep on the existing
dashboard-agent cron settles anything untouched for 30 minutes to
inconclusive, with the same wording the turn-level settle uses, guarded on
the row still being in_progress so a live turn always wins.
2026-08-04 13:23:56 +00:00
Katia Bulatova 81eefe447b Merge remote-tracking branch 'origin/main' into feat/dashboard-agent-flows 2026-08-03 17:33:34 +00:00
Katia Bulatova fbd6df33b4 feat(webapp): Themes + contrast settings update (#4206)
Adds System Preferences, Dark and Light themes, gated by the
`hasThemeSwitcher` feature flag (off by default — dark stays the default
theme for everyone).

Old theme is now "Classic"and set as default. 
"System preferences" theme has both Light and Dark modes and uses your
laptop settings to use a correct one.
It has less color accents (specifically less colored text), and they are
the same for both modes, only grayscale values change between them. And
Light/Dark themes can be used separately.

New Contrast setting is available for System Preferences, Dark and Light
themes - it changes the contrast for the whole app. All new visual
Settings live in Account.
2026-08-03 19:29:33 +02:00
nicktrn 3fba04573d fix(supervisor): hold the last backpressure verdict when a read fails (#4444)
The dequeue brake released the moment its signal became unreadable.
`refresh()` caught any error from `source.read()` and set the verdict to
`null`, which `computeEngaged()` treats as not-engaged — so a few failed
reads dropped an engaged brake, silently, with no log and no metric.

That handling was symmetric while the risk is not. A source that has
stopped answering correlates with the pressure the brake exists for, so
releasing on read failure gives up protection at exactly the wrong
moment; holding too long only costs throughput.

Now a failed read keeps the last verdict instead of discarding it. The
verdict then ages normally, so the existing `maxVerdictAgeMs` check
becomes the grace window and still bounds how long a dead source can
hold the brake — a permanently unreachable source releases it rather
than pinning dequeuing forever. Because `computeEngaged()` only consults
staleness for an *engaged* verdict, a released one is unaffected and
stays released.

The default grace moves from 15s to 120s, comparable to how long the
brake normally stays engaged.

One guard worth calling out: holding is only safe when something bounds
it, so when `maxVerdictAgeMs` is unset the previous discard behaviour is
kept. Otherwise an unbounded hold could pin the brake indefinitely.

Read failures were previously invisible — the catch block neither logged
nor counted. Adds a `read_failures_total` counter, plus an error log on
the transition into failure rather than once per tick, since the refresh
loop runs every second.

The post-release ramp needs no change: it anchors off the
engaged-to-released transition, so a grace-window release still ramps
back up instead of snapping to full rate, which is what you want after a
blind period.

Tests cover holding while reads fail, releasing past the max age, and
the existing unbounded-config paths are unchanged.
2026-08-03 18:06:23 +02:00
nicktrn 8f9db53350 feat(supervisor): configurable tolerations for run pods (#4491)
## Summary

Self-hosted Kubernetes deployments can now add tolerations to run pods,
so runs
can schedule onto tainted nodes. Previously the only way to do this was
to patch
the supervisor.

`KUBERNETES_RUNNER_TOLERATIONS` takes a comma separated list of
`key=value:effect`, or `key:effect` to tolerate any value. It applies to
every
run pod, and for runs from a schedule tree it merges with the existing
`KUBERNETES_SCHEDULED_RUN_TOLERATIONS`. Left unset, nothing changes: no
tolerations are added and the pod spec leaves the field off entirely.

The Helm chart takes it as a list:

```yaml
supervisor:
  config:
    kubernetes:
      runnerTolerations:
        - dedicated=runs:NoSchedule
        - spot:NoExecute
```

## Naming

The issue proposed `KUBERNETES_WORKER_TOLERATIONS`. This ships as
`KUBERNETES_RUNNER_TOLERATIONS` instead, because `RUNNER_*` is already
the prefix
for run pod settings (`RUNNER_HEARTBEAT_INTERVAL_SECONDS`,
`RUNNER_ADDITIONAL_ENV_VARS`, and `DOCKER_RUNNER_NETWORKS` for the
Docker
equivalent), whereas "worker" refers to the supervisor itself throughout
this app.

## Validation

Keys and values are checked against the Kubernetes naming rules when the
supervisor starts, so `dedicated=prod runs:NoSchedule` fails immediately
with a
message naming the offending entry. Without that check a bad value is
accepted at
startup and then rejected by the API server on every pod create, which
stops all
runs with the cause buried in an API error.
`KUBERNETES_WORKER_NODETYPE_LABEL` is
trimmed and validated for the same reason: surrounding whitespace is not
valid in
a label value, so a padded value fails every pod create today.

## Node selector off switch

`KUBERNETES_WORKER_NODETYPE_LABEL` accepts an empty string to skip the
node
selector entirely, so runs schedule on any node. This already worked and
the Helm
chart has always shipped it empty, but it was not documented. It is now.

The issue also asked for general node affinity configuration. That is
not
included: the node selector off switch plus tolerations covers the
reported
problem, and a free form affinity setting is a much larger config
surface to
commit to.

Fixes #4458
2026-08-03 15:40:41 +00:00
Katia Bulatova 586dc2d250 Dashboard Agent: Watch (background condition watches + wake notifications + alerts) (#4456)
Stacked on #4418 — the diff against that branch is the complete Watch
feature, extracted so the base agent PR can land without it.
2026-08-03 17:36:48 +02:00
Eric Allam 9d57aff542 fix(webapp): make the Queues hero charts environment-wide (#4486)
## Summary

The four charts above the queues table aggregated over **at most the 25
queues on the current page**. They reused the loader's already-paginated
queue array as a ClickHouse `queue IN (...)` filter, so paging or
re-sorting changed the values, and a name search matching nothing
blanked the whole chart row. The stat tiles above them were already
environment-wide, so the two rows disagreed.

They now read `env_metrics`, the environment-level rollup that already
exists for exactly this (the built-in Queues dashboard and the health
report read it). That is both correct and queue-count-independent: no
`GROUP BY queue` across an entire environment, and no client-side
summing.

Note this is not only a paging artifact: page 1 under-reported too. On
the seeded environment below, page 1 read 82% saturation against a true
87%, because the environment's running total is not the sum of one page
of per-queue gauges.

Three related fixes ride along.

**Scheduling delay and throttling sawed to zero.** Both are
event-driven, so at the 10-second bucket a short range picks, most
buckets hold no samples at all and were drawn as `0ms`. Measured over a
1-hour window: **232 of 349 buckets had no scheduling-delay samples**. A
bucket where nothing started is not a bucket where nothing waited, so
the line was both ugly and wrong. TRQL grows a `minBucketSeconds` floor,
plumbed through the metric resource route, and the hero tiles set 60s.
Buckets that still have no samples render as a gap instead of a dive to
zero.

**The floor must not feed a width-dependent headline.** Two of the four
headlines are not peaks, so widening the plotted buckets moved them:

- **Throttled** is a share of buckets that saw any throttling, so a
single brief throttle came to mark a whole minute instead of ten
seconds: the same seeded events read 17% at 10s and 85% at 60s.
- **Scheduling delay p95** is a percentile, and merging quantile states
over a wider bucket yields a p95 between the sub-buckets' own. Two 240s
samples among twenty in one 10-second sub-bucket give a worst-of-six p95
of 240,000ms against a merged 60-second p95 of 5,000ms — a 48x
understatement of a headline whose tooltip claims it is the worst in the
window.

Both charts keep the floor, since a readable line was the point of it.
Their headlines now come from a second query at the range's natural
bucket width, via an optional `readout` on the tile, so each means what
its tooltip says regardless of how the plotted buckets are sized.
Saturation and backlog are genuinely width-invariant (a max of maxes is
the same at any width), so they are unchanged and issue no extra query.
Both caught by Devin in review; I had wrongly lumped p95 in with the
peaks.

**Charts reported a hydration mismatch on every render.** Recharts
resolved victory-vendor's CJS entry on the server and its ESM entry in
the browser. Those bundle different d3-shape builds, and the CJS one
predates d3-path's digit rounding, so every server-rendered curve
carried full-precision coordinates while the client rounded to 3
decimals:

```
Server: M0,3C0.9305555555555555,3,1.8611111111111112,3,...
Client: M0,3C0.931,3,1.861,3,...
```

Bundling recharts for SSR makes both sides resolve the same ESM build.
Verified: 45 of 45 server-rendered chart curves now match the client,
and the page loads with an empty console.

## Verification

An isolated stack with 40 seeded queues (20 heavily loaded, 20 idle) and
90 minutes of 10-second buckets written into `queue_metrics_raw_v1`, so
the real materialized views built `queue_metrics_v1`, `env_metrics_v1`
and the 5m rollup. Ground truth for the environment: 260 running against
a limit of 300 (**87% saturation**), 800 queued.

| | before | after |
| -- | -- | -- |
| Saturation, page 1 | 82% peak | **87% peak** |
| Saturation, page 2 | 5% peak | **87% peak** |
| Backlog / delay, page 2 | "No activity" | **800 peak / 59.5s** |
| Name search matching nothing | all four charts blank | charts stay
environment-wide |
| Metric refetches on a page change | 4, each painting a skeleton | **0,
no skeleton** |
| Buckets drawn as 0ms with no samples | 232 of 349 | **0** |
| Throttled readout | 17% | **17%**, unchanged by the wider buckets |
| Worst-p95 readout source | plotted buckets | **natural width**, so a
sub-minute spike is not averaged away |
| Crosshair reach, hovering one detail-page chart | 2 of 4 others | **4
of 4** |
| SSR chart curves mismatching the client | 45 | **0** |

The bucket floor was measured across ranges: it widens 10s to 60s at 30m
and 1h, and is correctly a no-op at 12h (300s) and 7d (3600s). One extra
request per page load, for the throttled readout.

The built-in Queues dashboard, which reads `env_metrics` independently,
agrees at 86.7% and 260 of 300.

`internal-packages/tsql` suite green (612 tests), including 5 new ones
for the floor that fail without it. Webapp typecheck, oxfmt and oxlint
clean. Spot-checked the Run metrics dashboard and the per-queue detail
page for SSR regressions from bundling recharts: both render, console
clean.

The queue detail page carries the same event-driven series, so its
scheduling delay, throttling and per-key mean delay take the same
treatment.

## Screenshots

<img width="2540" height="580" alt="after-page1-charts"
src="https://github.com/user-attachments/assets/6cd23f9c-e7fd-4918-bcfa-b1d3340b16d1"
/>

## Rollout

Already behind the per-organization `queueMetricsUiEnabled` flag, so
only gated orgs see any of it. Blast radius is chart values on one page
plus the SSR bundling of recharts; rollback is a revert with no data
migration.

## Stated limitations

- `wait_ms_count` and the quantile state both only count `wait_ms > 0`,
so "nothing started in this bucket" and "everything started instantly"
are indistinguishable in storage. Both render as a gap. Distinguishing
them needs a schema change, which is not in this PR.
- The queue name search deliberately no longer narrows the charts. It
only did so incidentally and incorrectly before (first 25 matches, and
blanked on zero matches). Search-scoped charts would need the full
unpaginated matching set and a server-side aggregate; worth its own
ticket if we want it.
- Bundling recharts for SSR grows the server bundle slightly. That is
the cost of both sides resolving one d3-shape build.
- The plotted delay line is a smoothed 60-second view, so a sub-minute
spike above the one-minute warning threshold can fail to colour the line
even though the headline reports it and colours itself.
- Every chart inside one synced group shares the floor, because the
hover crosshair is a reference line on a category x-axis and only draws
where the hovered bucket exists in the other chart's own data. That
costs the queue detail page's gauges some resolution (1 minute instead
of 10 seconds) in exchange for the crosshair working across the row.

Separately, while taking the screenshots I found a pre-existing
rendering bug unrelated to this change: a **perfectly flat** saturation
series draws no line at all (the readout still shows the right
percentage), which looks like the threshold gradient's offset
degenerating when the series min equals its max. It reproduces on
`main`, so it is not a regression here and I have left it alone; filed
as its own issue.

Refs TRI-12784
2026-08-03 16:19:50 +01:00
Katia Bulatova 447223e81f feat(webapp): page context and suggested prompts for every env page
Only the runs, errors, queues and deployments pages described themselves
to the dashboard agent; everything else fell to "other" and offered the
generic chips. Add handle mappers for the remaining 37 env-scoped routes
and 24 page kinds in the contracts, so each page offers an explain and a
docs question about what it actually shows.

Investigate and status chips stay gated on loader data: a scheduled task
with no schedule attached, all its schedules disabled, a paused queue, a
batch whose runs failed, a wait token past its timeout, a bulk action
still running, a spent quota, a prompt pinned to an override, a session
whose run failed. Loader data only, no added queries, no new signals.
2026-08-03 15:16:58 +00:00
Katia Bulatova 859f30e224 fix(webapp): report message catalogs survive the production bundle (#4488)
GET /api/v1/reports/health threw `no catalog registered for report
"health"` in production (fine in dev): the catalog registered itself as
a side effect of a bare import, which the SSR build tree-shakes under
`"sideEffects": false`. Verified on the built server bundle — main's is
missing the catalog, this branch's carries it.

Fix: catalogs are values on the report registry entries; the resolver
reads them from there and the mutable register-at-import step is gone.
2026-08-03 16:12:03 +02:00
Katia Bulatova 1751c97ca1 feat(webapp): page context for the errors, queues and deployments pages
Only the runs list and run detail described themselves to the dashboard
agent, so every other page fell to "other" and offered generic chips.
Add handle mappers for the errors list, an error group, the queues list,
a queue, the deployments list and a deployment — loader data only, no
added queries — plus list page kinds in the contracts and an optional
deployment status. Investigate chips now appear for an unhealthy queue
and a deploy that didn't land.
2026-08-03 13:46:18 +00:00
Chris Arderne 763b5dc582 feat(webapp): enforce scopes for environment API keys (#4389)
## Summary

Environment API keys backed by the additional-key table can authenticate
API requests using their stored effective scopes. Revoked and expired
keys are rejected, branch environments retain their existing routing
behavior, and last-used timestamps are updated on a throttled
best-effort basis.

## Design

API route builders receive the resolved ability and reject restricted
keys on routes without an authorization declaration. Existing
deployment, environment variable, queue, run, task, batch, session, and
waitpoint routes declare the resources they access.

Trigger and batch responses return server-signed public access tokens,
so additional keys never need access to the environment signing secret.
Root-key rotation also keeps public tokens valid for the existing grace
window.

## Feature notes
- Root environment keys remain unrestricted for backward compatibility.
Additional keys enforce their persisted scopes and fail closed on routes
   without an authorization declaration.
- Machine-key requests never exchange one credential for another.
Additional keys cannot retrieve the root key, and rotated root keys are
not upgraded
   during their grace window.
- Public JWT validation remains host-owned, while installed RBAC plugins
continue to supply root-key abilities.
- Unfiltered session and run listings preserve existing broad task-read
behavior. Filtered requests enforce the supplied task identifiers.
- Related-run summaries remain embedded in run retrieval for API
compatibility. Retrieving or mutating a related run independently still
requires
   permission for that run.
- Queue management authorizes at collection scope, matching the queue
permissions currently issued.
- Batch responses deliberately include server-signed public access
tokens for all clients. Selected-task credentials continue using their
original
   credential for per-item authorization.
- Two-phase batches authorize declared task identifiers before creation
and authorize every streamed item. Streaming paths that cannot declare
the
   complete task set remain fail closed.
- Authentication telemetry records successful credential resolution
separately from subsequent resource-authorization failures.
- API keys are high-entropy random tokens. SHA-256 is intentionally used
for deterministic indexed lookup, not password hashing.

## Deployment notes

The schema migration must be present before this code is deployed.
Because bearer resolution runs on every authenticated request, deploy
the resolver with additional-key lookup disabled, verify root-key and
public-token parity, then enable lookup before any additional keys can
be issued.

The multi-task authorization tightening changes the result for narrowly
scoped tokens that request tasks outside their grants. Observe
would-deny results before enforcing that check. Request-idempotency keys
are also newly isolated by environment and task, so a retry crossing the
deployment boundary may execute once more before old cache entries
expire.

## Follow-ups

- [x] Add a system-wide kill switch for additional-key lookup, defaulted
off for the initial deployment.
- [x] Add authentication observability by credential kind, result,
latency, and lookup path without recording credential values.
- [ ] ~Add would-deny observability and an independent enforcement
switch for multi-task authorization.~
- [ ] ~Add an independent switch for server-issued batch tokens while
root-key parity is verified.~
- [ ] Confirm every API route reachable by a restricted key has an
explicit authorization declaration or intentionally fails closed.
- [x] Verify root-key rotation, revoked-key grace, and public-token
validation through each bearer resolver path.
2026-08-03 14:00:29 +01:00
Eric Allam 5f29ae49ab feat(webapp): default the queue metrics period to 1 hour and remember it (#4438)
## Summary

The Queues list and queue detail pages opened on a 1 day window, and
went back to it every time you navigated between queues or reloaded.
They now default to the last hour, and the period you pick is remembered
across navigations and refreshes.

## Design

The last period is stored in a `queueMetricsPeriod` cookie, written
client-side whenever a `period` lands in the URL and read by both
loaders. A cookie rather than localStorage because the queues list
renders its per-queue metrics columns server-side: with localStorage the
page would paint the 1 hour default and then re-fetch, and the picker
would flash the wrong window.

Both pages resolve the window once, in one place, and pass it down:

```ts
period: resolveQueueMetricsPeriod({
  period: value("period"),   // a usable period in the URL wins
  from: value("from"),       // an absolute range means "no period"
  to: value("to"),
  defaultPeriod,             // otherwise the remembered default from the loader
}),
```

That keeps the picker pill and every chart query on the same value, so
no call site falls back to its own default. Periods the picker could
never produce (a hand-edited `?period=garbage`, or a window past the 30
day retention) fall back to the default, and the picker renders the
resolved window rather than the raw search param so the label can't
disagree with the data. Absolute from/to ranges, including drag-to-zoom,
are not remembered, since they would pin later visits to a window that
has gone stale.

While wiring that up: the two queue-metric queries that go straight to
ClickHouse (the list table and the concurrency-keys endpoint) never
applied the org's `queryPeriodDays` limit, so a hand-typed `?period=`
read further back than the plan allows. Everything behind
`/resources/metric` is already clipped that way by `executeQuery`; both
of these now clip with the same limit, capped at the retention window,
and the plan cap is resolved once per load and handed to the page
instead of each route deriving its own copy from the client-side
subscription.

Verified on both pages: default with no cookie is 1 hr, picking 6 hrs
survives navigating away and back to a param-free URL and a hard reload,
clearing the cookie returns to 1 hr, an oversized period falls back
without being remembered, and an absolute range still renders as a
range.
2026-08-03 10:09:34 +01:00
Katia Bulatova 9d48f42d55 feat(webapp): blank-state hero and fullscreen mode for the agent panel
- the empty chat centers a hero: sparkles icon, 'Ask Trigger' at blank-slate
  title size with the Beta badge, a one-line subtitle, a three-row composer
  with the send button inside, and the suggested prompts as a wrapping row
  of buttons colored by meaning (action indigo, status secondary, explain
  tertiary, docs the docs style) — one slot-to-variant mapping
- an Expand button next to Close takes the panel over everything right of
  the nav bar, like a page: no route, no modal, no remount — the chat keeps
  its transport and draft text; content stays mounted underneath; the
  transcript column gets a prose max-width; the preference persists
- storybook: hero states at panel and fullscreen widths
2026-08-03 08:58:17 +00:00
Katia Bulatova dfd575b955 feat(webapp): Investigate entry points on degraded queues and waiting runs (TRI-12862)
The queue detail page offers Investigate when the queue is at capacity with
runs waiting or the head-of-line wait passes the existing warning threshold;
the run page's waiting widget offers it whenever it renders. Both post the
visible request in the user's own voice through the existing button.
2026-08-02 21:40:18 +00:00
Katia Bulatova 65a4942a6f Merge remote-tracking branch 'origin/main' into feat/dashboard-agent-flows 2026-08-02 21:13:09 +00:00
Iss 8f66af6e18 fix(webapp): stop the sidebar feedback popover from canceling the submit (#4445)
The Help & Feedback → "Contact us" form in the sidebar intermittently
failed to send. The `<Feedback>` dialog was nested inside the Help
popover, so clicking **Send** closed the popover and unmounted the form
mid-submit — canceling the `POST /resources/feedback` before it went
out. The message was silently lost (the success toast still shows). A
race, so it "worked sometimes"; the standalone "I'm stuck!" path was
unaffected.

**Fix:** host the Feedback dialog *outside* the popover (same pattern as
`AskAIRoot`) and open it from the menu item, so closing the popover no
longer tears down the form. `Feedback` gains an optional controlled
`open`/`setOpen` mode; existing `button`-triggered usages are unchanged.

## Changes

- `Feedback.tsx` — optional controlled `open`/`setOpen`; `button` now
optional.
- `HelpAndFeedbackPopover.tsx` — "Contact us…" opens a `<Feedback>`
hosted outside `PopoverContent`.
- `.server-changes/fix-sidebar-feedback.md` — user-facing note.

## Testing

Webapp typecheck passes. Sidebar "Contact us…" now sends on every
attempt (Network: `POST /resources/feedback` → `204`, never
`(canceled)`); "I'm stuck!" and the `?feedbackPanel=` open path
unchanged.
2026-08-02 14:49:09 +01:00
Katia Bulatova 19d6625efa refactor: extract the Watch feature to feat/dashboard-agent-watch
The base agent branch now ships without watches: schedule_watch, the tick
loop, wake delivery, the expiry sweep, watch alerts (email template, alert
type, unsubscribe), the watches table and its migrations, and every UI
surface (chips, wake banner, toast, unread dot, watching status) are gone.
The complete feature lives on feat/dashboard-agent-watch, stacked on this
branch. The review stand (seeder, heartbeat, guidebook) stays here.
2026-08-01 19:52:36 +00:00
Eric Allam f9c8d518c7 perf(webapp,run-engine,database): resolve the newest worker and deployment by createdAt (#4452) 2026-08-01 11:32:21 +01:00
Eric Allam 0445b8ec27 fix(webapp,clickhouse): keep the rest of a ClickHouse batch when one run or span has un-ingestable JSON (#4358)
## Summary

A single run output, trace span, or payload carrying JSON that
ClickHouse can't ingest (for example nesting past its depth limit) used
to fail the whole insert batch, so unrelated runs and spans silently
disappeared from the runs list, traces, and logs. This keeps the rest of
the batch and handles the offending row instead of dropping everything
around it.

## Fix

Recovery is per-table, matched to what each table needs:

- **Runs** (`task_runs_v2`) keep their status. We follow ClickHouse's
failing-row hint to strip just the un-ingestable JSON column(s) so the
run still lands (its output reads from Postgres on the detail page), up
to a configurable limit (`RUN_REPLICATION_MAX_POISON_STRIPS_PER_BATCH`,
default `1`). Past the limit we stop and land the batch with
`allow_errors` in a single pass, skipping the remainder. Cost stays a
fixed handful of inserts no matter how large or poisoned a flush is.
- **Trace events and payloads** (high volume, append-only) recover with
a single `allow_errors` insert: the good rows land in one pass and only
the un-ingestable rows are skipped.

Before falling back, a lightweight sanitizer still repairs what it can
losslessly (lone UTF-16 surrogates, out-of-range integers) so a
repairable row lands in full.

To read the failing-row hint we patch `@clickhouse/client-common`: its
error parser truncates the server response and discards the `(at row N)`
position, so the patch preserves the full text for the recovery path to
read.
2026-08-01 09:17:20 +01:00
nicktrn b42e5c3771 fix(supervisor): count pods from a limit=1 list instead of an aggregate metric (#4442)
The pod-count backpressure source read
`apiserver_storage_objects{resource="pods"}` from an apiserver
`/metrics` scrape. That gauge is a periodically-refreshed cached count,
and it is served by whichever apiserver replica the scrape lands on —
replicas disagree with each other at the same instant, by enough to
swamp the engage/release hysteresis band. Engage and release timing was
therefore partly a function of scrape routing.

This replaces it with a single `limit=1` list of the workload namespace
and computes `remainingItemCount + items.length`. One pod object
transferred, no informer, no watch cache.

Two request-shape constraints are load-bearing and called out in the
code: passing a label or field selector makes the apiserver omit
`remainingItemCount` entirely, and setting `resourceVersion` serves a
cached count rather than a quorum read. Neither is passed.

`remainingItemCount` is only set when the list is truncated, so
`_continue` is the truncation signal — if it is absent the returned page
is the whole collection and `items.length` is already exact. If the list
*is* truncated and the count is missing or implausible, the fetcher
throws rather than guessing.

Failure semantics are unchanged: a throw lands in the monitor's existing
catch, exactly as the previous parse did. The hysteresis, verdict shape,
and gauge are untouched. RBAC is unchanged — the existing role already
grants `pods: list`.

The `/metrics` non-resource grant in the deployment role becomes unused,
and the scrape-timeout env var is now a slight misnomer. Both left alone
deliberately: the grant may be wanted again for other apiserver signals,
and renaming the var would need a coordinated config change for no
behavioural gain.

Tests cover the not-truncated, truncated, missing-count, negative-count
and timeout paths.
2026-07-31 19:45:37 +01:00
Eric Allam f10bc23785 perf(run-engine,run-store): one execution snapshot per triggered run (#4419)
A non-delayed run used to get two execution snapshots the moment it was
triggered: `RUN_CREATED` nested in the run-create transaction,
immediately followed by `QUEUED` from its own `BEGIN`/`INSERT`/`COMMIT`.
It now gets a single `QUEUED` snapshot written inside the create, and
the trigger path only publishes to the queue. One fewer row per run on
`TaskRunExecutionSnapshot`, and one fewer round trip on the trigger hot
path.

`EnqueueSystem` gains a `publishRun` seam that enqueues without writing
a snapshot. Every re-enqueue path (waitpoint resume, checkpoint restore,
delayed enqueue, pending version, retry requeue) still calls
`enqueueRun` and writes its own `QUEUED`, so only the first enqueue
changes. The `QUEUED` snapshot still commits before the queue message,
so a dequeue sees a dequeueable status exactly as before.

Two things for reviewers. Nesting the write skips
`createExecutionSnapshot`, which is what emits
`executionSnapshotCreated` and therefore the run timeline's `[engine]
QUEUED` entry, so the trigger path now emits it directly, the same way
the dequeue and attempt-start paths already do for their nested creates.
And `RUN_CREATED` is still written when a dequeued run has no background
worker yet, so the status and both `statuses.ts` helpers stay live and
existing rows keep reading correctly.

Delayed runs are untouched: `DELAYED` then `QUEUED` are two genuinely
different moments and stay two snapshots.

Rollback is a revert. Create-and-enqueue happen in one request in one
process, so no in-flight run needs both code paths to agree during a
rollout.


One note for whoever debugs this path later. The `QUEUED` snapshot now
commits before the queue publish, so a failed publish leaves the run
recorded as `QUEUED` with no queue message. That state was already
reachable, since the publish was never part of the snapshot transaction,
but it used to be recorded as `RUN_CREATED`, which was distinctive
because it never otherwise persisted. `QUEUED` with no message is
indistinguishable from a run waiting on a concurrency slot, so
trigger-time publish failure is now one more cause of an apparently
stuck queued run.
2026-07-31 16:12:07 +01:00
Eric Allam c72ebf9084 fix(webapp,run-engine): stop batchTriggerAndWait hanging when item streaming never completes (#4397)
## Summary

`batchTriggerAndWait()` could leave a parent run waiting forever. The
2-phase batch API blocks the parent on the batch's waitpoint as soon as
the batch is created, but the batch is only sealed at the end of item
streaming. If streaming never completed, nothing sealed the batch,
nothing completed the waitpoint, and the parent stayed suspended with no
timeout and no way to recover.

Supersedes #4016, which added the reaper alone.

## Fix

Admission for item streaming was being decided twice. Batch creation
passes its own rate limiter, which fixes `expectedCount` and blocks the
parent, and then the item stream had to pass the general API limiter as
well, competing with unrelated traffic. A second limiter could therefore
veto work the first had already committed the parent to. Creation now
mints a bounded grant that the item stream spends, so an admitted batch
can finish streaming. The grant is capped per batch rather than
exempting the path, and every failure mode (no grant, spent grant,
unreachable store) falls back to the normal limiter.

That makes stranding much rarer but not impossible, since a request
timeout or a crash can still end streaming for good. So a seal-timeout
reaper aborts any batch still unsealed after `BATCH_SEAL_TIMEOUT_MS` and
completes the parent's waitpoint with an error, letting
`batchTriggerAndWait()` reject instead of hang. It is race-safe against
a late seal, and it is only scheduled for batches that actually block a
parent, so fire-and-forget batches cost nothing.

Finally, the batches page used to report "Batch completion checked." for
these batches while doing nothing, because the completion path returns
early on an unsealed batch. It now says the batch cannot be resumed.

Rate limiting is no longer the reason a batch strands, so the reaper's
default stays at 30 minutes, comfortably above the SDK's worst-case
stream-retry budget.

## Verification

Unit and container tests cover the grant cap, the bypass ordering (it
runs after the authorization check, so it can never skip
authentication), and the reaper's abort, seal race, idempotency, and
no-waitpoint cases.

Also verified end-to-end against a running stack. With the general limit
exhausted, batch creation and other API calls returned 429 while a
granted batch still streamed and sealed; an ungranted batch id was rate
limited rather than bypassed; and the grant cut off exactly at its
configured attempt count. Reproducing the stranded state on a real
parent run, the batch was aborted at the timeout, the waitpoint
completed with an error, and the parent resumed and finished instead of
hanging. A parentless batch left unsealed was untouched well past the
reaper window.

## Verified against deployed runs

The reaper was proven end to end with a real deployed run (locally-run
supervisor, containerised
run) and a real network fault, rather than a simulated one: toxiproxy
severs the phase 2 item
stream mid-flight so every SDK stream retry genuinely fails, while phase
1 still succeeds. Only
the batch calls traverse the fault, so control-plane traffic is
untouched.

The reproduction is the shape that actually strands a parent: the task
catches the
`BatchTriggerError` the SDK throws and carries on, so the phase 1 block
outlives the thrown error
and the parent hangs at its next suspension point.

With the reaper disabled, the parent sat in `EXECUTING_WITH_WAITPOINTS`
for over 24 minutes holding
two blockers, and stayed stuck across a full infrastructure restart:

```
 type     | status    | has_timeout
 BATCH    | PENDING   | f            <- orphan, completedAfter NULL
 DATETIME | COMPLETED | t            <- the wait already elapsed
```

With the reaper enabled the same task under the same fault completed in
about 75 seconds with zero
blockers left, the batch `ABORTED`, and its waitpoint completed carrying
the error.

Two conditions are required to observe this at all, which is worth
knowing for any future test:
the run must be deployed rather than `trigger dev` (dev runs execute in
process and finish while
still holding blocker rows), and the wait after the caught error must
exceed the checkpoint
threshold, or it is served in process and never suspends.

### Why completing the batch waitpoint is sufficient

`batchTriggerAndWait` runs create, then stream, then wait. A phase 2
failure throws before the wait
is ever reached, and the reaper only fires on an unsealed batch, so the
parent is never suspended
awaiting the batch when it runs. The parent therefore does not need a
synthetic result, only to stop
being blocked. Note this reasoning depends on that ordering: if the wait
were ever reached with an
unsealed batch, completing the batch waitpoint alone would not settle
the caller.

## Follow-ups

- Batches stranded before this ships still need a one-off recovery; the
reaper only schedules at creation time.
- That same property leaves a gap if the process dies between creating
the batch and scheduling the job. A periodic sweep would close it, but
wants a supporting index.
- When a partially streamed batch aborts, children already enqueued keep
running while the parent fails. Left as-is deliberately, since
cancelling triggered work is a bigger semantic call.
2026-07-31 11:55:25 +01:00
Katia Bulatova c929c0fb6a feat(webapp): TRI-12763 agent UI polish
- header aligned with the title bar; custom +/x icons; history as a title
  dropdown popover with shorthand ages and delete confirm
- composer: tighter corners, icon-only send/stop, context path above the field
- chat bubbles: assistant text is plain prose, user bubbles grey
- agent-identity.ts single source for name/icon (TODO: character icon from design)
- Help & Feedback: 'Ask Agent' (cmd+J) via open-request bridge; Kapa AI fully removed
- free plan: 20-message cap with upgrade gate and countdown notice
- storybook canvas matches the panel background; chart min-height and padding
2026-07-31 09:04:14 +00:00
claude[bot] debfa2b733 feat(webapp): impersonation consent page and a view-as-user toggle (#4421) 2026-07-30 21:28:44 +01:00
Oskar Otwinowski 4efe0a07c4 fix(webapp): create dev environments for SSO and Directory Sync members (#4426)
Members added by SSO just-in-time provisioning or Directory Sync never
got
their per-member DEVELOPMENT environments - only invite acceptance and
project creation created them. `trigger dev` returned "Environment not
found" for those members and the dashboard had no dev view.

ensureOrgMember now queues provisioning for every membership it settles,
so
both paths are covered and members missing environments are repaired on
their next sync. Provisioning runs as a common-worker job to keep
sign-in
and directory webhooks off the per-project write loop. A failed enqueue
surfaces for Directory Sync, whose worker retries the idempotent effect,
and is swallowed for sign-in, where the next login enqueues again.
Environment creation now tolerates a concurrent creator so the
project-creation loop and the job cannot collide on the unique index.

Also fixes environment resolution ignoring dev-environment ownership: a
member without their own dev environment could be handed a colleague's
and
have it persisted as their dashboard preference.
2026-07-30 21:50:15 +02:00