Commit Graph

320 Commits

Author SHA1 Message Date
Katia Bulatova abaeebbefa fix(webapp): make a repeated watch card close leave the revision alone
`settleInvestigationStateAndCloseCard()` bumped the revision before it looked at
the transcript, so a redelivered or replayed action settled the row a second time
while the append refused the duplicate card. The row then held revision 2 and the
transcript's terminal card revision 1: the live run rendered 2, a refresh rendered
1, for a tool advertised as idempotent on the action id.

The transaction now locks the investigation, checks the tenancy triple, locks the
chat, and looks for the message id before it writes anything. An action already in
the transcript returns the stored card and the current revision untouched.

Lock order stays investigation then chat, matching `persistTurn` and the sweep's
`settleInvestigationAndCloseCard`; taking the chat first here would deadlock
against them.

A missing or deleted chat was indistinguishable from an already-closed card —
both came back `closed: false` — and the settle committed anyway, which is exactly
the terminal-row-without-a-card the transaction exists to prevent. It is now
`{ ok: false, error: "chat_missing" }`, and the watch lane logs it as the race it
is rather than a fault.
2026-08-06 13:56:42 +00:00
Katia Bulatova afc0cfaac8 fix(webapp): close a consented watch investigation's card atomically
The watch lane settled the investigation with `settleOpenInvestigations`, then
appended the terminal card as a separate write — and swallowed that write's
error, logging it and reporting success. So the row went terminal, the card
never arrived, the stale sweep stopped selecting the row because it was no
longer `in_progress`, and the user was left on "Working…" with nothing able to
repair it. Nothing in production calls the action again on its own, which is
exactly why the swallowed error mattered.

The lane now writes through `settleInvestigationStateAndCloseCard`: the terminal
revision and the closing card commit in one transaction, under the lane's own
message id, and the error propagates so the task's retry is a real retry. If the
card cannot be rendered the settle rolls back, leaving the `in_progress` row the
sweep still selects. The duplicated revision bump is gone with it — one outcome
is now one revision.

The regression test drives a failing close in one action; it no longer proves
recovery by calling the action a second time by hand.
2026-08-06 13:23:15 +00:00
Katia Bulatova f6f753c45a fix(webapp): merge the transcript under the row lock instead of replacing it
`persistTurn` and `persistMessages` stored the whole `messages` array they were
handed. The array is the snapshot the turn started from, so anything another
process appended in between was deleted: a wake delivery, a watch consent
record, or the terminal card of an investigation the stale sweep had just
settled. That last one is unrecoverable — the row is already terminal, so the
sweep never selects it again, and the panel is back to "Working…" for ever.

Both writes now read the row under `select ... for update` inside the
transaction and merge by stable message id: incoming order is kept, a stored
message the snapshot does not have goes at the end, and no id appears twice. A
message with no id falls back to its content so it cannot be carried over twice.

An append-only `chat_messages` table is the better long-term shape; merging
under the lock is enough for this architecture and needs no migration.
2026-08-06 13:23:14 +00:00
Katia Bulatova 37c56aaa15 test(webapp): pin the settlement failure window and the open panel
Covers the two halves the earlier pass left open. Against a real database, a
stale card that cannot be rendered now proves the sweep rolls the settle back
with it: the row stays `in_progress` at revision 0, the chat stays empty, and
the row is still in the next run's selection.

The panel half is covered over `liveProgress`, the code that decides whether
"Working…" is shown: a mounted panel holding the unconcluded card re-reads the
stored transcript, merges by stable id, and the progress line goes away without
a reload. Re-reading repeatedly cannot add a second copy of the card.
2026-08-06 12:42:27 +00:00
Katia Bulatova 657f11daf0 docs(webapp): say which UAT flow the project-wide answer is preserved for 2026-08-06 12:18:36 +00:00
Katia Bulatova 1759618b7f fix(webapp): keep the project-wide answer for an environment-agnostic user-actor token
Binding a delegated token to its environment claim on the project-wide
routes also refused any token that carries no claim at all, which the
public PAT exchange used by MCP and the CLI is allowed to issue. Line the
project-wide helper up with its neighbours: a claimless dashboard-agent
token is still refused, everything else stays project-wide.
2026-08-06 12:18:35 +00:00
Katia Bulatova e62d8c66f3 style(webapp): format the route image CSP audit 2026-08-06 12:18:15 +00:00
Katia Bulatova 60a958414b test(webapp): fail if a route sets an over-broad image CSP 2026-08-06 12:18:14 +00:00
Katia Bulatova 26f93da1a9 fix(webapp): have the stale-investigation sweep close the card it settles
The sweep settled the row and appended nothing, so it visibly fixed nothing: the
chat kept rendering the last card it had, which was still "Working…". The settle
now returns the state and revision it wrote, and the sweep appends that as the
closing card revision on the chat — id-deduped on
`investigation-settlement:{id}:{revision}`, so a retried run can neither stack a
second card nor open a second investigation.

The append is scoped by chat id: a sweep runs off any session and has no user in
context, unlike the turn lane.
2026-08-06 11:34:59 +00:00
Katia Bulatova 49e8fb578c fix(webapp): tell a watch's creator about their own email alerts, not a colleague's
The create-watch response matched any watch-alert channel in the project, so a
second member was told they were subscribed while the mail went to the first.
The channel's deduplication key is the only record of whose it is, so state
resolution, subscribe and unsubscribe now share one owner lookup.
2026-08-06 11:27:43 +00:00
Katia Bulatova c4d3f585e3 fix(webapp): replay a watch submission's recorded email outcome instead of re-deciding it
replay() re-ran subscribe() for an already-recorded `created` submission. A retry that
succeeded could flip the ledger to `enabled`, but the confirmation in the transcript is
append-once, so the user kept being told email was unavailable while the system believed
it was on. The replay now reads recordedExternalNotification() and takes no external
decision. An attempt that dies before subscribing or before its outcome is recorded
leaves the row `pending`, and the normal creation path subscribes on the retry.
2026-08-06 11:27:42 +00:00
Katia Bulatova 79482dea37 chore(webapp): drop the comments the tests already say 2026-08-06 10:49:05 +00:00
Katia Bulatova 3336a47d1e test(webapp): give the container-backed user-actor scope tests their own timeout 2026-08-06 10:32:58 +00:00
Katia Bulatova f700174b44 fix(webapp): ceiling an exchanged environment JWT by what the delegated token can do
A user-actor token declaring no scope cap could mint an environment JWT with any
scopes it asked for. The exchange now clamps the minted scopes to the actor's own
ability, so a capless token is read-only, and only mints for its claimed environment.
2026-08-06 10:32:56 +00:00
Katia Bulatova d00b1266e4 fix(webapp): bind a delegated token's environment claim on the project-wide routes
The environments and runs listings are project-wide, so an environment-scoped
user-actor token could read every environment its user can reach. Both routes now
resolve the claim into a mandatory filter, and a conflicting request filter is
refused rather than overridden.
2026-08-06 10:32:54 +00:00
Katia Bulatova 210863ea76 fix(webapp): separate a watch's last look from its last check in the batch order
A tick that could read nothing now moves the group's fairness key only, so a watch
with a permanently broken reader stops crowding out the rest of an over-cap group.
Dueness and the streak facts still follow the last real check.
2026-08-06 10:28:05 +00:00
Katia Bulatova 22bf6d919a fix(webapp): cancel the watch a losing submit's winner does not name
A refusal that won the race kept the reserved watch id, so the user was told
nothing was created while that watch stayed active.
2026-08-06 10:28:04 +00:00
Katia Bulatova 09d1b55e22 fix(webapp): require the watch card's request id instead of falling back to its condition
A per-condition fallback identified the condition rather than the submit, so a
re-watch could replay a stale terminal outcome from the retention window.
2026-08-06 10:28:03 +00:00
Katia Bulatova cd112ba2a3 test(webapp): assert the document img-src has no wildcard host 2026-08-06 10:27:59 +00:00
Katia Bulatova 933aeb9ea4 test(webapp): stub the agent proxy's environment lookup, not a Prisma row
The hand-built row stopped satisfying every column the authenticated-environment
mapper reads, so both proxy cases failed. Stubbing the lookup keeps the fixture
independent of the row's columns.
2026-08-06 09:44:50 +00:00
Katia Bulatova dd90b6e968 fix(webapp): never render an image the agent's model authored
A remote image in model output fetches on render, so its URL is an exfiltration channel. No image is rendered at all, and the document CSP now names the hosts images may come from.
2026-08-06 09:42:43 +00:00
Katia Bulatova a0640ba248 style(webapp): format the watch tenancy changes 2026-08-06 09:40:30 +00:00
Katia Bulatova a63708c6e6 test(webapp): assert the derived chat id before the submit outcome 2026-08-06 09:40:29 +00:00
Katia Bulatova 5176e3e8c4 test(webapp): pin the watch ledger's tenancy, the once-only alert and the owner-scoped unsubscribe 2026-08-06 09:40:29 +00:00
Katia Bulatova 012ce0de53 test(webapp): drop an unsafe optional chain from the sweep boundary test 2026-08-06 09:39:17 +00:00
Katia Bulatova bd2559a7a8 fix(webapp): record only a real evaluation from a batch tick
An unreadable check is not an observation, so writing it moved the watch down the
rotation for a whole cadence and overwrote the facts a streak lives in.
2026-08-06 09:39:16 +00:00
Katia Bulatova e4852090aa perf(webapp): resolve the sweep's authorization and readers once per group
An incident expires a whole group at once, which is the worst time to repeat the
same authorization read per row. Bounded concurrency caps one tenant's share.
2026-08-06 09:39:15 +00:00
Katia Bulatova d2511bd264 fix(webapp): hand the sweep's final check the previous facts
A stateful condition is a transition across checks, so a boundary evaluation
without the last observation can never see a stall and records a reset.
2026-08-06 09:39:15 +00:00
Katia Bulatova d12e6a85ca style(webapp): merge duplicate imports on the user-actor auth path 2026-08-06 09:24:09 +00:00
Katia Bulatova f58a9dcc89 fix(webapp): only let the browser set the agent's page context
The chat turn's metadata is now a whitelist: the page the user is on is all the browser can set, and every identity, tenancy and token field is filled in by the server.
2026-08-06 09:24:09 +00:00
Katia Bulatova d51cb9a736 fix(webapp): stop the OSS fallback handing a delegated token blanket access
Without an RBAC plugin, a user-actor token got the same permissive ability a personal access token gets. It now gets only what its own scope cap allows, and reads only when it declares no cap. Personal access tokens are unchanged.
2026-08-06 09:24:08 +00:00
Katia Bulatova 9d45284112 fix(webapp): keep the agent token's environment claim on the actor through every route builder
The claim rode beside the authentication result, so a route builder that only forwarded the user id lost it. It now travels on the authenticated actor itself, and a route's resolved org/project/environment is checked against it — a mismatch fails closed.
2026-08-06 09:24:07 +00:00
Katia Bulatova f5a146b30c fix(webapp): cap the agent's body on any method and on a mixed-case path 2026-08-06 08:50:11 +00:00
Katia Bulatova 85b98ab8dc fix(webapp): mint the agent's token for the caller's own environment, or not at all 2026-08-06 08:20:52 +00:00
Katia Bulatova 50fd10bf1b fix(webapp): replay a watch submission's recorded outcome instead of re-running it
A retried card submit was only repairable while the first attempt's watch was
still active. Once it had fired, expired, or answered in one shot, the retry
re-evaluated the condition and created a second operation.

A watch_submissions ledger, keyed (chat_id, client_request_id), is now written
before the condition is read and carries the outcome once there is one. A retry
looks it up first: a recorded outcome is replayed, a different draft under the
same id conflicts, and only a pending row proceeds - converging on the watch id
reserved up front rather than creating another.
2026-08-06 07:14:40 +00:00
Katia Bulatova 4ff8859e1d fix(webapp): keep the agent's alert delete to channels the watch type is on 2026-08-06 01:54:46 +00:00
Katia Bulatova cbc2edc879 test(webapp): pin the watch submit's ordering, its retry repair and its refusal record 2026-08-06 01:54:45 +00:00
Katia Bulatova 28b1f3896c fix(webapp): cap the agent's request body while it streams, not after
The message-size checks ran after the body had been read, so a request without a
content-length was buffered and parsed in full before being refused. An ingress
cap on the agent's paths now counts the bytes as they arrive, and the chat proxy
reads its body with a ceiling instead of reading it whole first.
2026-08-06 01:53:17 +00:00
Katia Bulatova d719c9b2d3 fix(webapp): start the wake poll for a browser that has only an active watch
The page load reported unread wakes only, so a fresh browser whose watch was
created elsewhere and hasn't fired yet never started polling: the wake landed
without a toast or a dot until a reload. The loader now returns the active-watch
presence too, in one read per page load.
2026-08-06 01:53:16 +00:00
Katia Bulatova e599f1299c fix(dashboard-agent): rotate an over-cap watch group instead of starving it
The batch took the 500 soonest-expiring watches of a group every tick, so a group
larger than the cap could leave the rest unchecked until the first 500 expired.
The group is now ordered least-recently-checked first, with a generated
cadence_minutes column and an index so the due predicate no longer re-parses the
spec JSON per tick.
2026-08-06 01:53:15 +00:00
Katia Bulatova 7f13add1a0 feat(webapp): retire judged-turn rows after 30 days
The table is append-only quality data with no reader, so the sweep now drops rows past
the period in one bounded statement per run.
2026-08-06 00:11:21 +00:00
Katia Bulatova d6debe8fb7 test(webapp): cover the user-actor token's environment scope on every route that accepts one 2026-08-05 23:34:45 +00:00
Katia Bulatova 1dd85367b7 perf(webapp): answer the wake poll with one narrow read per database 2026-08-05 22:17:42 +00:00
Katia Bulatova e109a243f5 fix(webapp): read the alert channel and the watch target on the primary 2026-08-05 22:17:41 +00:00
Katia Bulatova 0d415c04ef chore(webapp): strip the step narration from the agent and report tests 2026-08-05 21:35:59 +00:00
Katia Bulatova 78aea57480 test(webapp): assert the card and the text report stay one report
Landmarks derived from the layout spec must appear in the same order in the server-rendered card
and in the markdown, over healthy, degraded, untrustworthy and empty reports.
2026-08-05 21:17:33 +00:00
Katia Bulatova 5aef047d87 fix(webapp): cap the query API's scope at the credential's environment
/api/v1/query took its tenant scope from the request body, so a public access
token minted for one environment could read the whole organization by asking for
organization scope. The credential is now the ceiling and a wider ask is refused;
secret keys and the membership-authorized Query page are unchanged.
2026-08-05 21:13:40 +00:00
Katia Bulatova a303cb0648 refactor(webapp): one presenter for all watch and view-block wording
Every watch sentence (card, banner, toast, email, Slack, webhook) now comes from
app/presenters/v3/dashboardAgent. The condition is written once per kind in four
registers instead of in four switches across three files, and a plain-text block
renderer lets the non-React surfaces say what the card says.
2026-08-05 21:13:39 +00:00
Katia Bulatova c331843dfc chore(webapp): shorten the watch service and test comments 2026-08-05 16:21:30 +00:00
Katia Bulatova a5dd3f13b6 chore(webapp): shorten the report, waiting-run and agent route comments 2026-08-05 16:21:30 +00:00