Commit Graph

4891 Commits

Author SHA1 Message Date
Katia Bulatova 9b4e4734a1 perf(dashboard-agent): share one 1h-TTL prompt cache between head start and the agent
Head start passed `system` as a bare string, so Anthropic neither wrote nor read
the cache: the ~18.4k-token prefix was billed in full and the agent's step 2 then
paid for a fresh write. Both breakpoints now carry a 1-hour TTL.

The two prefixes were also not identical — `get_run` sat in a different position in
the agent's tool set than in the canonical schema-only one, so they could never have
shared a cache. Reordered, and a fingerprint over system text + tool definitions is
logged per model call alongside the provider's cache-write, cache-read and uncached
input token counts, so a future drift is visible instead of silent.
2026-08-05 22:16:30 +00:00
Katia Bulatova d4f0357ac1 chore(webapp): move the gallery's transcripts into the storybook fixtures and drop demo-chats 2026-08-05 22:12:44 +00:00
Katia Bulatova 0760f1b55d refactor(webapp): give the agent cards one card primitive 2026-08-05 22:12:42 +00:00
Katia Bulatova 3f85acc2a7 fix(webapp): give the diagnosis card links a focus ring via a TextLink variant 2026-08-05 22:12:41 +00:00
Katia Bulatova 5a80876b16 fix(webapp): ring the unread dot against the navbar surface 2026-08-05 22:12:40 +00:00
Katia Bulatova 9b32bfda66 fix(webapp): let agent chips follow the interface-contrast control 2026-08-05 22:12:39 +00:00
Katia Bulatova fd83288906 fix(webapp): drop the metric sparkline to its own line on a narrow panel 2026-08-05 22:12:38 +00:00
Katia Bulatova 3773f31d22 fix(webapp): route the report's stale-data badge through AgentBadge 2026-08-05 22:12:37 +00:00
Katia Bulatova 5800e8929d chore: drop the unrendered demo cards and fold the release notes back into one 2026-08-05 21:49:37 +00:00
Katia Bulatova 8d3676a874 chore: oxfmt the lines the removed comments let collapse 2026-08-05 21:38:32 +00:00
Katia Bulatova 9270caf715 chore(webapp): cut the watch service and route comments to their invariants 2026-08-05 21:38:24 +00:00
Katia Bulatova 5153cca045 chore: restore the comments an off-by-one over-trimmed 2026-08-05 21:38:13 +00:00
Katia Bulatova 9b7ca8c1de chore: bring the last three-line comments down to two 2026-08-05 21:37:46 +00:00
Katia Bulatova e6ec1d2d11 chore(webapp): delete the agent panel and view rationale comments 2026-08-05 21:37:45 +00:00
Katia Bulatova a4f8ed1906 chore(webapp): trim the seed-script narration 2026-08-05 21:37:09 +00:00
Katia Bulatova 0739337c24 chore: cut the report, waiting-run and package comments to their facts 2026-08-05 21:37:08 +00:00
Katia Bulatova 93bc6c158d chore(webapp): clear the demo fixture and storybook commentary 2026-08-05 21:36:19 +00:00
Katia Bulatova 0d415c04ef chore(webapp): strip the step narration from the agent and report tests 2026-08-05 21:35:59 +00:00
Katia Bulatova cb70b83059 chore(webapp): drop the prose from the chat markdown css rules 2026-08-05 21:35:24 +00:00
Katia Bulatova bd6835f9d9 chore(webapp): drop the page-context notes from the route files 2026-08-05 21:35:23 +00:00
Katia Bulatova 17e11c771f chore(webapp): cut the chat-layout header and stop the test asserting on it 2026-08-05 21:35:22 +00:00
Katia Bulatova f3237decb7 chore(webapp): drop the unused demo report card 2026-08-05 21:21:52 +00:00
Katia Bulatova 78aea57480 test(webapp): assert the card and the text report stay one report
Landmarks derived from the layout spec must appear in the same order in the server-rendered card
and in the markdown, over healthy, degraded, untrustworthy and empty reports.
2026-08-05 21:17:33 +00:00
Katia Bulatova 880d94ae3a feat(webapp): build the report card from the shared layout
The card no longer formats values, resolves codes or classifies footer codes of its own, and a
genuinely-unknown state now gets a neutral icon instead of a green tick.
2026-08-05 21:17:33 +00:00
Katia Bulatova f6865d581a feat(webapp): render the markdown and ANSI report from the shared layout
The text surfaces gain the card's structure (headline, hero evidence, why: block, aggregated
read:, trust flag) and keep their alignment: fixed sparkline column, aligned notes.
2026-08-05 21:17:32 +00:00
Katia Bulatova 547ec711d8 feat(webapp): declare the report layout once, shared by the card and the text surfaces
Section order, section labels, the tone -> glyph vocabulary, the footer styles and every
string a renderer places now live in report-layout.ts. Renderers decide typography only.
2026-08-05 21:17:32 +00:00
Katia Bulatova 5aef047d87 fix(webapp): cap the query API's scope at the credential's environment
/api/v1/query took its tenant scope from the request body, so a public access
token minted for one environment could read the whole organization by asking for
organization scope. The credential is now the ceiling and a wider ask is refused;
secret keys and the membership-authorized Query page are unchanged.
2026-08-05 21:13:40 +00:00
Katia Bulatova 1af895bb47 fix(dashboard-agent): keep a failed turn in the conversation
A turn that ends in an error was only a stream event, so reloading history showed
a turn that just stops. It is now appended to the transcript through the same
id-deduped path a wake uses, in the same wording the live stream shows, and the
panel drops its live callout once that record is the last message.
2026-08-05 21:13:39 +00:00
Katia Bulatova a303cb0648 refactor(webapp): one presenter for all watch and view-block wording
Every watch sentence (card, banner, toast, email, Slack, webhook) now comes from
app/presenters/v3/dashboardAgent. The condition is written once per kind in four
registers instead of in four switches across three files, and a plain-text block
renderer lets the non-React surfaces say what the card says.
2026-08-05 21:13:39 +00:00
Katia Bulatova 04687767c9 chore: oxfmt the blank lines left by the removed comments 2026-08-05 16:21:32 +00:00
Katia Bulatova c331843dfc chore(webapp): shorten the watch service and test comments 2026-08-05 16:21:30 +00:00
Katia Bulatova a5dd3f13b6 chore(webapp): shorten the report, waiting-run and agent route comments 2026-08-05 16:21:30 +00:00
Katia Bulatova 6ac42141e6 chore(webapp): shorten the suggested-prompt and demo-fixture comments 2026-08-05 16:21:29 +00:00
Katia Bulatova d1fff9c45b chore(webapp): shorten the agent panel and view comments 2026-08-05 16:21:28 +00:00
Katia Bulatova 16682d565c chore(webapp): the page-context comments say only what the handle doesn't 2026-08-05 16:21:25 +00:00
Katia Bulatova 27a9263f5e chore(webapp): tighten the chat-prose theme comments 2026-08-05 16:21:25 +00:00
Katia Bulatova 8349bf4384 chore(webapp): shorten the agent logo, spinner and scenario-kit comments 2026-08-05 16:21:24 +00:00
Katia Bulatova f428759269 chore(dashboard-agent): drop the M0 verdicts file, inline its reasons at the call sites 2026-08-05 16:21:23 +00:00
Katia Bulatova a4cd475eb7 chore(webapp): drop the unused stale-window import from the watch sweep 2026-08-05 15:41:09 +00:00
Katia Bulatova 1989336e68 fix(dashboard-agent): a dead chain is re-armed on its own cadence, and never stops owing a wake 2026-08-05 15:41:08 +00:00
Katia Bulatova 93c5344df1 test(webapp): the batch check shares one authorization and one report read 2026-08-05 15:41:07 +00:00
Katia Bulatova 5d90eb1eea feat(webapp): the batch check answers a whole environment's watches at once 2026-08-05 15:41:06 +00:00
Katia Bulatova d3d33a6aef chore: oxfmt the generated drizzle snapshot and the presenter test import 2026-08-05 14:52:29 +00:00
Katia Bulatova 37e4c234c2 fix(webapp): the chat name is on the row before the turn settles
The panel reloaded its chat list twice per turn — once on settle, then
again on a 5s timer — purely because the generated name landed after the
first reload and the list would otherwise sit on "New chat".

The name is now written before the client settles instead. Generation
starts in onTurnStart, so it runs alongside the model answering, and is
awaited in onBeforeTurnComplete — the last hook before the turn-complete
chunk closes the stream. onTurnComplete cannot do this: it fires after
that chunk, which is exactly why the write used to miss.

Costs nothing in practice (the cheap title model finishes long before the
answer does) and a failure only loses the generated name. The delayed
reload and its timer are gone, and a test pins the ordering — it fails if
the await is dropped.
2026-08-05 14:52:28 +00:00
Katia Bulatova 4606b94e0f feat(webapp): the watch sweep retires watches that ended a week ago
A terminal watch is read by nothing once its wake has landed — the chip is
gone, dedup only looks at active rows, and the outcome's facts live in the
chat transcript from then on. So the row was pure accumulation.

The sweep now drops terminal rows whose last event was over 7 days ago, in
one guarded statement bounded to 500 per run. Guards: terminal status only,
delivery already settled (a row that still owes a wake is never taken), and
the age measured from the newest timestamp on the row so a late delivery
restarts the clock. Watches only — chats, messages and investigations are
untouched.
2026-08-05 14:52:28 +00:00
Katia Bulatova 0b66440e99 perf(webapp): a finished report is reusable for 90s
A report is ~9 ClickHouse queries, and the callers that dominate its
volume are periodic rather than interactive: every watch tick in an
environment asks for the same health verdict, so one sweep recomputed it
once per watch.

Caches the interpreted view model per (report, environment, period) for
90s — under the tick cadence, well over a burst. In-process and keyed by
environment id, so nothing crosses a tenant. The single-flight map stays:
it is what holds the first callers of a cold key to one load, and a
rejected load is still never cached.
2026-08-05 14:52:27 +00:00
Katia Bulatova 0045062261 perf(webapp): the wake poll sleeps in hidden tabs and no longer runs in lockstep
A background tab has no dot the user can see and nowhere to put a toast,
so it stops asking; becoming visible triggers one immediate catch-up and
restarts the cadence from there. Each delay carries fresh jitter, so many
open tabs stop hitting the same second.
2026-08-05 14:52:27 +00:00
Katia Bulatova 95fa1b64b5 chore(webapp): drop the agent-examples review seeder 2026-08-05 14:43:17 +00:00
Katia Bulatova 75aba7f24c feat(webapp): the scenario kit runs against any local project and environment 2026-08-05 14:43:17 +00:00
Katia Bulatova e740c16184 fix(webapp): give the remix build the same heap headroom as typecheck
The webapp outgrew the default heap on CI runners — both E2E workflows died
in webapp:build with allocation failures.
2026-08-04 23:43:22 +00:00