Files
Katia Bulatova 4569657923 feat(webapp): dashboard agent — chat, reports, investigate (#4418)
## What & why

This is the system behind the Dashboard Agent — an assistant that
answers questions about a project's runs, errors, queues, deploys and
health, and can investigate failures end to end.

The agent runs as a chat.agent task in its own Trigger project. It has
no access to the main database or ClickHouse; all platform data is read
through the public API using a delegated, read-only user token.

Everything here is behind `canAccessDashboardAgent` and inert with the
flag off. The UI that mounts the panel lands in #4529.

## Stack

`#4418` (this, base) ← `#4529` UI ← `#4525` Watch ← `#4516` storybook
gallery. The scenario/contract reference for the whole stack is
`internal-packages/dashboard-agent/GUIDEBOOK.md` (it lands on the Watch
branch): it states, per feature, what makes each thing happen and where
that is decided.

## What's inside

**Agent runtime and tools** — `internal-packages/dashboard-agent`:
prompt, tool set (API reads, TRQL query, docs, navigation,
evidence/investigations, repo source), conversation compaction, a
prompt-prefix token budget pinned by snapshot test, and sampled
LLM-judged turn evals. The package cannot import webapp server code,
which is what makes the "no DB access" claim structural rather than a
convention.

**Contracts** — `internal-packages/dashboard-agent-contracts`:
`trigger://` URIs, intents, and the block envelope every rendered card
travels in.

**Conversation store** — `internal-packages/dashboard-agent-db`: drizzle
over postgres-js in its own `trigger_dashboard_agent` Postgres schema,
plus one additive migration.

**Auth boundary** — the user-actor token gains an optional environment
claim; one guard (`userActorEnvironment.server.ts`) enforces it so
routes don't each re-derive the rule. Token minting, cap ceiling, and
the RBAC fallback path for self-hosted.

**Transport** — webapp resource routes that mint the token and proxy
each turn, and SDK-side mid-turn reconnect.

**Public API the agent reads through** — orgs, projects, environments,
runs, queue metrics, workers, a run's commit metadata, repo snapshot,
reports, and `POST /api/v1/query`.

**Reports** — the health report's layout is declared once and shared by
the card, the markdown surface and the JSON/MCP surface, so the same
report reads the same in the dashboard, the terminal and an editor.

**Block renderers** — the report and investigation cards the flows above
already emit (`app/components/dashboard-agent/`). The panel that hosts
them, and the rest of the chat UI, is #4529.

**Query safety and CSP** — see below.

## Key decisions

- **The agent is a separate Trigger project, not webapp code.** It reads
platform data over the public API with a delegated user-actor token
whose `cap` ceilings it to read scopes. No Prisma, no ClickHouse, no
webapp imports.
- **The PAT-only auth helper now refuses user-actor tokens.** This is an
intentional behavioral change: its callers consume only a bare userId
and do not enforce delegated-token capabilities. Actor-aware routes
continue through the scoped route builders instead.
- **RBAC fallback builds a delegated token's ability from its own cap**,
never the blanket ability a PAT gets (read-only when the token declares
none). Without this, the agent's read-only cap would buy a write JWT on
self-hosted.
- **Org creation checks RBAC only for user-actor tokens, and only after
the env gate**, so an install with `ORG_CREATION_API_ENABLED` off
returns 404 rather than 403, and an ordinary PAT never consults an
ability the route has no org to scope. Both orderings are pinned by
test.
- **The query path is read-only in depth.** TRQL rejects write
statements at the grammar level (they don't parse, rather than being
filtered), ClickHouse runs with `readonly=1`, and the org/project/env
filters are injected server-side from the credential — the request body
cannot widen scope. An unparseable query denies instead of falling
through to the permissive resource.
- **Document-wide img-src CSP.** Remote images are an
outbound-request/exfiltration surface, so the policy permits only
own-origin/data/blob, the required SSO avatar hosts, and the favicon
endpoint. Operators can add exact origins through CSP_IMG_SRC_ALLOWLIST;
wildcard hosts and bare schemes are intentionally not allowed.
- **The chat transport reconnects on a mid-turn EOF**
(`@trigger.dev/sdk`). A body that ends without a turn-complete is
terminal only when the server says `X-Session-Settled: true`; otherwise
the transport resubscribes from `lastEventId` with bounded backoff, and
any record re-earns the budget. Previously a closed long-poll window or
a proxy restart left the reply stuck as if still generating.
- **Conversations live in their own datastore**, schema-scoped and
foreign-key-free (it references `organizationId`/`userId` by id, because
in cloud it is a different database). It is a display read-model for the
History tab and transport resume; `chat.agent`'s object-store snapshot
remains the model's source of truth.
- **Deterministic first.** Reports and health checks contain no LLM —
they are computed from the same data the dashboard shows, and the model
only narrates and links them. That is what makes a number in an answer
auditable.

## Testing

- 63 new test files, run with `pnpm run test --filter webapp` and
per-package vitest. Heaviest coverage on the auth boundary
(`userActorPatOnlyBoundary`, `userActorTokenClaimsAndScopes`,
`contextlessPatRoutes`, `rbacFallbackBranch`), TRQL read-only, the
report layout, and the SDK reconnect.
- The agent package has a separate eval lane (`pnpm run test:evals`,
`vitest.eval.config.ts`) that hits the real model, so it never runs in
`pnpm test`.
- Live-tested against a local stack scenario by scenario; the GUIDEBOOK
lists the condition each behaviour is expected under, which is what
those runs were checked against.

## Changelog

`.server-changes/dashboard-agent.md`, plus changesets for
`@trigger.dev/core` (report schemas), `@trigger.dev/sdk` (chat
reconnect) and the CLI's `mint-token` help text.
2026-08-11 18:56:14 +02:00

5.3 KiB

@internal/dashboard-agent-db

The conversation datastore for the in-dashboard agent, isolated from the main Prisma database. Drizzle (postgres-js) over a dedicated trigger_dashboard_agent Postgres schema.

  • Cloud: a separate PlanetScale Postgres database. The app connects over a pooled connection (DASHBOARD_AGENT_DATABASE_URL); migrations run over a direct (non-pooler) connection (DASHBOARD_AGENT_DIRECT_URL), since a transaction-mode pooler can't run the migrator.
  • OSS / self-host: falls back to the main DATABASE_URL (and DIRECT_URL for migrations); the tables live in the dedicated trigger_dashboard_agent schema, isolated from Prisma's public.

The schema is foreign-key-free — it references main entities (organizationId, userId) by id only, because in cloud it lives in a different database.

Why a separate store

The agent runs as an ephemeral Trigger task and must have no access to the main database or ClickHouse (those go through the API). This is its own low-blast-radius store: the agent connects directly here to persist conversations, and the webapp connects here for the History tab. Conversation history correctness is owned by chat.agent's built-in object-store snapshot — this DB is a display read-model (list chats, render a past chat, resume the transport), never the model's source of truth.

Tables

  • chats — one row per conversation: org/user scope, title, metadata (the project/env context the chat ran in), and next_message_position, the allocator the transcript's ordering comes from. No transcript of its own. Soft-deleted via deleted_at, pinned via pinned_at, read-marked via last_read_at (NULL = never read, so every watch wake in it counts as unread).

  • chat_messages — the transcript, one row per message. Identity is (chat_id, message_id) and order is position, unique per chat and reserved from chats.next_message_position by the same single statement that reads it, so concurrent writers get disjoint contiguous ranges. role is lifted out of the payload so the message-quota count is an index scan.

    Three write modes, and only the third may change a message the chat already holds: a new message is a plain insert; a redelivered durable event (a watch wake, a settlement card) is ON CONFLICT DO NOTHING on (chat_id, message_id), so it leaves the recorded row untouched; a deliberate finalisation is finalizeChatMessage, which rewrites one body under a verified role and never moves the id or the position. So re-sending a whole turn snapshot is a no-op.

    Positions are monotonic, not gapless: a reservation whose insert then conflicts, or a batch that rolls back, leaves the slot unused. Only the relative order matters, so a gap is expected and harmless.

  • chat_sessions — live transport state keyed by chat_id: the session-scoped public_access_token and last_event_id for resume. Separate table so the secret token is isolated from list queries and the hot per-turn write stays off the conversation row's indexes.

  • chat_turn_evals — one row per judged turn, written by the dashboard-agent-eval-turn task: quality scores (grounded / answered / concise) and insight classification (intent, outcome, capability & docs gaps). Keyed on (chat_id, turn) so a re-delivered turn can't double-insert. A row holds the judge's derived verdict only — never the user's question, the agent's answer or any tool data. What is judged and what a row may carry is one file: @internal/dashboard-agent/src/eval-policy.ts. Rows are retired after 30 days by the webapp's dashboard-agent sweep. user_text and judge are legacy columns nothing writes any more.

  • investigations — the agent's revisioned working state for a diagnostic thread. Keyed by investigation_id so a follow-up can load one from the id alone; revision is bumped by a single atomic revision = revision + 1 update, and the chat_id/project_ref/environment_ref triple must match on every commit. state is intentionally untyped JSONB — the payload shape isn't frozen yet.

  • watches — "tell me when X happens", checked by a periodic task. status (active | fired | expired | cancelled) and delivery_status (not_required | pending | delivering | delivered) are guarded in the query layer with WHERE status = 'active' … RETURNING, so concurrent fire/expire/cancel resolves to one winner. The org/project/env/user identity is a snapshot taken at creation and never updated — a watch fires with exactly the access its creator had. identity is the dedup key for the watched thing: a partial unique index on (chat_id, project_id, environment_id, identity) WHERE status = 'active' is what actually prevents duplicates, since a read-then-insert check can't be race-proof. A chat may hold at most three active watches, enforced by counting and inserting in one transaction under a per-chat advisory lock.

Migrations

pnpm run db:generate   # generate SQL migration from src/schema.ts (offline)
pnpm run db:migrate    # apply migrations (direct url: DASHBOARD_AGENT_DIRECT_URL, falling back to DASHBOARD_AGENT_DATABASE_URL / DIRECT_URL / DATABASE_URL)

drizzle-kit is scoped to the trigger_dashboard_agent schema (schemaFilter), so pointing it at the main OSS database never touches Prisma's tables.