Files
Katia Bulatova 4569657923 feat(webapp): dashboard agent — chat, reports, investigate (#4418)
## What & why

This is the system behind the Dashboard Agent — an assistant that
answers questions about a project's runs, errors, queues, deploys and
health, and can investigate failures end to end.

The agent runs as a chat.agent task in its own Trigger project. It has
no access to the main database or ClickHouse; all platform data is read
through the public API using a delegated, read-only user token.

Everything here is behind `canAccessDashboardAgent` and inert with the
flag off. The UI that mounts the panel lands in #4529.

## Stack

`#4418` (this, base) ← `#4529` UI ← `#4525` Watch ← `#4516` storybook
gallery. The scenario/contract reference for the whole stack is
`internal-packages/dashboard-agent/GUIDEBOOK.md` (it lands on the Watch
branch): it states, per feature, what makes each thing happen and where
that is decided.

## What's inside

**Agent runtime and tools** — `internal-packages/dashboard-agent`:
prompt, tool set (API reads, TRQL query, docs, navigation,
evidence/investigations, repo source), conversation compaction, a
prompt-prefix token budget pinned by snapshot test, and sampled
LLM-judged turn evals. The package cannot import webapp server code,
which is what makes the "no DB access" claim structural rather than a
convention.

**Contracts** — `internal-packages/dashboard-agent-contracts`:
`trigger://` URIs, intents, and the block envelope every rendered card
travels in.

**Conversation store** — `internal-packages/dashboard-agent-db`: drizzle
over postgres-js in its own `trigger_dashboard_agent` Postgres schema,
plus one additive migration.

**Auth boundary** — the user-actor token gains an optional environment
claim; one guard (`userActorEnvironment.server.ts`) enforces it so
routes don't each re-derive the rule. Token minting, cap ceiling, and
the RBAC fallback path for self-hosted.

**Transport** — webapp resource routes that mint the token and proxy
each turn, and SDK-side mid-turn reconnect.

**Public API the agent reads through** — orgs, projects, environments,
runs, queue metrics, workers, a run's commit metadata, repo snapshot,
reports, and `POST /api/v1/query`.

**Reports** — the health report's layout is declared once and shared by
the card, the markdown surface and the JSON/MCP surface, so the same
report reads the same in the dashboard, the terminal and an editor.

**Block renderers** — the report and investigation cards the flows above
already emit (`app/components/dashboard-agent/`). The panel that hosts
them, and the rest of the chat UI, is #4529.

**Query safety and CSP** — see below.

## Key decisions

- **The agent is a separate Trigger project, not webapp code.** It reads
platform data over the public API with a delegated user-actor token
whose `cap` ceilings it to read scopes. No Prisma, no ClickHouse, no
webapp imports.
- **The PAT-only auth helper now refuses user-actor tokens.** This is an
intentional behavioral change: its callers consume only a bare userId
and do not enforce delegated-token capabilities. Actor-aware routes
continue through the scoped route builders instead.
- **RBAC fallback builds a delegated token's ability from its own cap**,
never the blanket ability a PAT gets (read-only when the token declares
none). Without this, the agent's read-only cap would buy a write JWT on
self-hosted.
- **Org creation checks RBAC only for user-actor tokens, and only after
the env gate**, so an install with `ORG_CREATION_API_ENABLED` off
returns 404 rather than 403, and an ordinary PAT never consults an
ability the route has no org to scope. Both orderings are pinned by
test.
- **The query path is read-only in depth.** TRQL rejects write
statements at the grammar level (they don't parse, rather than being
filtered), ClickHouse runs with `readonly=1`, and the org/project/env
filters are injected server-side from the credential — the request body
cannot widen scope. An unparseable query denies instead of falling
through to the permissive resource.
- **Document-wide img-src CSP.** Remote images are an
outbound-request/exfiltration surface, so the policy permits only
own-origin/data/blob, the required SSO avatar hosts, and the favicon
endpoint. Operators can add exact origins through CSP_IMG_SRC_ALLOWLIST;
wildcard hosts and bare schemes are intentionally not allowed.
- **The chat transport reconnects on a mid-turn EOF**
(`@trigger.dev/sdk`). A body that ends without a turn-complete is
terminal only when the server says `X-Session-Settled: true`; otherwise
the transport resubscribes from `lastEventId` with bounded backoff, and
any record re-earns the budget. Previously a closed long-poll window or
a proxy restart left the reply stuck as if still generating.
- **Conversations live in their own datastore**, schema-scoped and
foreign-key-free (it references `organizationId`/`userId` by id, because
in cloud it is a different database). It is a display read-model for the
History tab and transport resume; `chat.agent`'s object-store snapshot
remains the model's source of truth.
- **Deterministic first.** Reports and health checks contain no LLM —
they are computed from the same data the dashboard shows, and the model
only narrates and links them. That is what makes a number in an answer
auditable.

## Testing

- 63 new test files, run with `pnpm run test --filter webapp` and
per-package vitest. Heaviest coverage on the auth boundary
(`userActorPatOnlyBoundary`, `userActorTokenClaimsAndScopes`,
`contextlessPatRoutes`, `rbacFallbackBranch`), TRQL read-only, the
report layout, and the SDK reconnect.
- The agent package has a separate eval lane (`pnpm run test:evals`,
`vitest.eval.config.ts`) that hits the real model, so it never runs in
`pnpm test`.
- Live-tested against a local stack scenario by scenario; the GUIDEBOOK
lists the condition each behaviour is expected under, which is what
those runs were checked against.

## Changelog

`.server-changes/dashboard-agent.md`, plus changesets for
`@trigger.dev/core` (report schemas), `@trigger.dev/sdk` (chat
reconnect) and the CLI's `mint-token` help text.
2026-08-11 18:56:14 +02:00

117 lines
4.3 KiB
TypeScript

import {
parseTSQLSelect,
SyntaxError as TSQLSyntaxError,
type Field,
type JoinExpr,
type SelectQuery,
type SelectSetQuery,
} from "@internal/tsql";
/**
* Extract every known table a TRQL query reads — the FROM table, every JOIN in
* the chain, and any subqueries — for per-table JWT-scope authorization.
*
* `allowedTableNames` is the set of recognised table names (matched
* case-insensitively); anything not in it is ignored. Injected so this stays
* dependency-free (the caller derives it from the query schemas).
*
* Returns `null` when the query can't be parsed; callers MUST treat `null` as
* deny-by-default.
*/
export function detectQueryTables(query: string, allowedTableNames: Set<string>): string[] | null {
let ast: SelectQuery | SelectSetQuery;
try {
ast = parseTSQLSelect(query);
} catch (err) {
if (err instanceof TSQLSyntaxError) return null;
throw err;
}
const allowed = new Map(Array.from(allowedTableNames, (n) => [n.toLowerCase(), n]));
const seen = new Set<string>();
const scanned = new WeakSet<object>();
function visitSelect(q: SelectQuery): void {
// CTE bodies: `WITH r AS (SELECT ... FROM <table>) ...` — the table is
// read by the CTE even when the outer query only references the CTE alias.
if (q.ctes) {
for (const cte of Object.values(q.ctes)) {
scanForSubqueries(cte.expr);
}
}
// FROM / JOIN chain (tables + FROM-position subqueries).
if (q.select_from) visitJoin(q.select_from);
// Subqueries anywhere else (WHERE, SELECT list, GROUP BY, ORDER BY, etc.)
// can each embed a SELECT that reads a real table, e.g.
// `WHERE id IN (SELECT … FROM runs)`.
scanForSubqueries(q.select);
scanForSubqueries(q.where);
scanForSubqueries(q.prewhere);
scanForSubqueries(q.having);
scanForSubqueries(q.group_by);
scanForSubqueries(q.array_join_list);
scanForSubqueries(q.order_by);
scanForSubqueries(q.limit);
scanForSubqueries(q.offset);
scanForSubqueries(q.limit_by);
scanForSubqueries(q.window_exprs);
}
// Shape-agnostic walk of an expression subtree: descends every nested
// object/array and hands any embedded SELECT to the query visitors, so a new
// node shape can't silently reintroduce a detection gap. The WeakSet guards
// against back-reference cycles the AST might carry.
function scanForSubqueries(node: unknown): void {
if (node === null || typeof node !== "object") return;
if (scanned.has(node)) return;
scanned.add(node);
if (Array.isArray(node)) {
for (const item of node) scanForSubqueries(item);
return;
}
const expressionType = (node as { expression_type?: string }).expression_type;
if (expressionType === "select_query") {
visitSelect(node as SelectQuery);
return;
}
if (expressionType === "select_set_query") {
visitSelectSet(node as SelectSetQuery);
return;
}
for (const value of Object.values(node)) scanForSubqueries(value);
}
function visitSelectSet(q: SelectSetQuery): void {
visitAny(q.initial_select_query);
for (const node of q.subsequent_select_queries ?? []) {
visitAny(node.select_query);
}
}
function visitAny(q: SelectQuery | SelectSetQuery): void {
if (q.expression_type === "select_query") visitSelect(q);
else visitSelectSet(q);
}
function visitJoin(node: JoinExpr): void {
const tableExpr = node.table;
if (tableExpr) {
if ((tableExpr as Field).expression_type === "field") {
const name = (tableExpr as Field).chain[0];
const canonicalName =
typeof name === "string" ? allowed.get(name.toLowerCase()) : undefined;
if (canonicalName) seen.add(canonicalName);
} else if ((tableExpr as SelectQuery).expression_type === "select_query") {
visitSelect(tableExpr as SelectQuery);
} else if ((tableExpr as SelectSetQuery).expression_type === "select_set_query") {
visitSelectSet(tableExpr as SelectSetQuery);
}
}
// The `ON expr` can embed a SELECT that reads a real table, e.g.
// `JOIN x ON id IN (SELECT … FROM runs)`.
scanForSubqueries(node.constraint);
if (node.next_join) visitJoin(node.next_join);
}
if (ast.expression_type === "select_set_query") visitSelectSet(ast);
else visitSelect(ast);
return Array.from(seen);
}