## Summary
Adds read methods to `RunStore` (`findRun`, `findRunOrThrow`,
`findRuns`) and routes every Postgres read of `TaskRun` through them,
mirroring how writes already go through the store. Behavior-preserving:
each relocated read keeps its exact query, field selection, and database
client (writer, replica, or transaction). This lets `TaskRun` reads be
retargeted to a different backing store later without touching call
sites.
Stacked on #3981 (the write adapter); that PR is the base of this one.
## Scope
In scope: the run engine, webapp services, presenters, and route
loaders. Three reads that pulled `TaskRun` in through a parent model's
relation `include` (alert delivery, batch results, attempt-dependency
cancellation) are decomposed to fetch the run(s) through the store and
stitch them back, since a relation include would not follow `TaskRun` to
a new table.
Left reading the existing table (out of scope): the legacy MarQS paths,
the legacy trigger idempotency read, and one raw-SQL recovery script
(commented for revisiting at cutover).
## Notes
Reads default to the read replica; callers pass the writer or a
transaction client wherever the original read did, so writer-vs-replica
behavior is unchanged.
* Cleanup context and execution creation, cache stuff, add parent and root task run ids
* more efficient by using friendly IDs instead of doing joins
* metadata.root/parent now reference current run when run has no root/parent
* Adding changeset
* try to make test less flaky
* Clean imports
* Another attempt to fix the flaky test
* Fix usage by still passing durationMs and costInCents to the execution, just not the run.ctx
* OOM retrying on larger machines
* Create forty-windows-shop.md
* Update forty-windows-shop.md
* Only retry again if the machine is different from the original
* Various fixes for run engine v1
- Make sure there are connected providers before sending a scheduled attempt message, nack and retry if there are not
- Fail runs that fail task heartbeats when pending and locked
- More and better logging around shared queue consumer
- Fix bug when failing a task run with no attempt
* Prevent findUnique from bringing down our database
* magic links on span event errors
* prevent task monitor from processing errors handled elsewhere
* exclusively use internal error code enum for completion data
* add complete attempt service opts
* reattempts need to go via the queue for task controllers that may have exited
* only infer retry config if completed via crash or system failure
* enhance error before deciding if retriable
* retry on SIGTERM
* enable retry config helper for latest sdk
* don't retry heartbeat timeouts for now
* enable task monitor to update fatal errors
* add missing service
* update retry config since package version
* don't alter completion time when updating existing error
* refactor finalize run service
* refactor complete attempt service
* remove separate graceful exit handling
* refactor task status helpers
* clearly separate statuses in prisma schema
* all non-final statuses should be failable
* new import payload error code
* store default retry config if none set on task
* failed run service now respects retries
* fix merged task retry config indexing
* some errors should never be retried
* finalize run service takes care of acks now
* execution payload helper now with single object arg
* internal error code enum export
* unify failed and crashed run retries
* Prevent uncaught socket ack exceptions (#1415)
* catch all the remaining socket acks that could possibly throw
* wrap the remaining handlers in try catch
* New onboarding question (#1404)
* Updated “Twitter” to be “X (Twitter)”
* added Textarea to storybook
* Updated textarea styling to match input field
* WIP adding new text field to org creation page
* Added description to field
* Submit feedback to Plain when an org signs up
* Formatting improvement
* type improvement
* removed userId
* Moved submitting to Plain into its own file
* Change orgName with name
* use sendToPlain function for the help & feedback email form
* use name not orgName
* import cleanup
* Downgrading plan form uses sendToPlain
* Get the userId from requireUser only
* Added whitespace-pre-wrap to the message property on the run page
* use requireUserId
* Removed old Plain submit code
* Added a new Context page for the docs (#1416)
* Added a new context page with task context properties
* Removed code comments
* Added more crosslinks
* Fix updating many environment variables at once (#1413)
* Move code example to the side menu
* New docs example for creating a HN email summary
* doc: add instructions to create new reference project and run it locally (#1417)
* doc: add instructions to create new reference project and run it locally
* doc: Add instruction for running tunnel
* minor language improvement
* Fix several restore and resume bugs (#1418)
* try to correct resume messages with missing checkpoint
* prevent creating checkpoints for outdated task waits
* prevent creating checkpoints for outdated batch waits
* use heartbeats to check for and clean up any leftover containers
* lint
* improve exec logging
* improve resume attempt logs
* fix for resuming parents of canceled child runs
* separate SIGTERM from maybe OOM errors
* pretty errors can have magic dashboard links
* prevent uncancellable checkpoints
* simplify task run error code enum export
* grab the last, not the first child run
* Revert "prevent creating checkpoints for outdated batch waits"
This reverts commit f2b5c2ac42.
* Revert "grab the last, not the first child run"
This reverts commit 89ec5c8bfd.
* Revert "prevent creating checkpoints for outdated task waits"
This reverts commit 11066b4e74.
* more logs for resume message handling
* add magic error link comment
* add changeset
* chore: Update version for release (#1410)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Release 3.0.13
* capture ffmpeg oom errors
* respect maxAttempts=1 when failing before first attempt creation
* request worker exit on fatal errors
* fix error code merge
* add new error code to should retry
* pretty segfault errors
* pretty internal errors for attempt spans
* decrease oom false positives
* fix timeline event color for failed runs
* auto-retry packet import and export
* add sdk version check and complete event while completing attempt
* all internal errors become crashes by default
* use pretty error helpers exclusively
* error to debug log
* zodfetch fixes
* rename import payload to task input error
* fix true non-zero exit error display
* fix retry config parsing
* correctly mark crashes as crashed
* add changeset
* remove non-zero exit comment
* pretend we don't support default default retry configs yet
---------
Co-authored-by: James Ritchie <james@trigger.dev>
Co-authored-by: shubham yadav <126192924+yadavshubham01@users.noreply.github.com>
Co-authored-by: Tarun Pratap Singh <101409098+Wackyator@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* When resuming a batch, only do marqs operations once
* Made TaskRunDependency clearer in the Prisma schema
* New ResumeDependentParentsService service, use it from checkpoints
* WIP on making resuming more robust
* Turn the declarative schedules off because they make debugging other runs painful
* Resuming batches when there’s an attempt is working
* If there’s no attempt then create one
* Added a log if there are no span events to complete
* If Graphile addJob doesn’t return a row, log and return undefined. No throw
* Pass prisma into the ResumeDependentParentsService
* Removed the todos
* Pass Prisma through to the checkpoint service
* Fix for not checking the batch item correctly
* Fix for when a log flush times out and the process is checkpointed
* Fix for when a log flush times out and the process is checkpointed
* Another test run that does batches with failed subtasks
* Don’t call ResumeTaskRunDependenciesService anymore (we have a new service)
* Only resume if the run is in a final state
* If an attempt doesn’t exist, fix for creating queue with sanitized name
* If DEV then don’t resume using marqs/batches. The CLI manages it
* We don’t need to check the run status again, it’s in the main function now
* Added TaskRunAttempt taskRunId index
* Only allow calling ResumeDependentParentsService with a run ID
* Put the flushing back to what it was
* Add an error to the final attempt if there isn’t one
* Improved the checkpointResumer test task
* Using triggerAndWait or batchTriggerAndWait frees up concurrency
Normally it frees up just env and org concurrency. If it’s a recursive task then it will free up the run concurrency too (e.g. a task calling itself).
* Removed the filepath and export name from attempt spans
* Run inspector: only show an error if the run is in a finished state
* support named capture groups
* write crash errors to attempt.error
* make restored pod names unique per checkpoint
* use last eight characters of checkpoint id instead
* add more chaos monkey env vars
* Ignore unfreezable states
* prevent excessive queue config parsing errors
* handle dependency resume edge case
* better entry point logging
* ignore checkpoint cancellation timeouts
* add missing idempotency keys to wait for dep replays
* remove checkpoints between attempts
* fix retry container names on kubernetes
* add changeset
* fix types
* bring back internal duration timers
* WIP notes on each location where we’ll use finalize
* Initial FinalizeTaskRunService
* ExpireEnqueuedRunService uses FinalizeTaskRunService
* FailedTaskRunService uses FinalizeTaskRunService
* Allow passing in an include when finalizing the run
* CrashTaskRunService using FinalizeTaskRunService
* Remove comments
* Status is optional
* CancelAttemptService using FinalizeTaskRunService
* Import tidy
* CancelTaskRunService using FinalizeTaskRunService
* Import tidying
* CompleteAttemptService system failure switched to FinalizeTaskRunService
* Added more logging to Finalizing
* CompleteAttemptStatus COMPLETED_SUCCESSFULLY
* CompletedAttempt “SYSTEM_FAILURE”
* CompletedService final pair
* Use satisfies so we can derive types from the groups
* Only allow final states to be used with this service
* BaseService tx support, minor improvements
* Added TaskRun completedAt column
* When finalising a run set the completedAt date
* A note to discuss whether we need to set the completedAt to null
* Remove the note in the sharedQueueConsumer
* WIP worker TaskRunAttempt creation
* Handling failing task runs that cannot create an attempt for whatever reason
* Move the visibility queue stuff into a graphile job
* Fixed task runs with unsanitized queue names
* “Borrow” the code from alerts PR to get self hosted deployments working
* Add an admin API endpoint to get info about the shared marqs queue
* Allow admins to view any project metrics
* start adding lazy attempts to prod
* lazy attempt creation for prod workers
* resurrect prod stack traces
* add exception event to failed run spans
* simplify dependency resumes
* fix typecheck
* fix merge
* fresh process for all attempts
* always try sigterm first
* stop heartbeat timeout on non-inplace replace message
* add missing ack on checkpoint creation service failure
* bypass dequeue for retries with running worker
* respect retry delays
* crash runs with invalid run status for execution
* remove debug logs
* fix nack message
* fix version locking
* fresh attempt processes in dev and prod
* improve handling of ipc timeouts
* consider checkpoint failures on cancellation
* add basic chaos monkey to checkpointer
* changeset
* control forced checkpoint simulation via env var
* fix merge
* kill old attempt processes before checkpointing
* detailed perf logging for checkpointing
* add coordinator otlp endpoint example
* improve prod run cancellation
* rename supports lazy attempts migration
* fix graceful exit
* fix retry mechanics
* clear paused state before retry
* remove checkpoint image after push
* crash worker on unrecoverable errors
* refactor unrecoverable error emit
* switch to do hosted busybox image
* increase wait for duration ipc timeout
* add changeset for misc fixes
* fix merge
* fix retry delay span runId
* fix dev retries
* improve prod worker logging
* log checkpoint sizes
* add lazy attempts catalog entries
* Fixed merge issue: use zodFetch, not wrapZodFetch
* Revert "Fixed merge issue: use zodFetch, not wrapZodFetch"
This reverts commit d137e4e1fe.
* importEnvVars uses wrapZodFetch now
* add backwards compat for retries without checkpoints
* handle more cases of unrecoverable runs
* don't kill the child process if it shouldn't be killed
---------
Co-authored-by: nicktrn <55853254+nicktrn@users.noreply.github.com>
Co-authored-by: Matt Aitken <matt@mattaitken.com>