* upgrade @opentelemetry packages to the latest versions
* remove v2 only packages, will be moved to a dedicated repo
* remove more v2 code and run pnpm install
* use the npm yalt package in the webapp
* convert @trigger.dev/core to tshy
* Switch from jest to vitest in @trigger.dev/core
* Fixed core test
* move core-backend code into core subpath export
* convert @trigger.dev/sdk to tshy
* Removed hono
* move core-apps to core/v3/apps, remove core-apps, start converting cli-v3
* Fix up some of the commands
* cli now building and loadable
* using package-json-from-dist to get package version now in core and cli
* dev command WIP
* cleaned up some repetition and structure of the entry point stuff
* bringing back the background worker stuff
* Indexing of the v3 catalog
* getting closer to executing dev runs...
* centralize dev logging using event emitter
* Move indexing to it’s own entry point, simplify code
* dev runs working
* Get instrumentation to work with openai
* debugging achieved internally
* provide worker files as part of the worker creation on the server
* support for cjs and esm javascript
* Fixed timeout
* worker manifest now has the config path
* auto-upgrade config to non-deprecated alternatives
* Adding package preview release
* deployment WIP
* improve the syncEnvVars output and adapt resolveEnvVars
* WIP bun runtime
* WIP bun support
* seed tasks with the machine preset if listed in the config
* deploy run executions WIP, extracted TaskRunProcess into 1 place
* deployed tasks running and executing 🎉
* support for waits and better flushing & process cleanup
* Fixed the heartbeating
* Better warning messages
* Improve and unify the indexing between dev and deploy
* Support for external deps that need node-gyp to build
* build extensions can now install custom packages and run instructions in the image. Also prisma extension now works and also works with multiple schema files
* Add back in the main/types/module to sdk
* dev no longer is Ink/React, grace period for disconnections in dev
* Fix the changeset config
* More changeset fixes
* Remove config packages
* More changeset fixes
* Fixed typescript issues (needed to revert back to zod 3.22.3
* Fix pr_checks workflow
* Remove the prepare script
* Fixed tests and package versions
* Remove cli test script
* Remove packages from tailwind watch paths
* Add repo to public packages
* Just commit the generated files and do the building at dev time
* Try and get pkg.pr.new working
* Try again
* Fix emitDecoratorMetadata importing named export from typescript
* config file backwards compat with export const config
* Fixed issue where import errors weren’t coming through
* p-retry is a prod dep
* typescript needs to be a prod dependency for emitDecoratorMetadata
* Add better debug logging to help track down import-in-the-middle bug
* An external is only considered resolvable if it resolves to the same path as the collected external
* Fix runtime checks to allow >=18.20
* Move extensions to a new build package
* Fixed building packages in dockerfile
* Remove the e2e test from publish workflow for now
* Don’t treat pkg.pr.new versions has needing upgrading
* making sure config handleError works, and discovered path aliases don’t work in config files
* Strip empty string env vars so they accidentally override real values
* Couple of things
* Update version to use preview instead of beta
* Hopefully fix re-attempts with >30s delay
* Match socket emit messages to current latest in main
* Initial guide
* Go back to beta
* Go back to the preview, and update guide to use pr preview tags
* Go back to beta
---------
Co-authored-by: Matt Aitken <matt@mattaitken.com>
* Add new ProjectAlertType ‘TASK_RUN’ and then migrate channels to it
* WIP new alerts for run failures
* Consolidate failed run status into taskStatus.ts
* Added the read replica into BaseService, could be useful
* Send task run alerts code
* Revert "Added the read replica into BaseService, could be useful"
This reverts commit cc2348a40254ea3f9e11c0feaf303ecd8794433a.
* Allow adding task run alerts in the UI
* Use task run alert not attempt alert…
* Use the primary
* Test for checkpoints
* Make sourceTaskAttemptId optional on resumeBatchRun
* Removed all completions/executions logic from the shared queue consumer
* Removed the sourceTaskAttemptId from ResumeBatchRunService
* Revert "Removed all completions/executions logic from the shared queue consumer"
This reverts commit d35398d50463c81a1975bb5d5bcfca66a24b8ede.
* WIP on triggerAndWait…
* Fixed triggerAndWait continuing when a checkpoint completes
* Remove the ResumeAttempt code that fails attempts (was protecting against infinite restores)
* Removed messageBody.data.completedAttemptIds.length === 0 commented out code
* Don’t ack if there’s no batchRun
* Added the marqs?.replaceMessage back in but NOT when there’s no checkpoint. More logging
This is a fix for when some attempts fail
* Improvement to the test task that now randomly fails attempts
* When a checkpoint happens, only continue the attempt if it’s in the correct state
* Changeset for rollback in branch
* Set keepRunAlive to false when the dependent task isn’t finished
* Changeset manual version (to get inline with the hotfix branch)
* Changeset: Fixes for continuing after waits
* Latest lockfile (after manual changeset version)
* Remove the old log truncation
* Added TaskRUn logsDeletedAt column
* Accurate timestamps for the run inspector
* EnsureProperty type when you want to make a single property not nullable
* No logs and upgrade messages working
* Button can be autofocused
* Replay dialog code editor is autofocused
* Fix for wrapping of span duration
* v3 subscription endpoints
* Use pnpm linked billing package during development
* Moved v2 billing components into a subfolder
* Select plan using the real data
* Improved v3 plan display
* Use new api response that doesn’t require a Stripe call
* Free flow is working
* Added GitHub modal and verified badge
* Deleted old request v3 access component/route
* Allow setting classes on the Tooltip button
* Paid plans working
* Redirect from select plan if you’ve got v3 enabled
* Loading state improvements
* New admin API endpoint to set concurrency across multiple environments
* When projects are created, conditionally create staging based on the plan
* New billing page working with side menu and stripe portal
* Layout, formatting and some plan state improvements/fixes
* Don’t show the period if you’re on the free plan
* Temporary upgrade callout
* Refactored the platform code so it’s easier to call and doesn’t require a isManagedCloud check
* Send taskIdentifier to OpenMeter
* Side menu
* Upgrade prompts
* More improvements to the app-wide usage indicators
* Early work on usage graphs
* Added the usage bar for v3
* Moved code to presenter and now using defer
* Added the tasks table to usage
* Improved the v3 usage bar if theres’ no usage on a paid plan
* If no run data, still render a graph
* Usage page errors when defered loading fails
* Don’t show the public API key for v3, they’re not used and probably never will be
* Improved the upgrade callout and API keys page layout
* Only show the “reveal all” toggle if you have environment variables in the table
* Replaced Upgrade callout with a more generic InfoPanel component
* better panel width
* Show conditional upgrade prompts based on plan and number of schedules used
* Removed duplicate class
* Wider blank state panels for the scheduled page
* Wider info panel for the env var page
* Blank state now using the info panel
* Platform alerts prompt now using the InfoPanel
* Deploy blank state uses InfoPanel
* Github verified badge padding adjustment
* Improved the layout of the page, some style tweaks, organized imports
* Better default tooltip style
* could be undefined fix
* text fix + style updates
* Changed the billing icon in the side menu
* Billing page layout and style improvements
* plan tooltips don’t use dark variant
* Don’t highlight the plan on the billing page
* Tooltip underlines stand out more
* Fixed padding in the PageTitle
* Tooltips use the correct cursor
* Improved the plan banner on the billing page
* Fixed Header1 inconsistent font weight
* Fixed issue where input field focus states were being clipped
* Fixed large button not having large text size
* Added a link to the Get in touch copy and improved the connect to GitHub modal
* Select plan page uses the MainCenteredContainer
* Better logging from the Loops endpoint because this error finally got hit
* Move the ingestion of compute to the platform
* Reporting usage of invocations moved to the platform
* Get the entitlement before triggering a non-dev task
* Contact us enterprise plan button opens the feedback form
* Removed Github discussions link from the Feedback panel
* Swapped billing icon for credit card
* Show a Unlock staging panel on the env var page
* Updated staging environment colour
* Show a prompt to upgrade to get staging in the new env var modal
* Improved the edit env var modal
* Implement ability to disable org concurrency
* Use common logic for the plans
* Use the billing server to get the schedule limits
* Some schedules page fixes
* More convenient way of getting a limit
* Use the new schedule limit
* Team member limiting
* Made the limit visible on the team page
* Limit alerts
* Added an index for TaskRun.scheduleId
* Remove console.log on schedules page
* Added durations to the run table
* Tabular numbers
* Improved the usage page formatting
* Only admins see the compute column on the run table
* Include the base cost on the usage stats
* Moved the status to the sidebar
* Optional table header tooltip
* Allow InfoIconTooltips to have customizable content styles
* Added a tooltip to the duration header, changed no test to a dash
* Removed all references to signing up to v3 from the docs
* Switched @trigger.dev/billing to @trigger.dev/platform
* Passing up the variant for the InfoIconTooltip
* table tooltip max-width fixed
* Switch to the published @trigger.dev/platform 1.0.11
* v2 usage page title changed to include “v2"
* code theme has a transparent background so it works on any background
* duration columns now grouped together nicely at wide screen size
* Last duration column fills the width properly
* Fix for the per run price being in cents not dollars
* Show the total cost with 8 decimal places
* Show 8 decimal places in the usage graph tooltip
* Moved the UpgradePrompt to the v3 folder
* Prepare to use Shadcns chart helpers
* Much nicer chart
* Small tweaks to the graph
* Fix run table col spans for empty/loading messages
* We don’t need isManagedCloud in createProject
* Hide v3 usage/billing pages if there aren’t v3 projects in your org
* Removed unused tooltipStyle
* Usage bar now says “Included usage” instead of “Tier limit” if you’re paying
* Get the plan/usage data in parallel
* The usage page now has a month dropdown and all data is for that calendar month
* Ensure the passed date is the 1st of the month
* Use the machine presets from the platform package
---------
Co-authored-by: James Ritchie <james@jamesritchie.co.uk>
Co-authored-by: Eric Allam <eallam@icloud.com>
* v3: cancel subtasks when parent task runs are cancelled
* v3: recover from server rate limiting errors in a more reliable way
- Changing from sliding window to token bucket in the API rate limiter, to help smooth out traffic
- Adding spans to the API Client core & SDK functions
- Added waiting spans when retrying in the API Client
- Retrying in the API Client now respects the x-ratelimit-reset
- Retrying ApiError’s in tasks now respects the x-ratelimit-reset
- Added AbortTaskRunError that when thrown will stop retries
- Added idempotency keys SDK functions and automatically injecting the run ID when inside a task
- Added the ability to configure ApiRequestOptions (retries only for now) globally and on specific calls
- Implement the maxAttempts TaskRunOption (it wasn’t doing anything before)
* Adding some docs about the request options
* Fix type error
* Remove context propagation through graphile jobs
* Remove logger
* only select a subset of task run columns
* limit columns selected in batchTrigger as well
* added idempotency doc
* allow scoped idempotency keys, and fixed an issue with the unique index on BatchTaskRun and TaskRun
* Removed old cancel task run children code
* v3: Trigger delayed runs and reschedule them
* Create a `@trigger.dev/core/v3/schemas` export
* fixed the `@trigger.dev/core/v3/schemas` export
* Small docs tweak
* Add ttl option when triggering tasks, expire runs after ttl
Dev runs expire in 10m by default
* WIP
* Handle tasks that have failed but are being auto yielded
* Limit trace view to 25k event records, add a download run logs button
Also added two new indexes to TaskEvent:
```
/// Used on eventRepository.getTraceSummary()
@@index([traceId, startTime])
// Used for getting all logs for a run
@@index([runId])
```
* perf improvements on eventRepository.getSpan()
* v2: Add a 5 minute timeout for run execution requests in dev
* v3: Include presigned urls for downloading large payloads and outputs when using runs.retrieve
* v3: better handle large task payloads and outputs
* Change to 512KB
* v2: paginate trigger schedules endpoint
* v3: add 3MB limit on batch and single payloads
* Update task payload and output limits
* Starting to measure wall time and cpu time in the workers, and reporting that via otel and to completed task run attempts
* Move usage tracking outside of the executor
* WIP prod usage tracking
* WIP
* WIP custom fetch to openmeter
* Create a usage client
* WIP
* WIP
* Implement new machine preset stuff and send usage reports to OpenMeter from webapp
* WIP
* Expose usage info to the client
* Add usage and cost to TaskEvent
* Add ability to globally configure the task machine preset
* Report start run usage
* Change the machine docs to use presets
* setExpirationTime to 24h
* Removed logs
* Update machines.mdx
* Removed console.logs
* Handle revalidating JWT tokens
* Couple tweaks
---------
Co-authored-by: Matt Aitken <matt@mattaitken.com>
* Added maximumScheduleInstancesLimit column to Org, default to 20
* Docs on the schedule limits and improved soft-limit communication
* Added limit info to the schedules list page
* Created a task that creates schedules, useful for testing
* Make deduplicationKey required when creating/updating a schedule using the SDK
* New schedule button shows an alert if you’re over the limit
* Added timezone to the form and db
* WIP on the timezone dropdown for the create/edit schedule form
* Use the new filter search for timezones
* Made the timezone dropdown faster by fixing the virtualization
* The preview table is working and added a nice message about daylight savings
* Created a page where you can view the full list of timezones
The URL is included in the error message if you send an invalid time using the SDK
* Creating tasks with the timezone
* Added timezone support the the scheduler and the schedules list
* Added timezone support to more of the schedules UI
* The timezone comes through to scheduled runs with nice JSDocs
* Allow setting the timezone from the SDK
* Always have a timezone on a schedule
* Updated jsdocs
* Updated catalog example
* Changed the column to be a string, not null. Added the timezone across the SDK
* API endpoint for getting the timezones
* Added an SDK function to get the list of timezones
* Added timezones to the docs
* Changeset: Added timezone support to schedules
* Added support for testing timezone
* Tidied up imports
* Imports
* Imports
* Update limits.mdx
* Fixed a couple type issues and use the already exported zodfetch
---------
Co-authored-by: Eric Allam <eallam@icloud.com>
* add amin email regex env var
* fix displayed init command for self-hosted setups
* shared env var to disable telemetry in cli and webapp
* pin sdk version during init
* if specified, add api url to dev command shown after init
* improve checkpoint support detection
* control forced checkpoint simulation via env var
* add public init to providers
* better checkpoint support check for coordinator
* add docker to coordinator image
* update docker provider containerfile
* bump remaining containers to node 20
* add infra image build to default publish workflow
* lockfile
* remove concurrency group from infra workflow
* add docker provider to build matrix
* fix var subst
* checkpoint test is docker specific
* enable v3 projects by default on self-hosted instances
* fix v3 setup command again
* add default posthog key
* self-hosting docs
* add latest tags to versioned infra and webapp builds
* some checkpoint errors should skip retrying
* add changeset
* shorten paragraph
* some docs updates
* update tunnelling section
* add registry setup section
* use correct cli push flag
* add checkout to v3 branch
* update the worker machine setup steps
* fix infra build
* small docs update
* remove unused feature function
* Revert "remove unused feature function"
This reverts commit cfe07887a12b6893dca8ce499964481a9b3dc9db.
* fix self-hosted v3 feature gate
* add note about missing arm support
* simplify helper script syntax
* Switch to read replica: getEvent API endpoint
* Switch to read replica: v2 run list presenter
* Switch to read replica: Job presenter
* Switch to read replica: Job list presenter
* Switch to read replica: billing client
* Switch to read replica: OrgUsagePresenter
* Switch to read replica: OrgBillingPlanPresenter
* Switch to read replica: ScheduleListPresenter
* Switch to read replica: EventRepository taskEvent.findMany
* Proof of concept
* When ingesting events, if it’s already been delivered then don’t continue
* DeliverEvent: throw AlreadyDeliveredError and don’t retry if that’s thrown
* Test for duplicate event ids
* Return the original event so sendEvent doesn’t fail, don’t enqueue
* Add AlreadyDeliveredError to the logged out message
---------
Co-authored-by: Eric Allam <eallam@icloud.com>
* WIP
* Allow marqsv2 and v2 graphile to run in parallel
* Fix missing GraphileLogger import
* Fixed heartbeat after rebase
* Replace postgres based run counters with redis ones with a backfill
* Add back in the graphile logger
* Remove duplicate visibility timeout calls
* Clamp simple weighted strategy to max of 5
* Easier to create a rate limiter, use it in the ApiRateLimiter. Upgraded the Upstash package
* Always prefix any rate limiter in Redis with “ratelimit:”
* By default log when the rate limit is hit
* Added rate limiting to IngestSendEvent
* Log out the EventRecord id
* Increase events.deliverScheduled attempts
* INGEST_EVENT_RATE_LIMIT_MAX is optional
* Removed old API rate limit code
* IngestSendEvent rate limiter is optional. Moved outside of the DB transaction
* Log a message out when the rate limiter is created
* Return undefined if the rate limit has been crossed
* WIP worker TaskRunAttempt creation
* Handling failing task runs that cannot create an attempt for whatever reason
* Move the visibility queue stuff into a graphile job
* Fixed task runs with unsanitized queue names
* “Borrow” the code from alerts PR to get self hosted deployments working
* Add an admin API endpoint to get info about the shared marqs queue
* Allow admins to view any project metrics
* start adding lazy attempts to prod
* lazy attempt creation for prod workers
* resurrect prod stack traces
* add exception event to failed run spans
* simplify dependency resumes
* fix typecheck
* fix merge
* fresh process for all attempts
* always try sigterm first
* stop heartbeat timeout on non-inplace replace message
* add missing ack on checkpoint creation service failure
* bypass dequeue for retries with running worker
* respect retry delays
* crash runs with invalid run status for execution
* remove debug logs
* fix nack message
* fix version locking
* fresh attempt processes in dev and prod
* improve handling of ipc timeouts
* consider checkpoint failures on cancellation
* add basic chaos monkey to checkpointer
* changeset
* control forced checkpoint simulation via env var
* fix merge
* kill old attempt processes before checkpointing
* detailed perf logging for checkpointing
* add coordinator otlp endpoint example
* improve prod run cancellation
* rename supports lazy attempts migration
* fix graceful exit
* fix retry mechanics
* clear paused state before retry
* remove checkpoint image after push
* crash worker on unrecoverable errors
* refactor unrecoverable error emit
* switch to do hosted busybox image
* increase wait for duration ipc timeout
* add changeset for misc fixes
* fix merge
* fix retry delay span runId
* fix dev retries
* improve prod worker logging
* log checkpoint sizes
* add lazy attempts catalog entries
* Fixed merge issue: use zodFetch, not wrapZodFetch
* Revert "Fixed merge issue: use zodFetch, not wrapZodFetch"
This reverts commit d137e4e1fe.
* importEnvVars uses wrapZodFetch now
* add backwards compat for retries without checkpoints
* handle more cases of unrecoverable runs
* don't kill the child process if it shouldn't be killed
---------
Co-authored-by: nicktrn <55853254+nicktrn@users.noreply.github.com>
Co-authored-by: Matt Aitken <matt@mattaitken.com>
* WIP env var management API
* Add import env var API endpoint
* Adding docs and support for using both API keys and PATs when interacting with the env var endpoints
* WIP envvar SDK
* Uploading env vars in a variety of formats now works
* Finish env var endpoints and add resolveEnvVars hook
* Add changeset
* Fix: API rate limit error has the correct seconds until reset
* When a v2 run hits the rate limit, reschedule using the reset timestamp
* Still throw AutoYieldRateLimitErrors
* Reschedule runs from the rate limit
* The stress test timeout should be inside the task
* If the rate limit error is thrown, don’t retry the API request
* WIP on multi-select
* WIP on simple checkbox
* CheckboxWIthLabel and Checkbox
* Multi-selection of runs across pages is working
* Fix for selection on seconds page
* Focus the run filter on page load
* Don’t focus the checkbox
* BulkActionBar now shows/hides and has buttons
* Some state to stop escape clearing the selection when the modals are open
* Delete unused formData util
* Improvements to the page
* Created the replay resource action. It doesn’t do anything useful yet.
* Database schema created for BulkActionGroup/BulkActionItem
* The BulkActionService is creating the right data, now we need to process it
* WIP on bulk processing
* Added failed state and made the sourceRun required
* Bulk replaying is working
* WIP on bulk action filtering
* Fixed bulk filters displaying
* Filtering by batch is working
* Some fixes for the bulk id filtering
* Style tweaks
* Load the extra info in parallel
* Bulk canceling working
* Get the most recent 20 bulk actions to display in the filter menu
* Even if the run isn’t cancelable add it to the final list
* Maximum of 250 runs can be bulk actioned
* Don’t let them select more than the maximum (250 currently)
* Separate each bulk item action into it’s own separate graphile job to increase resiliency
---------
Co-authored-by: Eric Allam <eallam@icloud.com>