Run headless agent loops without a finite model-step ceiling unless the caller explicitly supplies --max-turns. Keep finite values validated and preserve the separate Fleet worker budget.
Remove the verifier harness's implicit 100-turn flag so long benchmark rollouts are not silently truncated. Verified with the focused TUI regression, all nine verifier harness tests, cargo fmt, targeted strict Clippy, and diff checking.
The README body already said v0.9.4 while the three eval examples still
passed --harness.version 0.9.1, which would resolve a two-release-old
runtime companion set.
Terminal-Bench latency analysis (FINISH-0.9.4 #52) had to infer
reasoning-token counts from wall time because the exec stream-json had
no per-turn usage: content deltas carry none, and only the terminal
metadata receipt reported cumulative totals.
The engine now emits Event::TurnUsage once per model call (turn-step)
when the provider reported usage for that call, carrying the step's
Usage plus stream wall-clock duration. The chat-completions adapter's
synthetic zeroed MessageStart is explicitly not treated as reported, so
providers that never send usage produce no event instead of fabricated
zeros.
codewhale exec --output-format stream-json maps it to a new additive
turn_usage event: turn (1-based), input_tokens, output_tokens, and
duration_ms always present; reasoning_tokens, prompt_cache_hit/miss/
write_tokens, and reasoning_replay_tokens omitted (never null, never
zero-filled) when the provider does not report them. Field names mirror
the terminal metadata receipt so consumers parse one vocabulary.
Existing event shapes are untouched. Consumers: the TUI ignores the new
engine event (its token surfaces run on cumulative TurnComplete usage),
the fleet ledger maps turn_usage to a Running liveness heartbeat for
thinking-heavy calls, and the verifiers harness whitelist accepts the
new type.
Tests: serialization shape + honest-absence unit tests, a pre-existing
event-tag contract guard, and an end-to-end wiremock integration test
locking both the usage-present shape (and the unchanged metadata->done
terminal contract) and the usage-absent skip.
extensions/vscode was 0.8.53, npm/runtime-sdk 0.8.60, and the
verifiers README claimed v0.9.1 while the workspace is 0.9.4. All
three now read 0.9.4. The release.yml version gate previously checked
only workspace + npm/codewhale, which is how the drift survived
release prep; it now also requires runtime-sdk and vscode package
versions to match the tag.
Evidence: cross-surface-tech-debt-audit-2026-08-03.md TL;DR 'Stale
version strings'; §11.3 version-sweep row.
Gate: node -p require(...).version -> 0.9.4 for both packages.
Apply `npm audit fix --package-lock-only` across npm workspaces to
resolve the 17 open Dependabot alerts (7 high, 10 moderate) on the
v0.9.1 merged tree:
- integrations/feishu-bridge: protobufjs 7.6.4 → 7.6.5
- extensions/vscode: brace-expansion 5.0.6 → 5.0.7, js-yaml 4.2.0 → 4.3.0,
fast-uri 3.1.2 → 3.1.4, linkify-it 5.0.1 → 5.0.2
- web: brace-expansion/js-yaml and other transitive dev deps updated to
patched versions; build, lint, tests, and `check:facts` still pass
- root package-lock: refreshed transitive lockfile metadata
Remaining npm audit findings:
- sharp <0.35.0 (via miniflare/next/wrangler) in web and root: no
non-breaking patch available; miniflare pins sharp 0.34.5. Website is
not deployed for v0.9.1, so exposure is build-time only.
- axios in feishu-bridge lockfile is already resolved to 1.18.1; the
Dependabot alerts appear stale against the current lockfile.
All affected workspace checks pass:
- integrations/feishu-bridge: `npm run check && npm run test` — 19 passed
- extensions/vscode: `npm run check` — compiles
- web: `npm run prebuild && npm run check:facts && npm test && npm run lint
&& npm run build` — green
Refs #4713
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Brings 65 upstream commits (web client #4423, OpenCode Go provider,
Windows/Linux ARM64 releases, xAI device OAuth, K3 route contracts,
tool heartbeat reattach, MCP hot-reload, permission posture cycle,
structured conversation export, PTY terminal tools) into the local
lane carrying the remaining milestone work (#4625, #4651, #4603, #4324,
#4628, #3934, #4632, #4647, #414, #4619, #4636, #2889, #4032, #4412,
#4598, #185, #4641, #4605).
Resolution notes:
- Purged 608 files of pr-4582 revert residue to their upstream state;
re-applied the 25 real local commits on top by hand.
- Kept the unified BashTool (#4625) on top of upstream's shell
hardening (EXEC_SHELL_WAIT_MAX_TIMEOUT_MS preserved).
- Kept the #185 wire-format dispatch on upstream's accessor-style
ReadyRouteCandidate + prepare_model_bound_request sanitization.
- Kept upstream's dispatch architecture (route-before-mutate); the
#4605 enter-lag fix is re-applied separately on this architecture.
- Kept upstream's SkillSource/plugin skill machinery with the #4632
privacy-safe path rendering on the native arm.
- coord.rs holds both upstream's coordination tools and the #4647
decision-record/write-scope types.
- Changelogs rebuilt: upstream's full 0.9.0 and 0.9.1 sections plus
the local 0.9.1 entries; dropped unverified claims.
Python integration package that adapts the Codewhale CLI for the
Verifiers v0.2.1 evaluation harness. Provides structured tool-call
extraction, receipt bounding, and terminal output normalization.
Integrate the underwater TUI, message-first Operate, Fleet and Workflow reliability, expanded model/provider catalog, exact custom-route restoration, docs-first site, localization, packaging, and release metadata for the v0.9.0 candidate.
Harden endpoint-bound credential provenance, approval and goal UX, Fleet attempt fencing and crash recovery, large-workspace mention discovery, Kimi budgeting, and release asset/version gates. Include the stopship Fleet and Workflow fixtures used by release dogfood.
Verified with workspace fmt/check/clippy/tests on Rust 1.88, release-script and npm suites, 18-crate publish dry run, production web build, Docker build check, secret scan, dependency audit, and protected-state hash validation.
Safety checkpoint of multi-agent cutover work (fleet roster core,
whaleflow-js runtime, /fleet roster view, audit fix lanes). Tree may
not compile: the agent-tool profile edit in tools/subagent/mod.rs was
interrupted mid-edit (credit exhaustion). Checkpoint precedes repair.
Harvested from PR #3640 by @pkeging.
Preserves the contributor's deployment and security guide while aligning commands with the files currently shipped in the repository. The harvest replaces stale helper-script references with the actual runtime and npm commands, and adds the approval-timeout variable to the bridge env template.
feat(bridge): support natural-language approval responses
Thanks @pkeging for splitting this into a standalone bridge PR, adding the keyword regression tests, and dogfooding the WeCom bridge flow from mobile.
Users can now reply with Chinese/English approval keywords
(e.g. "允许", "可以", "好", "ok", "yes") instead of
copying the approval_id from the prompt.
- Add isApprovalResponse / isDenyResponse helpers to lib.mjs
- Track latest pending approval per chat in a Map
- Intercept prompt messages that match approval/deny keywords
- Add approvalTimeoutMs config (env CODEWHALE_APPROVAL_TIMEOUT_MS)
- Update approval prompt with natural-language hint
Ref: #2967 (WeCom bridge resilience hardening)
Align the Feishu/Lark and WeCom bridge docs with the branded Telegram bridge and
remove DeepSeek-era residue:
- DEEPSEEK_RUNTIME_TOKEN -> CODEWHALE_RUNTIME_TOKEN, DEEPSEEK_ALLOW_UNLISTED ->
CODEWHALE_ALLOW_UNLISTED, /etc/deepseek -> /etc/codewhale, /ds -> /cw.
- Name the allowlist var explicitly (CODEWHALE_CHAT_ALLOWLIST for Feishu;
WECOM_CHAT_ALLOWLIST for WeCom) and document the first-pairing flow.
- Add a concise trust-boundary note per bridge: the chat platform only sees the
prompts/status/approvals the bridge sends; workspace, shell, and the runtime
HTTP listener stay on localhost behind the runtime token.
All env-var names and the /model and /cw commands verified against the bridge
source (feishu-bridge/src reads CODEWHALE_CHAT_ALLOWLIST/CODEWHALE_ALLOW_UNLISTED
with DEEPSEEK_* compat fallback). Weixin left unchanged (already branded).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Move the duplicated JSON thread-map store into bridge-core with bridge-specific options for message dedupe, Telegram action tokens, and WeCom private state-file modes. Keep thin local subclasses where startup-order tests and bridge exports expect ThreadStore names.
Verification:
- npm --prefix integrations/bridge-core run check && npm --prefix integrations/bridge-core test
- npm --prefix integrations/telegram-bridge run check && npm --prefix integrations/telegram-bridge test
- npm --prefix integrations/feishu-bridge run check && npm --prefix integrations/feishu-bridge test
- npm --prefix integrations/wecom-bridge run check && npm --prefix integrations/wecom-bridge test
- npm --prefix integrations/weixin-bridge run check && npm --prefix integrations/weixin-bridge test
Add integrations/bridge-core with shared pure helpers for env parsing, command parsing/action mapping, message splitting, runtime error compaction, active-turn detection, and preserved chat state.
Re-export or wrap those helpers from Telegram, Feishu, WeCom, and Weixin bridge libs while keeping transport-specific identity, validation, keyboards, and protocol handling local. Add Weixin lib tests so its helper behavior is guarded.
Add a maintainer follow-up for the WeCom bridge harvest: private thread-map permissions, public error reporting back to the chat, malformed SSE tolerance, and tests for the local state store.
Follow-up to PR #3370 by @pkeging.
So a remote-setup-generated Feishu bundle (emits CODEWHALE_*) wires up end-to-end; legacy
DEEPSEEK_* deployments keep working via fallback. .env.example now leads with CODEWHALE_*.
Matches the telegram-bridge convention + the RFC namespace migration (item 1).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When running in a Feishu thread-enabled group (话题群), every bot
response — status messages, approval prompts, streaming progress,
turn results — was sent via the Lark SDK's `create` API which spawns
a new standalone topic. The user sees a cluttered group with orphan
topics for each intermediate bot message.
Root cause: `sendText()` only called `client.im.message.create()`
with a bare `chat_id`, never passing any reply context. The Feishu
`reply` API was completely unused.
Fix (two changes, one site each):
1. **lib.mjs — incomingIdentity()**: expose `parentId`, `rootId`,
`threadId` from the raw Feishu message event so callers can
determine thread context. (Not consumed directly yet, but
available for future use.)
2. **index.mjs**:
- `handleIncomingMessage()`: store the latest incoming
`messageId` as `replyToMessageId` in the per-chat thread store.
- `sendText()`: look up `replyToMessageId` from the thread store;
when present, call `client.im.message.reply()` instead of
`create()`. This keeps ALL bot responses nested under the
original user message inside the same topic.
No config changes needed. New chats automatically start using the
reply path; existing chats without a `replyToMessageId` in the store
fall back to the old `create` behaviour.
/ 修复飞书话题群中 bot 消息新建独立话题的问题。所有回复改为使用 reply API
/ 在原话题内嵌套回复,而非通过 create API 创建新话题。
* docs: v0.8.46 CHANGELOG — platform archives, palette, sub-agents, sandbox, web install, search fixes
Closes#2188
* feat(v0.8.46): quick fixes — palette, model picker Esc, sub-agent sidebar, shell chip, model name casing, CVE bump (#2212)
* fix: bump qs to >=6.15.2 for CVE-2026-8723
Add qs override in feishu-bridge package.json to force transitive
dependency resolution to >=6.15.2, addressing CVE-2026-8723.
Refs: #2198
* fix: Esc in model picker applies last-highlighted choice
Previously Esc reverted to the initial model when the user hadn't
moved the selection. Now Esc always applies the currently highlighted
model and thinking-effort tier, making Esc consistent with Enter.
Also updates the picker footer hint from 'Esc cancel' to 'Esc apply'.
Refs: #2196
* feat: show '⏳ shell running' chip in TUI footer
Adds a footer_shell_chip function that displays a '⏳ shell running'
status chip in the footer's right cluster whenever a foreground shell
command is active via exec_shell. The chip is always visible regardless
of user-configured status items.
Refs: #2194
* feat: auto-collapse finished sub-agents in sidebar
When a sub-agent completes (status = 'done'), its detail lines
(id, steps, duration, progress) are now hidden in the sidebar agents
panel. Only the summary label line is shown, keeping the sidebar
compact. Running agents still show full detail.
Refs: #2195
* feat: refresh Whale dark palette for better contrast
Improve contrast and layer separation in the Whale dark theme:
- Deepen base background for more depth (10,17,32)
- Lighten panel (22,34,56) for clearer distinction from bg
- Lighten elevated surface (36,52,78) for better elevation
- Lighten selection (48,68,100) for clearer selected state
- Boost text hint (138,150,174) and dim (118,130,156) readability
- Brighter border (52,88,145) for better edge definition
- Update tool surface colors for consistency
Refs: #2197
* fix: preserve model name casing in normalize_model_name_for_provider
When the user enters a model name like 'DeepSeek-V4-Flash', the
normalizer was lowercasing it to 'deepseek-v4-flash' via the
canonical_official_deepseek_model_id function. Now the normalizer
preserves the caller's casing when the input already matches a known
model id case-insensitively. Compact aliases like 'deepseek-v4pro'
are still rewritten to 'deepseek-v4-pro'.
Refs: #2109
* feat(web): install download tile with arch detection, SHA256, China mirrors + companion binary fix (#2213)
* fix(web): download both codewhale and codewhale-tui binaries in install snippets
The SNIPPETS map only fetched one binary per platform, causing the
dispatcher to fail with MISSING_COMPANION_BINARY. Every arch now
downloads both codewhale AND codewhale-tui side-by-side.
- macOS/Linux: added second curl + combined chmod/xattr/mv for tui
- Windows: added second Invoke-WebRequest for codewhale-tui.exe
- VERIFY: PowerShell now hashes both binaries; Unix --ignore-missing
covers all present binaries in a single sha256sum pass
* feat(web): add install download tile with arch detection, SHA256, and China mirrors (#2192)
* feat(sandbox/linux): process hardening — PR_SET_DUMPABLE, NO_NEW_PRIVS, RLIMIT_CORE (#2214)
* feat(sandbox/linux): add process hardening module — PR_SET_DUMPABLE, NO_NEW_PRIVS, RLIMIT_CORE (#2183)
* feat(sandbox/linux): seccomp filter + bwrap passthrough
- seccomp: BPF filter whitelisting safe syscalls, denying ptrace/mount/kexec
and other dangerous syscalls. Uses raw BPF instructions via libc prctl to
avoid external dependencies (#2182).
- bwrap: optional bubblewrap passthrough when /usr/bin/bwrap is present
and [sandbox] prefer_bwrap=true in config. Creates read-only rootfs with
write access limited to the working directory (#2184).
- landlock detect_denial extended to recognize seccomp SIGSYS/"Bad system
call" patterns alongside existing Landlock EACCES/EPERM detection.
- SandboxManager gains prefer_bwrap field; set_prefer_bwrap on ShellManager.
- EngineConfig gains prefer_bwrap field, wired through main/ui/runtime_threads.
- Diagnostics now reports bwrap_available and cgroup_version.
- config.example.toml documents the prefer_bwrap key.
Pre-existing clippy fixes picked up in the same build:
- collapsible_if in ui.rs version-check
- cmp_owned in goal.rs test
- consecutive str::replace in normalize_auth_mode
Closes#2182, closes#2184
* docs: add cross-links to issue and PR templates in CONTRIBUTING.md (#2215)
- Link .github/ISSUE_TEMPLATE/bug_report.md and feature_request.md from
the Reporting Issues section
- Link .github/PULL_REQUEST_TEMPLATE.md from the Pull Request Guidelines
section
* feat(release): bundle platform archives with install scripts (#2216)
- Add bundle job to release workflow that creates per-platform archives
(tar.gz for Linux/macOS, .zip for Windows) containing both codewhale
and codewhale-tui binaries plus install scripts
- Create install.bat (Windows) — copies binaries to %USERPROFILE%\bin
- Create install.sh (Unix) — copies binaries to ~/.local/bin
- Windows gets a portable .zip variant without install script
- Release notes updated to promote archives as primary download method
- Individual binaries retained for npm wrapper and scripting
Closes#2193
* fix(web_search): fall back to DuckDuckGo when Bing returns zero results (#2130)
When the configured search provider is Bing and the query returns zero
results (common for technical/compound queries), fall through to the
DuckDuckGo path instead of reporting empty. A provenance message is
surfaced: "Bing returned no results; used DuckDuckGo fallback".
Also adds Security and Code of Conduct cross-links to CONTRIBUTING.md
per the sub-agent renovation (#2203).
* docs: SANDBOX.md threat model + RFCs for persistence and MCP + SandboxExecutor trait
- docs/SANDBOX.md: complete threat model describing each platform's sandbox
(Seatbelt, Landlock, seccomp, process hardening, bwrap, Windows v1).
Covers defense-in-depth layering, config keys, denial detection, limitations.
- docs/rfcs/2189-persistence-sqlite.md: RFC for SQLite migration (drafted by sub-agent)
- docs/rfcs/2190-mcp-modularization.md: RFC for MCP crate split into
protocol/client/server with OAuth support
- crates/tui/src/sandbox/policy.rs: SandboxExecutor trait definition and
SafetyLevel→SandboxPolicyBehavior mapping function with tests
Closes#2180, closes#2186, closes#2189, closes#2190
* feat: sandbox parity tests + remove sub-agent 100-turn cap
- Add sandbox parity tests covering platform detection, denial patterns,
bwrap preference, and policy consistency across modes (#2187)
- Remove arbitrary 100-turn sub-agent cap: DEFAULT_MAX_STEPS changed
from 100 to u32::MAX. Sub-agents now run until they produce a final
text response, are cancelled by the parent, or hit a configured
explicit budget (#2034)
Closes#2187, closes#2034
The bridge only supported a single global model (DEEPSEEK_MODEL env /
default_text_model in config.toml). Users who wanted a different model
for a particular Feishu group had to restart the bridge with a different
env var — impractical and disruptive.
This commit adds per-chat model switching so users in a group chat can
run "/model <name>" to switch the model for all future threads and
turns in that chat, without affecting other chats.
Changes:
- **lib.mjs — commandAction()**: handle "model" command → { kind:
"set_model", modelName }.
- **index.mjs**:
- setChatModel(chatId, modelName): store/clear per-chat model in the
thread store. "/model default" resets to the bridge-level default.
- ensureThread(): read per-chat model from store when creating a
new runtime thread; fall back to config.model.
- runPrompt(): read per-chat model for each turn submission
(independent of the thread-creation model), so a /model change
takes effect on the very next message.
Usage:
/model deepseek-v4-flash — switch this chat to Flash
/model deepseek-v4-pro — switch this chat to Pro
/model default — reset to bridge default
Resolution priority (per-chat): global default < per-chat /model.
/ 新增 /model 命令,支持按飞书群设置独立模型。
/ 群内输入 /model <name> 切换,/model default 恢复全局默认。
When running in a Feishu thread-enabled group (话题群), every bot
response — status messages, approval prompts, streaming progress,
turn results — was sent via the Lark SDK's `create` API which spawns
a new standalone topic. The user sees a cluttered group with orphan
topics for each intermediate bot message.
Root cause: `sendText()` only called `client.im.message.create()`
with a bare `chat_id`, never passing any reply context. The Feishu
`reply` API was completely unused.
Fix (two changes, one site each):
1. **lib.mjs — incomingIdentity()**: expose `parentId`, `rootId`,
`threadId` from the raw Feishu message event so callers can
determine thread context. (Not consumed directly yet, but
available for future use.)
2. **index.mjs**:
- `handleIncomingMessage()`: store the latest incoming
`messageId` as `replyToMessageId` in the per-chat thread store.
- `sendText()`: look up `replyToMessageId` from the thread store;
when present, call `client.im.message.reply()` instead of
`create()`. This keeps ALL bot responses nested under the
original user message inside the same topic.
No config changes needed. New chats automatically start using the
reply path; existing chats without a `replyToMessageId` in the store
fall back to the old `create` behaviour.
/ 修复飞书话题群中 bot 消息新建独立话题的问题。所有回复改为使用 reply API
/ 在原话题内嵌套回复,而非通过 create API 创建新话题。
Add the Feishu/Lark long-connection bridge, Tencent Lighthouse runbooks, CNB mirror guidance, CNB tag release pipeline, and China-friendly update fallback documentation for the v0.8.37 line.