Compare commits

...

88 Commits

Author SHA1 Message Date
jakevin 3076c12d6c chore: bump version to 1.7.4 (#1045)
Release / release (push) Has been cancelled
2026-04-15 15:50:30 +08:00
Howard 44147e54c1 feat(youtube): add feed, history, watch-later, subscriptions, playlist, like, unlike, subscribe, unsubscribe (#1029)
* feat(youtube): add feed, history, watch-later, subscriptions, playlist, like, unlike, subscribe, unsubscribe

* fix(youtube): normalize subscriptions channel fields

* docs(skills): add youtube command coverage

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-15 12:43:27 +08:00
槑囿脑袋 677e37b7a4 feat(xiaoyuzhou): add episode download and transcript support (#1031)
* feat(xiaoyuzhou): add episode audio download

* feat(xiaoyuzhou): add transcript download support

* docs(xiaoyuzhou): clarify credential file requirement

* fix(xiaoyuzhou): remove env credential fallback
2026-04-15 12:35:27 +08:00
Harvey Yue d48c71b993 feat(binance): depth shows both bids and asks (#1019)
* feat(binance): depth shows both bids and asks

* test(pipeline): cover root data access after inline select

* fix(binance): preserve map select context and register manifest entries

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-15 12:30:31 +08:00
DavidDuang 6fbeda951e feat: add hot stock ranking adapters for eastmoney, tdx, ths (#1025)
* feat: add hot stock ranking adapters for eastmoney, tdx, ths

Add three new site adapters for Chinese stock hot rankings:
- eastmoney/hot-rank: 东方财富热股榜
- tdx/hot-rank: 通达信热搜榜
- ths/hot-rank: 同花顺热股榜

All use Strategy.COOKIE browser mode with page.evaluate() DOM scraping.
Each includes co-located tests (13 tests total, all passing).

* fix(tdx,ths): add symbol validation and deduplication in evaluate()

Add seen Set for deduplication and skip entries with empty symbol/name,
matching the pattern already used in eastmoney/hot-rank.js.

* fix: refine hot-rank selectors based on browser inspection

- eastmoney: use table.rank_table tbody tr with td index-based extraction,
  fix name from a[title] to avoid post content contamination
- tdx: use div.top-cell[data-code] data attributes for reliable extraction,
  add tags column from div.tips-item.gnbk
- ths: use card-based layout selectors, remove price column (not in UI),
  extract tags from div.tag.PFSC-R

* fix(hot-rank): align tdx and ths columns with actual output

* fix: register hot stock ranking adapters

---------

Co-authored-by: dengjingren <dengjingren@cn.wilmar-intl.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-15 12:24:41 +08:00
jakevin 9bcdaaa0be fix(external): use safe npm install for dws (#1033) 2026-04-15 12:12:13 +08:00
zhengyu db70a3aaf3 fix(deamon&extension): preserve network capture and surface extension mismatch diagnostics (#1030)
* fix: preserve network capture and surface extension mismatch diagnostics

Older Browser Bridge installs can still connect to the daemon while
missing two capabilities we now rely on: the network-capture actions
and the extension version handshake. That created three user-facing
failure modes with real impact:

1. `opencli explore ...` crashed with `Unknown action: network-capture-start`
   against an old extension, so exploration stopped before any site
   analysis finished.
2. `opencli doctor` and `opencli daemon status` could show a healthy
   connection even when the extension never reported a version, which
   hid the compatibility problem and sent users toward the wrong fix.
3. After reloading a new extension, `explore` could still report
   `Endpoints: 0 total, 0 API` because `handleNavigate()` detached the
   debugger before top-level navigation and cleared the active network
   capture state right before the page load we needed to observe.

Fix this in two layers:

- Teach `Page` to treat unsupported `network-capture-*` actions as an
  old-extension compatibility case. It now warns once, memoizes the
  unsupported state, and returns empty capture data instead of throwing.
- Teach `doctor` and `daemon status` to treat "connected but version
  unknown" as a warning instead of a healthy state, so version-handshake
  failures are visible immediately.
- Preserve the debugger attachment while network capture is armed, so
  the initial navigation keeps the capture state alive and the extension
  can record requests from the first page load.

Before:

- `opencli explore ...` -> `Error: Unknown action: network-capture-start`
- `opencli doctor` -> `[OK] Extension: connected` / `Everything looks good!`
- `opencli daemon status` -> `Extension: connected` even when the
  extension version was missing
- `opencli explore ...` after reloading the extension -> `Endpoints: 0 total, 0 API`

After:

- `opencli explore ...` on an old extension -> warns once and continues
- `opencli doctor` -> `[WARN] Extension: connected (version unknown)`
- `opencli daemon status` -> `Extension: connected (version unknown)`
- `opencli explore ...` on the reloaded extension keeps network capture
  armed across navigation instead of clearing it before the page load

* fix: reset network capture flags on closeWindow()

Prevents stale _networkCaptureUnsupported flag from persisting across
sessions when the user reinstalls or reloads the extension mid-session.

* fix: startNetworkCapture returns boolean to prevent false-positive on old extensions

When the extension doesn't support network-capture-*, startNetworkCapture()
now returns false instead of silently resolving. This ensures browser open/
network correctly falls back to the JS interceptor on old extensions.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-15 12:07:31 +08:00
jakevin 0040081f2b fix: auto-restart stale daemon and improve connection error messages (#1028)
* fix: auto-restart stale daemon and improve connection error messages

When daemon is running but extension never connected (stale daemon started
before extension was installed), the CLI now auto-restarts the daemon to
give the extension a fresh WebSocket endpoint, instead of just waiting
and then telling the user to install the extension.

Also improves error messages across cli.ts, bridge.ts, and doctor.ts to
suggest "opencli daemon stop && opencli doctor" as the quick fix, since
that's what actually resolves the issue.

* fix: version-aware stale daemon detection and improved error messages

- Daemon /status now includes `daemonVersion` field
- bridge.ts: when daemon is running but extension not connected, checks
  daemonVersion vs CLI version. Only auto-restarts if version mismatch
  (stale daemon from older CLI). Same-version daemon shows improved error
  message with "opencli daemon stop && opencli doctor" hint.
- doctor.ts: explicitly identifies stale daemon (version mismatch) in
  diagnostics report, shows daemon version in status line
- cli.ts: error message changed to suggest "opencli daemon stop && opencli doctor"

* fix: treat missing daemonVersion as stale, verify shutdown before respawn

- Missing daemonVersion (pre-version daemon) is now treated as stale,
  covering the most common user scenario (old daemon without version field)
- After requestDaemonShutdown(), poll until daemon actually stops (port
  released) before spawning new one, with 3s timeout
- If shutdown request fails, log warning instead of silently proceeding
- doctor.ts also treats missing daemonVersion as stale with clear message

* fix: fail explicitly when stale daemon replacement fails

- If shutdown request fails or port isn't released within 3s, throw
  'Stale daemon could not be replaced' instead of blindly spawning on
  an occupied port
- Add tests for all three stale-daemon branches: same-version (no
  restart), missing daemonVersion (stale), mismatched version (stale)

* fix: use type-based error dispatch in browserAction instead of string matching

browserAction() now checks `instanceof BrowserConnectError` first and
renders both message and hint, instead of string-matching on message
content. This ensures stale daemon errors ("Stale daemon could not be
replaced") surface the actionable hint to the user.
2026-04-15 11:33:26 +08:00
AstroHan 16d597cfce fix(doubao): harden ask response parsing (#933)
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-14 20:43:12 +08:00
flizzywine ba3a674d7b feat(grok): add image command for grok.com image generation (#906)
* feat(grok): add image command for grok.com image generation

Add `opencli grok image <prompt>` which submits a prompt via the existing
grok.com browser session and returns the generated image URLs from the
latest assistant bubble.

Because assets.grok.com URLs are gated by Cloudflare and cannot be
downloaded with a plain HTTP client, the --out flag triggers an in-page
fetch(credentials: 'include') so the browser session's cookies and
referer are attached, then writes the decoded blob to disk.

Flags:
- --new       start a fresh chat before sending
- --timeout   max seconds to wait for the image (default 240)
- --count     minimum number of images to wait for before returning
- --out       directory to save downloaded images

Ships with unit tests for the helpers (isOnGrok, normalizeBooleanFlag,
dedupeBySrc, imagesSignature, extFromContentType, buildFilename).

* fix(grok): harden image composer and bubble detection

* fix(grok): harden image flow and docs

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-14 20:42:43 +08:00
warkcod 0e38fd8c37 Feat/douban book subject (#993)
* chore: ignore local worktrees

* feat(douban): support book subject details
2026-04-14 20:41:14 +08:00
AstroHan cd48917a39 fix(xiaohongshu): require signed note URLs (#996)
* fix(xiaohongshu): require signed note urls

* chore: drop generated manifest from pr
2026-04-14 20:40:58 +08:00
CissiBot 45d6f5b09f feat(uiverse): add Uiverse code and preview adapters (#1000)
* feat(uiverse): add code and preview adapters

* fix(manifest): register uiverse commands

* docs(uiverse): add usage examples

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-14 20:38:38 +08:00
XavierCai 3ebc46f978 feat(bilibili): favorite command supports specifying fid (#1013)
* feat(bilibili): favorite command supports specifying fid

* fix(bilibili): sync favorite help and docs contract

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-14 20:37:54 +08:00
Benjamin Liu 6e29845dc3 fix(plugin): install monorepo sub-plugin dependencies when not hoisted (#1007)
Closes #722
2026-04-14 17:25:50 +08:00
mademing68092354-glitch 88bce1becf fix(chatgpt): support Chinese UI for model selector (#1006)
When ChatGPT macOS app is set to Chinese language, the "Options"
button label becomes "选项". This change checks for both English
and Chinese labels to find the button.

Co-authored-by: mad <mademing@maddeMac-mini.local>
Co-authored-by: Claude <noreply@anthropic.com>
2026-04-14 17:21:54 +08:00
jakevin ca68f3999b feat: Ref-Backed Locator for browser actions (#1016)
* feat: implement Ref-Backed Locator for browser actions

Introduces a unified target resolution system with fingerprint
verification and structured error diagnostics.

Snapshot phase:
- Each interactive element now gets a fingerprint (tag, role, text,
  ariaLabel, id, testId) stored in window.__opencli_ref_identity
- Zero overhead: metadata is already available during DOM walk

Resolution phase (new target-resolver.ts):
- Numeric input → ref path with fingerprint verification
- CSS-like input → querySelectorAll with uniqueness check
- No more silent first-match: ambiguous selectors are rejected

Error model (new target-errors.ts):
- stale_ref: element identity changed since snapshot
- ambiguous: CSS selector matched multiple elements (with candidates)
- not_found: element not in DOM or invalid input
- All errors include actionable hints for AI agents

base-page.ts:
- click() and typeText() now use two-phase resolve-then-act
- Existing CDP fallback for click preserved

* feat: migrate scrollTo to unified resolver pipeline

scrollTo now uses the same two-phase resolve-then-act pattern as
click and typeText, getting fingerprint verification and structured
error diagnostics (stale_ref/ambiguous/not_found) for free.

* fix: address review — stronger fingerprint verification & surface TargetError in CLI

1. Fingerprint verification now uses the full identity vector (tag, id,
   testId, ariaLabel, role, text) instead of just tag/role/text. Strong
   identifiers (id, testId) are decisive; remaining signals use majority
   voting. Fixes false negatives where same-tag elements swapped.

2. browserAction() now renders TargetError with code, hint, and
   candidates list instead of just the message string.

* fix: migrate get/select/type-autocomplete to unified resolver

- browser get text/value/attributes now resolve via resolveTargetJs
  instead of raw querySelector, getting fingerprint verification and
  structured errors for free
- browser select uses selectResolvedJs on __resolved element
- type command's autocomplete detection uses isAutocompleteResolvedJs
  on the already-resolved element
- Fix empty-string text prefix match: fp.text="Login" + text="" no
  longer falsely passes fingerprint check
2026-04-14 16:57:05 +08:00
jakevin 847c8317b6 fix(twitter): register lists command in manifest (#1011) 2026-04-14 10:37:34 +08:00
forvendettaw 741bcf9b6e Add bookmark_count field to bookmarks command (#1010)
* Add bookmark_count field to bookmarks command

Extract bookmark_count from legacy object in Twitter GraphQL
Bookmarks response. Add to returned tweet object and table columns.

* fix(manifest): sync twitter bookmarks columns

---------

Co-authored-by: Hermes Agent <hermes@lei.zong>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-14 10:27:20 +08:00
dependabot[bot] 44388d21fc chore(ci): bump softprops/action-gh-release from 2.6.1 to 3.0.0 (#1002)
Bumps [softprops/action-gh-release](https://github.com/softprops/action-gh-release) from 2.6.1 to 3.0.0.
- [Release notes](https://github.com/softprops/action-gh-release/releases)
- [Changelog](https://github.com/softprops/action-gh-release/blob/master/CHANGELOG.md)
- [Commits](https://github.com/softprops/action-gh-release/compare/v2.6.1...v3.0.0)

---
updated-dependencies:
- dependency-name: softprops/action-gh-release
  dependency-version: 3.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-14 10:18:45 +08:00
dependabot[bot] 745ce459d1 chore(deps): bump undici from 8.0.2 to 8.1.0 (#1003)
Bumps [undici](https://github.com/nodejs/undici) from 8.0.2 to 8.1.0.
- [Release notes](https://github.com/nodejs/undici/releases)
- [Commits](https://github.com/nodejs/undici/compare/v8.0.2...v8.1.0)

---
updated-dependencies:
- dependency-name: undici
  dependency-version: 8.1.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-14 10:18:32 +08:00
dependabot[bot] a5cd0dc307 chore(deps): bump @types/node from 25.5.2 to 25.6.0 (#1004)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.5.2 to 25.6.0.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.6.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-14 10:18:24 +08:00
dependabot[bot] beabed4bad chore(deps): bump vitest from 4.1.2 to 4.1.4 (#1005)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.2 to 4.1.4.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.4/packages/vitest)

---
updated-dependencies:
- dependency-name: vitest
  dependency-version: 4.1.4
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-14 10:18:15 +08:00
jakevin fa208ec761 docs: sync Highlights cleanup across all doc surfaces (#1009)
- docs/index.md: update feature cards to match README Highlights
- docs/zh/index.md: sync Chinese feature cards
- docs/guide/getting-started.md: align Highlights section
- README.zh-CN.md: rename "为什么是 OpenCLI" to "亮点", align with EN
2026-04-14 09:24:33 +08:00
jakevin 56a727cc04 docs: remove empty Why OpenCLI section and clean up Highlights (#1008)
- Remove the empty "Why OpenCLI" heading
- Rename "CLI All Electron" to "Desktop App Control" for clarity
- Remove "Anti-detection built-in" (exposes implementation details)
- Remove "Broad coverage" (duplicates intro and Built-in Commands table)
- Merge "Self-healing setup" and "Dynamic Loader" out (minor features)
- Rename "External CLI Hub" to "CLI Hub" for brevity
2026-04-14 09:19:08 +08:00
jakevin feedaf93b4 fix: remove duplicate extension zip from releases (#1001)
* fix: remove duplicate extension zip from releases

The release and build-extension workflows were creating both
opencli-extension.zip and opencli-extension-v{version}.zip (identical
content), causing both to be uploaded. Keep only the versioned filename.

* docs: update extension zip filename to versioned format

Update all references from opencli-extension.zip to
opencli-extension-v{version}.zip to match the workflow change.
2026-04-13 23:47:58 +08:00
jakevin 9ebb921c89 chore: prune legacy config switches (#998) 2026-04-13 23:28:30 +08:00
jakevin 9ac2e1d8ef chore: bump version to 1.7.3 (#997)
Release / release (push) Has been cancelled
2026-04-13 23:12:50 +08:00
SherlockSalvatore 2aee4caa10 feat(mubu): add Mubu adapter with 5 commands (#964)
* feat(mubu): add mubu (mubu.com) adapter with 5 commands

Commands: doc, docs, notes, recent, search.

- Uses COOKIE strategy; API calls via in-page XHR with Jwt-Token
  from localStorage (matches the web app's own mechanism).
- Renders node trees to Markdown (default) or plain text;
  supports tables, tasks, images, emoji, mentions, strikethrough,
  underline, and nested structures.
- notes supports flexible time ranges: single day, month, year,
  or custom --from/--to spans, plus a --list overview mode.
- search returns full-text matches with hit count and snippets
  for both folders and documents.

* fix(manifest): register mubu commands in runtime manifest

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-13 16:50:32 +08:00
jakevin 323fe8857c refactor: unify OPENCLI_VERBOSE and DEBUG=opencli (#991)
* refactor: unify OPENCLI_VERBOSE and DEBUG=opencli into one mechanism

Three debug output levels (verbose/debug/diagnostic) was redundant.
Merge DEBUG=opencli into OPENCLI_VERBOSE so `-v` flag controls all
verbose/debug output through a single mechanism.

- log.verbose() now checks both OPENCLI_VERBOSE and DEBUG=opencli
- log.debug() becomes an alias for log.verbose() (backward compat)
- boss/utils.js verbose helper simplified to check OPENCLI_VERBOSE only
- DEBUG=opencli still works as fallback (no breaking change)

* fix(boss): preserve debug fallback for verbose logs
2026-04-13 16:48:06 +08:00
jakevin ff6563d12a Fix automation window not closing on command failure (#992)
The error path in executeCommand did not call page.closeWindow(),
leaving the automation window open until the extension's idle timer
fires. On Windows, MV3 service worker suspension makes this timer
unreliable, causing windows to linger indefinitely.

Now closeWindow is called after diagnostic collection but before
rethrowing, ensuring the window is closed on both success and failure.
2026-04-13 16:47:47 +08:00
jakevin c42b040af4 Rename chatgpt adapters: desktop → chatgpt-app, web → chatgpt (#989)
* Rename chatgpt adapters: desktop → chatgpt-app, web → chatgpt

Aligns with existing `-app` suffix convention (discord-app, doubao-app):
- clis/chatgpt/ (desktop, AppleScript) → clis/chatgpt-app/
- clis/chatgptweb/ (browser, chatgpt.com) → clis/chatgpt/
- electron-apps.ts: chatgpt → chatgpt-app
- Updated all docs and README references

Closes #283

* Fix review findings: update cli-manifest.json and skill docs

- cli-manifest.json: update site/modulePath/sourceFile from chatgpt to chatgpt-app
- skills/opencli-usage/desktop.md: update commands from chatgpt to chatgpt-app
2026-04-13 14:33:32 +08:00
jakevin 79a15e8353 Remove unused OPENCLI_SKIP_FETCH env var (#987)
The adapter sync already has version caching (skips if same version)
and makes no network requests, so this opt-out flag adds no value.
2026-04-13 14:09:10 +08:00
jakevin 6d769ff354 docs: document undocumented environment variables (#983)
Add missing env vars to both README and README.zh-CN:
- OPENCLI_SKIP_FETCH: skip adapter sync on global install
- OUTPUT: override output format (json/yaml/table)
- DEBUG=opencli: internal debug logging
- DEBUG_SNAPSHOT: DOM snapshot debug output
2026-04-13 14:01:46 +08:00
jakevin 988ed19223 fix: clean up stale .yaml adapter files from older versions (#953) (#986)
* fix: clean up stale .yaml adapter files from older versions (#953)

Users upgrading from v1.6.x retain .yaml adapter files in
~/.opencli/clis/ that trigger "Ignoring YAML adapter" warnings on
every run. The hash-based sync only tracks .js files, so these
legacy .yaml files are never cleaned up.

Add a cleanup step (3b) that removes .yaml/.yml files from user
adapter directories when the corresponding site exists in the
official package (i.e., the site has been migrated to .js).

* fix(fetch-adapters): narrow stale yaml cleanup
2026-04-13 13:17:27 +08:00
jakevin 51bc48ec61 feat: decouple extension version from CLI version (#985)
* feat: decouple extension version from CLI version

Extension and CLI had tightly coupled version numbers (both 1.7.2),
requiring manual sync across 3 files on every release. This decouples
them so each can release independently.

Changes:
- Extension version reset to 1.0.0 with independent versioning
- Extension sends compatRange (e.g. ">=1.7.0") in hello message
  so doctor can check CLI/extension compatibility
- Daemon stores and exposes extensionCompatRange via /status
- Doctor uses compatRange for compatibility checks (falls back to
  major-version check for older extensions without compatRange)
- Doctor shows extension update availability from cached GitHub
  Releases data
- release.yml always builds and attaches extension zip to every
  CLI release, so users always find both in the same release page
- build-extension.yml triggers on ext-v* tags (not v*) to avoid
  duplicate builds

* fix: version extension release assets
2026-04-13 12:43:34 +08:00
jakevin 72bc86cf41 fix: code audit round 2 — safety, hot-reload, error diagnostics (#982)
* fix: code audit round 2 — pruneEmptyDirs, evaluateWithArgs, hot-reload, error cause chain

1. pruneEmptyDirs: use path.relative() instead of startsWith() to prevent
   false boundary matches on overlapping directory names
2. evaluateWithArgs: add safe evaluate method that auto-serializes args via
   JSON.stringify, preventing injection by design
3. Hot-reload: detect mtime changes on user adapter files in daemon mode,
   invalidate module cache so edits take effect without restart
4. toEnvelope: preserve error cause chain in verbose mode for better
   production debugging

* fix: address review feedback on code audit round 2

- pruneEmptyDirs: resolve() paths before relative() check
- evaluateWithArgs: validate keys are valid JS identifiers
- hot-reload: only bust ESM cache on reload, not first load
- toEnvelope: move cause serialization into toEnvelope itself
  so all consumers (AI agents, MCP tools) get cause chain
2026-04-13 09:36:53 +08:00
jakevin ffb61c51ea fix: address code audit findings (C1-C4, I1, I4, I6) (#981)
* fix: address code audit findings (C1-C4, I1, I4, I6)

Security:
- C1: Fix page.evaluate injection in browser type/select commands and
  6 adapter files by using JSON.stringify for user input interpolation
- C2: Close WebSocket on CDP connect timeout to prevent resource leak
- C3: Reject CDP connect promise on Page.enable failure instead of
  silently swallowing the error

Reliability:
- C4: Guard against corrupted adapter-manifest.json hashes to prevent
  false-positive override deletion
- I1: Throw on pre-navigation failure instead of warn-and-continue
- I4: Use Map<string, Promise<void>> for lazy module loading to prevent
  concurrent double-imports of the same adapter

Performance:
- I6: Replace O(n) registry alias cleanup with O(k) direct deletion

* fix: address self-review findings on PR #981

- C1: add quotes around CSS selector attribute values in browser
  type/select to match other commands (get text/value/attributes)
- C2: clear this._ws in timeout handler to prevent race with open event
- C4: refine corruption guard — treat null/undefined hashes as empty,
  only skip sync for truly invalid types (string, number, array)
2026-04-13 09:24:01 +08:00
AstroHan 5dcbf92a59 fix(douban): classify tv search results correctly (#979) 2026-04-13 08:41:19 +08:00
AstroHan 83dce2430e fix(xiaohongshu): harden anti-detection flows (#980) 2026-04-13 08:40:48 +08:00
Tony Simons 4d1fa8a6e2 feat(clis/chatgptweb): add ChatGPT web image generation command (#973)
* feat(clis/chatgptweb): add ChatGPT web image generation command

Add `opencli chatgptweb image` command that generates images using
ChatGPT web (GPT-4o image generation) and saves them locally.

Features:
- Navigates to chatgpt.com/new with full page reload to ensure clean state
- Uses Playwright's page.type() for reliable text input in TipTap editor
- Closes sidebar if open (covers the chat composer on some layouts)
- Polls for response completion (handles thinking/throttling states)
- Extracts generated images from DOM (backend-api/estuary/content URLs)
- Downloads and saves as PNG/JPEG files to user-specified directory
- Supports --op for output directory and --sd to skip download

Files:
- clis/chatgptweb/image.js: CLI command definition
- clis/chatgptweb/utils.js: DOM helpers, send/wait/export functions

Works cross-platform (Linux/macOS/Windows) via OpenCLI browser automation.

* fix(chatgptweb): stabilize image generation flow

* docs(chatgptweb): add browser adapter guide

---------

Co-authored-by: Tony Simons <tony@tonysimons.dev>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-12 21:24:03 +08:00
Harvey Yue aa47de726d feat(bilibili): add feed-detail and enhance feed command (#974)
* feat(bilibili): add feed-detail and enhance feed command

* docs: add binance adapter documentation

* docs: add feed-detail command to bilibili docs

* docs: sync bilibili adapter contract for feed-detail

* docs: add ke adapter page for doc coverage

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-12 21:19:12 +08:00
runzhliu 232b6828be feat(ke): add Beike (贝壳找房) adapter with ershoufang, xiaoqu, zufang, chengjiao commands (#975)
Support browsing second-hand houses, neighborhoods, rentals, and
transaction records on ke.com with city/district/price filtering.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-12 16:21:59 +08:00
Ivan Xia 20e024a001 feat(maimai): add talent search with multi-dimensional filters (#977)
* feat(maimai): add talent search with multi-dimensional filters

Add maimai.cn talent search adapter with support for:
- Keyword search (query)
- Company filtering (multiple companies supported)
- School filtering (with 985/211 options)
- Location filtering (province/city)
- Work experience and education level filters
- Industry and position filters
- Direct chat availability
- Sort by relevance, activity, work years, or education

Features:
- Reuses Chrome login session for authentication
- Extracts candidate info: name, job title, company, work history
- Shows work years, education, age, active status
- Displays skill tags and mutual friends count

* fix docs and strategy for maimai adapter

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-12 16:16:55 +08:00
Alex Yang 01957bf654 feat(discord-app): add delete command to remove a message by ID (#976)
* feat(discord-app): add delete command to remove a message by ID

Adds a new `delete` command for the discord-app CLI that deletes a
message in the active channel by its snowflake ID. Uses the UI strategy
to hover the message, open the "More" menu, click "Delete Message", and
confirm the deletion dialog.

* docs: add binance adapter doc and update discord doc with delete command
2026-04-12 16:07:14 +08:00
jakevin 315cc59f8a chore: bump version to 1.7.2 (#972)
Build Chrome Extension / build (push) Has been cancelled
Release / release (push) Has been cancelled
2026-04-11 22:14:43 +08:00
jakevin ebe6be945e fix(zsxq): update topic test for group_id parameter added in #963 (#971)
The test mock was missing the evaluate call for getActiveGroupId,
which was added when #963 introduced the group_id parameter.
2026-04-11 22:12:26 +08:00
iiilin b0a019121e feat(weibo): support for-you and following feed types (#959)
* feat(weibo): support for-you and following feed types

* docs: clarify weibo feed types

---------

Co-authored-by: iiilin <19162130+iiilin@users.noreply.github.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-11 22:04:01 +08:00
Paul Zhu c63bad34b0 feat(twitter): add lists command to retrieve user lists (#958)
* feat(twitter): add lists command to retrieve user lists

Add twitter/lists command that fetches Twitter/X lists for a user.
Supports:
- Lists with member and follower counts
- Private/public mode detection
- Default to current user if no user specified
- Works for any Twitter user

* docs: add lists command to twitter commands in README

Add twitter lists command to Built-in Commands table in both
English and Chinese README files

* fix(twitter): parse lists from card DOM instead of locale-specific page text

---------

Co-authored-by: isanwenyu <isanwenyu@users.noreply.github.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-11 21:58:46 +08:00
Hoshea 1ffe12f85e fix(zsxq): accept topic_id as string in getTopicFromResponse (#963)
* fix(zsxq): accept topic_id as string in getTopicFromResponse

The ZSXQ API returns topic_id as a string, but getTopicFromResponse()
only checked for typeof === 'number', causing it to fall through and
return null. This made 'opencli zsxq topic <id>' fail with NOT_FOUND
for all valid topic IDs.

* fix(zsxq): use group-scoped topic endpoint instead of bare /v2/topics/{id}

The ZSXQ API requires topics to be fetched within their group context.
Change /v2/topics/{id} -> /v2/groups/{groupId}/topics/{id} for both
the detail and comments endpoints. Also adds optional --group_id arg.
2026-04-11 21:49:02 +08:00
jakevin 30f216b2c1 fix: include adapter tests in default npm test (#969)
* fix: include adapter tests in default npm test

`npm test` only ran unit + extension projects, so adapter tests
(clis/**/*.test.js) were never exercised by the default test command.
Add --project adapter so they run alongside unit and extension tests.

* test: include adapter project in default npm test
2026-04-11 21:33:50 +08:00
jakevin 3c088da53e refactor: smart sync adapters — hash-based diff instead of full copy (#966)
* refactor: smart sync adapters instead of full copy (#sparse-override)

Replace unconditional full-copy of all adapters to ~/.opencli/clis/ with
hash-based smart sync that only copies files whose content has changed.

Changes:
- fetch-adapters.js: use SHA-256 content hashes to skip unchanged files;
  store per-file hashes in adapter-manifest.json
- discovery.ts: simplify ensureUserAdapters() to only create the directory
  (no longer triggers full copy on first run)
- main.ts: fix fast completion to check manifest file existence instead of
  directory existence (sparse override may have empty user dir)
- cli.ts: add `opencli adapter eject/reset/status` commands for managing
  local adapter overrides
- engine.test.ts: add tests for empty user dir and ensureUserAdapters

* fix: address review blockers — site-level sync + reset --all

1. Fix `adapter reset --all`: change <site> from required to optional
   argument so --all can be used without specifying a site name.

2. Change smart sync from file-level to site-level granularity:
   if any file in a site has changed upstream, overwrite the entire
   site directory. This matches the agreed product semantics — local
   modifications to any file in a site are replaced when upstream
   updates that site.

* fix: delete old site dir before writing updated adapter files

When a site has upstream changes, delete the entire site directory
first, then write the new version. This prevents stale files from
older versions lingering in the user directory.

* fix: reset --all preserves custom sites, only removes official overrides

Blocker 3 fix: reset --all now checks BUILTIN_CLIS to identify official
sites and only deletes those, preserving user-created custom sites.

* refactor: sparse sync deletes local overrides instead of copying new versions

Changed fetch-adapters.js semantics per team agreement:
- When an official site has upstream changes, DELETE the local override
  instead of copying the new version into ~/.opencli/clis/
- Runtime automatically falls back to package baseline
- ~/.opencli/clis/ becomes a true sparse override layer

* fix: reset <site> rejects custom sites, only allows official overrides

Single-site reset now checks BUILTIN_CLIS before deleting, matching
the same protection that reset --all already has.

* fix: reset <site> allows custom sites per product decision

Per @WAWQAQ: explicit single-site reset should work on custom sites too.
Differentiate messaging: official sites say "using official baseline",
custom sites say "removed custom site".

reset --all still only removes official overrides (bulk safety).

* fix: reset --all deletes all local sites including custom per product decision

Per @WAWQAQ: --all should clear the entire local working cache,
including custom sites. Single-site reset already handles both types.
2026-04-11 21:28:40 +08:00
jakevin 00e200b0b7 migrate: move binance adapters from src/clis/ to clis/ (#967)
Binance was the only adapter left in src/clis/ after the TS→JS
migration (PR #928). Move all 11 adapters and the test file to
clis/binance/, strip TypeScript syntax from the test, and switch
the test import to the @jackwener/opencli/pipeline package export.
2026-04-11 21:21:52 +08:00
jakevin 1fd578c404 chore: bump version to 1.7.1 (#965)
Build Chrome Extension / build (push) Has been cancelled
Release / release (push) Has been cancelled
2026-04-11 20:28:23 +08:00
jakevin b0fd95279c docs: add CHANGELOG.md entry for v1.7.0 (#955)
Comprehensive release notes covering all changes since v1.6.1:
- Breaking changes: Node >= 21, YAML deprecated, .ts no longer loaded,
  error output as YAML envelope, tabId → targetId, operate → browser
- 10+ new adapters, 15+ adapter enhancements
- Major refactors: JS-first adapters, registry validation, strategy normalization
- Performance: P0 optimizations, fast-path completion, browser pipeline
- Upgrade guide with step-by-step migration instructions
2026-04-11 13:51:51 +08:00
jakevin 892ddcb19f docs: fix stale .ts adapter references in skills and guides (#954)
* docs: fix stale .ts adapter references in skills and guides

All adapters are now .js files. Update references in:
- opencli-oneshot SKILL.md (7 instances)
- opencli-autofix SKILL.md (1 instance)
- opencli-explorer references (15 instances)
- electron-app-cli guide (5 instances)

* docs: fix stale YAML/TS references in READMEs and zh docs

- Plugin type column: YAML/TS → JS (all 4 plugins have JS conversion PRs)
- synthesize command: "YAML adapters" / "TS adapters" → "JS adapters"
- zh index: remove "YAML 声明式" reference
- zh README: add missing vk plugin entry

* docs: fix remaining .ts references found in review

- electron-app-cli.md: "TypeScript desktop adapter" → "desktop adapter", file layout .ts → .js
- adapter-templates.md: section title "提取 utils.ts" → "提取 utils.js"
- opencli-explorer SKILL.md: "写 following.ts" → "写 following.js"
2026-04-11 13:13:27 +08:00
jakevin ee8d7cce77 docs: fix stale adapter counts and .ts reference (#950)
- README.md: "70+ pre-built adapters" → "87+"
- docs/comparison.md: "73+ sites" → "87+", ".ts adapter" → ".js adapter"
2026-04-11 13:13:21 +08:00
jakevin 420dc0f3c8 fix: DEBUG_SNAPSHOT should work without DEBUG=opencli (#952)
log.debug() requires DEBUG=opencli to output, which means
DEBUG_SNAPSHOT=1 alone no longer shows snapshot fallback diagnostics.
Use process.stderr.write directly since the DEBUG_SNAPSHOT guard
already controls when this diagnostic fires.
2026-04-11 13:13:14 +08:00
jakevin 0f90b42f71 fix: warn users when .ts adapters are found but not loaded (#951)
Users who created custom .ts adapters in ~/.opencli/clis/ will see
their commands silently disappear after upgrading to the JS-only
version. Add an explicit warning so they know to convert to .js.
2026-04-11 13:13:07 +08:00
jakevin 2469d12efd fix: resolve alias target correctly in validate command (#949)
The alias resolution logic checked `!registry.has(target)` before
calling `registry.get(target)`, which always returned undefined.
Moreover, aliases registered as `site/alias` keys meant `registry.has`
returned true, skipping the block entirely. The canonical name was
never resolved, so `validate site/alias` silently checked 0 commands.

Simplify to always resolve via `registry.get(target)` which handles
both canonical keys and alias keys correctly.
2026-04-11 13:13:02 +08:00
Harvey Yue 25014f1067 fix(bilibili): add missing domain for following cli (#947) 2026-04-11 12:36:41 +08:00
jakevin 63b7b291ab fix: clean up stale .ts adapter files during upgrade (#948)
Older versions (pre-1.7.1) shipped adapters as .ts files. When users
upgrade to a .js-only version, the old .ts files are left orphaned in
~/.opencli/clis/. Add a cleanup step that removes .ts files when a
corresponding .js official adapter exists.
2026-04-11 12:36:02 +08:00
jakevin 4a0b8054b2 fix: batch quality improvements — dedupe completion, unify logging, fix docs (#945)
* fix: batch quality improvements — dedupe completion, unify logging, fix docs

1. Extract shared completion code (BUILTIN_COMMANDS + shell scripts) into
   completion-shared.ts, eliminating duplication between completion.ts and
   completion-fast.ts.

2. Replace console.error/warn/log with log.* from logger.ts in:
   - daemon.ts (7 occurrences)
   - runtime.ts (1 occurrence)
   - cli.ts browserAction error handler (3 occurrences)
   - base-page.ts snapshot fallback (1 occurrence)
   - download/index.ts cookie warning (1 occurrence)
   - commands/daemon.ts (2 occurrences)

3. Fix Node version in build-extension.yml: 20 → 22 (matches package.json >=21)

4. Fix error handling consistency: tap.ts now throws CliError instead of bare Error

5. Remove 31 duplicate rows in docs/adapters/index.md (grok, gemini, yuanbao,
   notebooklm, doubao, weread + 25 more entries duplicated without .md suffix)

6. Update skill version: opencli-usage SKILL.md 1.6.9 → 1.7.0, adapter count 79 → 87

* fix: update daemon.test.ts to match logger migration

Tests now spy on process.stderr.write (used by log.*) instead of
console.log/console.error (no longer used by daemonStop).

* fix: address review feedback on PR #945

1. base-page.ts: restore DEBUG_SNAPSHOT env guard — log.debug uses a
   different env var (DEBUG=opencli), so keep the original gate to
   avoid breaking existing users.

2. daemon.ts: remove dead `prefix` variable left over from console.error
   migration.
2026-04-11 01:45:00 +08:00
jakevin 575986c656 perf: P0 performance optimizations (#944)
* perf: P0 performance optimizations — VM context reuse, startup parallelization, stealth caching

1. Reuse VM sandbox context in pipeline template engine instead of creating
   a new vm.createContext() on every expression evaluation. This eliminates
   ~0.3ms per call in map/filter loops over large arrays.

2. Cache sanitizeContext() results via WeakMap keyed by object reference.
   In pipeline loops, `args` and `data` are the same object across all
   iterations — the expensive JSON round-trip now runs only once per step.

3. Parallelize independent startup I/O: built-in CLI discovery now runs
   concurrently with ensureUserCliCompatShims and ensureUserAdapters,
   saving ~30-50ms on cold start.

4. Cache the stealth JS string (350 lines, pure static) after first
   generation — every subsequent goto() reuses the cached string.

* fix: address review feedback on P0 perf optimizations

1. sanitizeContext: cache JSON string instead of parsed object to prevent
   sandbox mutation from polluting subsequent calls
2. VM sandbox: clean non-whitelisted properties before each execution to
   prevent cross-expression state leakage
3. Startup parallelization: document registry overwrite semantics and
   confirm no shared-state race between parallel tasks
2026-04-11 01:31:34 +08:00
jakevin 110e047b3c refactor(validate): switch from YAML to registry-based validation (#943)
* refactor(validate): switch from YAML scanning to registry-based validation

The validate/verify commands only scanned YAML files, which are no
longer supported. Rewrite to validate commands from the in-memory
registry populated by discoverClis(), aligning with the JS-first
adapter architecture.

New checks: missing description, browser commands without domain,
pipeline step name typos, commands without func/pipeline, duplicate
arg names, and positional arg ordering.

* fix(validate): treat lazy-loaded commands as valid

Manifest-registered commands have _lazy=true and no func/pipeline
until execution time. Recognize this as a valid execution form.

* fix(validate): warn on empty registry, support alias targets

- Emit warning when registry is empty instead of silent PASS
- Resolve alias targets to canonical key before filtering
2026-04-11 01:17:14 +08:00
jakevin a9d21f3de0 fix: project hygiene — docs, lint, daemon restart (#942)
* fix: project hygiene — docs, lint, daemon restart, code fence

- Update Node version requirement from >= 20 to >= 21 in 7 doc files
  (README, README.zh-CN, installation guides, troubleshooting)
- Update adapter count from 79+ to 87+ in READMEs
- Remove duplicate `lint` script (identical to `typecheck`)
- Fix TESTING.md CI matrix: Node ['22'] instead of ['20', '22']
- Fix autofix SKILL.md code fence escaping (\``` → ~~~)
- Add daemon restart to postinstall so updated adapters are picked up
- Fix preuninstall to respect OPENCLI_DAEMON_PORT env var

* fix: align docs and skills with JS-first adapter contract

Adapters are now .js files (not .ts). Update all references across:
- README.md, README.zh-CN.md, CONTRIBUTING.md
- docs/guide/getting-started.md, docs/index.md
- skills/opencli-browser/SKILL.md, skills/opencli-explorer/SKILL.md

The runtime (discovery.ts) only loads .js from user clis/ directories,
and `opencli browser init` generates .js scaffolds. Documentation was
still teaching users to create .ts files.

* fix: update CI matrix to Node 22 only (drop Node 20)

package.json requires Node >= 21 (styleText dependency). The CI matrix
was still testing Node 20 which doesn't meet this requirement.

* fix: revert incorrect daemon restart from postinstall

The daemon (browser bridge) only handles CDP communication — it has no
knowledge of adapters. Adapter discovery, loading, and execution all
happen in the CLI process, which is fresh each invocation. The
_loadedModules cache in execution.ts is process-local and not a real
staleness concern. Remove the unnecessary restartDaemon() call.
2026-04-11 00:50:05 +08:00
jakevin 383d28fcf7 refactor: normalize strategy into runtime fields at registration time (#941)
Strategy is a 5-value enum (PUBLIC/COOKIE/HEADER/INTERCEPT/UI) that
the execution path was reading at two points — resolvePreNav() and
shouldUseBrowserSession() — to make decisions that are already fully
expressible by the existing `browser` and `navigateBefore` fields.

This commit introduces normalizeCommand() inside registerCommand(),
which expands strategy into concrete runtime fields at registration
time. After normalization, execution code never reads cmd.strategy.

normalizeCommand expansion rules:
  - strategy → browser: PUBLIC defaults to false, others to true.
    Explicit browser value always wins.
  - strategy + domain → navigateBefore:
    · COOKIE/HEADER + domain → 'https://{domain}' (pre-navigate)
    · Non-PUBLIC without domain → true (needs auth context, no URL)
    · PUBLIC → undefined (no auth needed)
    Explicit navigateBefore (false or string) always wins.

This matters because commands enter the registry from 4 sources
(cli(), manifest, generate-verified, tests), and previously only
cli() did strategy derivation. The other 3 constructed CliCommand
directly, leaving strategy as a runtime dependency. Now all sources
converge through registerCommand → normalizeCommand.

Changes:
  - registry.ts: add normalizeCommand(); simplify cli() to delegate
    all derivation to normalizeCommand via registerCommand()
  - execution.ts: resolvePreNav() no longer reads strategy; just
    reads the already-expanded navigateBefore field. Strategy import
    removed.
  - capabilityRouting.ts: shouldUseBrowserSession() checks
    cmd.navigateBefore (truthy = needs browser session) instead of
    cmd.strategy !== PUBLIC. Strategy import removed.
  - discovery.ts: manifest path no longer hardcodes browser default;
    delegates to normalizeCommand.
  - capabilityRouting.test.ts: test now reflects normalized command
    shape (navigateBefore: true for COOKIE without domain).

strategy is preserved as metadata on CliCommand — opencli list,
cascade probe, adapter generation, and documentation continue to
read it. Only the execution path stops consuming it.
2026-04-11 00:37:31 +08:00
jakevin f92f7571b4 chore: remove unused test-site.mjs script (#940)
Not referenced in package.json, CI, or documentation.
2026-04-11 00:19:02 +08:00
jakevin 93bd1eb88a docs: document autofix issue filing flow (#939) 2026-04-10 23:52:45 +08:00
jakevin 39a1d673ac feat(skill): add upstream issue filing step to opencli-autofix (#938)
Add Step 6 to the autofix skill: after a verified local fix, prepare a
GitHub issue draft and file it (with user confirmation) via `gh issue
create`. Pure skill/documentation approach — no new runtime code.

Closes the need addressed by #936 with zero code, zero tests to maintain.
2026-04-10 23:46:54 +08:00
jakevin cc18ed67b7 fix: sync package-lock.json to unblock CI (#937)
* fix: sync package-lock.json with package.json dependencies

package-lock.json was missing @emnapi/core@1.9.2 and
@emnapi/runtime@1.9.2 (transitive deps of @emnapi/wasi-threads),
causing `npm ci` to fail on all CI jobs.

* fix: resolve remaining CI failures after TS-to-JS adapter migration

- vitest.config.ts: update adapter project include/exclude from .test.ts
  to .test.{ts,js} to match converted adapter test files
- check-doc-coverage.sh: skip adapter directories containing only utility
  files (prefixed with _), fixing false positive for clis/slock/
- linux-do/topic-content.test.js: fix hardcoded reference to topic.ts
  (now topic.js after PR #928 migration)
2026-04-10 23:29:09 +08:00
jakevin 91c208c855 fix: address deep review findings (security, correctness, consistency) (#935)
* fix: address deep review findings (security, correctness, consistency)

1. Security: add path traversal guard for plugin manifest entry.path
2. Security: sanitize evaluate() index param via JSON.stringify
3. Correctness: fix startNetworkCapture idempotency (don't wipe entries on re-call)
4. Correctness: log pre-navigation failures instead of silently swallowing
5. Consistency: replace console.log/error with log module in commanderAdapter, external
6. Consistency: add PluginError class, convert user-facing plugin errors
7. Dedup: remove local isRecord() in plugin.ts, use shared utils.ts version
8. Clarify: document intentional double validateArgs call

* chore: remove unused chalk imports from external.ts and commanderAdapter.ts

* refactor: replace chalk with Node.js built-in util.styleText

- Remove chalk dependency, use `styleText` from `node:util` (stable in Node 21+)
- Bump engines to Node >= 21
- Update all 10 source files that used chalk
- Remove stale chalk mock from daemon.test.ts
- One fewer runtime dependency

* fix: tighten deep-review follow-up
2026-04-10 22:59:41 +08:00
jakevin 2288cf7149 fix: clean up legacy shim files and stale tmp files on upgrade (#934)
* fix: clean up legacy shim files and stale tmp files on upgrade

Add cleanup steps to fetch-adapters.js that run on every version upgrade:

1. Remove legacy compat shim files from ~/.opencli/ (registry.js,
   errors.js, utils.js, etc.) that were created by an older approach
   using file:// re-exports. Current approach uses node_modules symlink.
   Only deletes files containing "export * from 'file://" to avoid
   removing user-created files.

2. Remove legacy compat shim directories (browser/, download/, errors/,
   etc.) using the same safety check.

3. Clean up stale .plugins.lock.json.tmp-* files left behind by
   crashed processes. These accumulate over time (108 found on one
   machine) and clutter ~/.opencli/.

* fix: check every file in legacy shim directories before deleting

Instead of checking only the first file and deleting the entire
directory, now checks each file individually and only deletes files
matching the shim pattern. Directory is removed only if empty after
individual file cleanup.
2026-04-10 18:58:44 +08:00
jakevin 2457002167 chore: remove migration residuals (mapDistToSource, clean-yaml) (#931)
- Remove mapDistToSource() from diagnostic.ts — mapped dist/clis/
  paths back to clis/ but dist/clis/ no longer exists after JS-first
  migration. The function always returned null.
- Simplify resolveAdapterSourcePath() to check candidates directly
  without the dead dist→source mapping detour.
- Delete scripts/clean-yaml.cjs — walked dist/clis/ to delete YAML
  files, but dist/clis/ no longer exists.
- Remove clean-yaml script entry from package.json.
2026-04-10 16:41:44 +08:00
jakevin 714df646a5 fix(security): escape codegen strings and redact diagnostic body (#930)
1. candidateToJs: escape single quotes in site, name, domain, and arg
   name/type fields to prevent syntax errors in generated JS adapters.
   Previously only description and help fields were escaped.

2. diagnostic: pass network request body through redactText() to
   prevent sensitive data (JWT, bearer tokens) from leaking into
   repair context. responseBody/responsePreview already used
   sanitizeCapturedValue which calls redactText, but the body field
   only had truncation.
2026-04-10 15:29:32 +08:00
jakevin d2974a9ff6 refactor(adapters): convert adapter layer from TypeScript to JavaScript (#928)
* refactor(adapters): convert adapter layer from TypeScript to JavaScript

Core framework stays TypeScript; adapter layer moves to JS-first.
Adapters are essentially "executable config + browser scripts" that
barely use TS features — this simplifies the build/distribution pipeline
by removing the dist/clis/ intermediate compilation step.

Changes:
- Convert all 753 adapter files in clis/ from .ts to .js
- Update tsconfig to exclude clis/ from compilation
- Simplify build-manifest to scan clis/*.js directly (no dist/clis/)
- Update discovery, main, fetch-adapters to load JS adapters from clis/
- Update generate-verified to output .js artifacts
- Update package.json files field: dist/clis/ → clis/
- Fix all test files for the .ts → .js transition

* fix(main): use findPackageRoot for BUILTIN_CLIS path

The previous relative path (../../clis from __dirname) only worked for
dist/src/main.js but broke dev mode (tsx src/main.ts) where __dirname
is <repo>/src — resolving to /clis instead of <repo>/clis.

Use findPackageRoot() which works for both dev and prod paths.
2026-04-10 14:52:18 +08:00
jakevin b45a64d91d fix(build-manifest): import compiled JS from dist/clis/ instead of raw TS (#926)
* fix(build-manifest): import compiled JS from dist/clis/ instead of raw TS

Node's type stripping does not rewrite '.js' → '.ts' in import
specifiers, so dynamically importing .ts source files fails whenever
they contain relative imports like './utils.js'.

Switch to scanning dist/clis/ for compiled .js files after tsc runs.
This eliminates all 268 "Cannot find module" warnings and increases
manifest entries from 254 to 532 (previously half were silently skipped).

* fix: write manifest to dist/cli-manifest.json where runtime expects it

The runtime resolves BUILTIN_CLIS to dist/clis/ (relative to
dist/src/main.js), so discoverClis() looks for manifest at
dist/cli-manifest.json. Previously it was written to the package root
where the runtime never found it — manifest was effectively unused,
always falling through to filesystem scanning.
2026-04-10 12:58:34 +08:00
jakevin dbac7fc921 refactor(errors): unify error output as YAML envelope to stderr (#923)
* refactor(errors): unify error output as YAML envelope to stderr

Replace the 100+ line chalk renderError() switch-case with a single
YAML envelope output path. All errors now output a structured
{ok, error: {code, message, help, exitCode}} envelope to stderr,
regardless of TTY status.

This simplifies the error system from 5 mechanisms to 3:
1. Error Envelope (YAML → stderr) — unified error output
2. Exit codes (sysexits.h) — process exit semantics
3. Diagnostic (OPENCLI_DIAGNOSTIC=1) — autofix repair context

Removed: chalk error rendering, ERROR_ICONS map, classifyGenericError
regex classifier, BrowserConnectError-specific bridge status display.
Added: toEnvelope() utility, ErrorEnvelope type.

* refactor(errors): migrate adapters to throw CliError, update docs

- Migrate xueqiu adapters from return [{error,help}] to throw CliError
- xueqiu/utils.ts: fetchXueqiuJson now throws AuthRequiredError/
  CommandExecutionError instead of returning {error, help} objects
- Remove resolveColumns error fallback from output.ts (no longer needed)
- Add verbose stack trace support to error envelope
- Add ADAPTER_LOAD to AutoFix hint trigger codes
- Update skill docs (adapter-templates, explorer, oneshot, advanced-patterns)
  to recommend throw CliError pattern instead of return [{error, help}]

* fix: remove remaining dead error-forwarding in 4 xueqiu adapters + review fixes

- Remove `if ('error' in d) return [d]` from feed, hot, search, kline
  (fetchXueqiuJson now throws, so these were dead code)
- Add `stack?: string` to ErrorEnvelope interface (removes type cast hack)
- Fix adapter-templates.md: use AuthRequiredError instead of plain Error

* fix: migrate barchart/quote and yahoo-finance/quote to throw CliError

Last two adapters that silently returned [] on error instead of
throwing CommandExecutionError.

* fix: self-review fixes — doc evaluate crash, error messages, kline consistency

- adapter-templates.md: getServerContext was throwing AuthRequiredError
  inside a function serialized into page.evaluate() (browser has no
  CliError). Reverted to return {error} sentinel + func() body throw.
- yahoo-finance/quote, barchart/quote: include symbol in fallback error msg
- xueqiu/kline: throw EmptyResultError instead of returning [] for
  consistency with other xueqiu adapters
2026-04-10 03:20:53 +08:00
jakevin 309dadcf46 refactor(adapters): migrate pipeline adapters to func() + { error, help } pattern; docs: skill improvements (#922)
* docs(skills): add Tier 2.5 localStorage Bearer, SPA discovery, and test standards

From real-world experience building slock.ai CLI adapters:

- oneshot: add network-empty diagnosis, SPA baseURL bundle search, Tier 2.5
  localStorage Bearer template (with multi-tenant X-Server-Id pattern),
  updated auth quick-reference, file path note, opencli browser verify test flow
- explorer: add Tier 2.5 to decision tree and strategy table, update test section
  with opencli browser verify + Done standard, fix Step 5 path to ~/.opencli/clis/,
  add 4 new pitfall rows (SPA HTML, 400 context header, empty network, wrong dir)

* docs(skills): fix path conflict + add anti-change patterns from real adapters

Fix reviewer blocking issue:
- Remove the contradictory "~/.opencli/clis/" note that mixed user-local and
  repo-contributor workflows; replace with explicit two-scenario callout in
  Step 4, Step 5, pitfall table, and oneshot test section
- Template comments in oneshot restored to clis/<site>/<name>.ts (repo path)

Add "抗变更模式" section to explorer, based on opencli's own production code:
- Pattern 1: dynamic queryId discovery (twitter/shared.ts resolveTwitterQueryId)
  — scan loaded JS bundle by operationName (stable) to find queryId (unstable)
- Pattern 2: semantic DOM priority fallback (web/read.ts)
  — article > [role=main] > main > class-hint > body, pick largest text block
- Pattern 3: ordered selector array + timestamp comments (xiaohongshu/publish.ts)
  — first-match wins, comment records UI version and observed attribute values
- Pattern 4: nullish-coalescing field multi-path (xiaohongshu/user-helpers.ts)
  — covers camelCase/snake_case variants without assuming fixed key name

* docs(explorer): split SKILL.md into reference sub-documents

- Shrink main SKILL.md from 994 to 270 lines — core workflow only
- Extract all TS templates (Tier 1~4, pagination) to references/adapter-templates.md
- Add error handling standard: { error, remedy } pattern (remedy > hint)
- Add Tier 2.5 localStorage Bearer template with multi-tenant X-Server-Id example
- Extract cascading requests, tap debug, verbose mode, anti-change patterns to references/advanced-patterns.md
- Extract record workflow to references/record-workflow.md

* docs(skills): fix verify command — split by dev scenario

browser verify only reads ~/.opencli/clis/, not repo's clis/.
Split all verify instructions:
- Repo 贡献: npm run build + opencli <site> <cmd>
- 私人 adapter: opencli browser verify <site>/<name>

Fixes blocker in explorer:L209, L224 and oneshot:L286, L298

* docs(adapter-templates): add utils.ts extraction pattern for same-site adapters

* docs(skills): add decision matrix, stop conditions, sync comments

explorer: add path decision matrix before core workflow
oneshot: add explicit stop/switch conditions (when to escalate to explorer)
both: add keep-in-sync comment on the two-scenario verify block

* feat(slock): extract utils.ts + apply { error, help } pattern; docs: remedy→help

slock/utils.ts: new — getSlockContext(), resolveChannelId()
  - Shared token + workspace resolution, no more 4-line duplication
  - UUID regex (/^[0-9a-f]{8}-...$/) replaces fragile !includes('-')
  - Returns { error, help } instead of throwing

tasks.ts / members.ts / send.ts:
  - Import from utils.ts, remove all duplicated auth boilerplate
  - All errors return [{ error, help }], no more throw
  - members.ts: add limit arg (was unbounded before)

docs: rename remedy → help across all skill references

* refactor(adapters): migrate pipeline adapters to func() with { error, help } pattern

- slock: agents, channels, messages, servers now use getSlockContext/resolveChannelId
  from utils.ts; error handling uses { error, help } return instead of bare throws
- linux-do: export fetchLinuxDoJson from feed.ts; migrate search, topic, categories,
  tags, user-posts, user-topics from pipeline+throw to func() using fetchLinuxDoJson
- xueqiu: add utils.ts with fetchXueqiuJson helper; migrate hot, feed, search, stock,
  watchlist, hot-stock, groups, kline, earnings-date from pipeline+throw to func()

* fix(output): show error rows in table/csv/markdown when columns declared

When a command declares columns (e.g. ['rank', 'title', 'value']) but
returns an error row ({ error, help }), the declared columns would
render empty cells. Now resolveColumns detects the error key and falls
back to the row's actual keys, making diagnostics visible in all output
formats.

* chore: remove slock adapters from this PR

Slock adapters should be in a separate PR, not bundled with the
adapter refactor and skill docs improvements.
2026-04-10 02:29:35 +08:00
jakevin 56f371fbad docs(skills): improve oneshot & explorer with real-world SaaS patterns (#921)
* docs(skills): add Tier 2.5 localStorage Bearer, SPA discovery, and test standards

From real-world experience building slock.ai CLI adapters:

- oneshot: add network-empty diagnosis, SPA baseURL bundle search, Tier 2.5
  localStorage Bearer template (with multi-tenant X-Server-Id pattern),
  updated auth quick-reference, file path note, opencli browser verify test flow
- explorer: add Tier 2.5 to decision tree and strategy table, update test section
  with opencli browser verify + Done standard, fix Step 5 path to ~/.opencli/clis/,
  add 4 new pitfall rows (SPA HTML, 400 context header, empty network, wrong dir)

* docs(skills): fix path conflict + add anti-change patterns from real adapters

Fix reviewer blocking issue:
- Remove the contradictory "~/.opencli/clis/" note that mixed user-local and
  repo-contributor workflows; replace with explicit two-scenario callout in
  Step 4, Step 5, pitfall table, and oneshot test section
- Template comments in oneshot restored to clis/<site>/<name>.ts (repo path)

Add "抗变更模式" section to explorer, based on opencli's own production code:
- Pattern 1: dynamic queryId discovery (twitter/shared.ts resolveTwitterQueryId)
  — scan loaded JS bundle by operationName (stable) to find queryId (unstable)
- Pattern 2: semantic DOM priority fallback (web/read.ts)
  — article > [role=main] > main > class-hint > body, pick largest text block
- Pattern 3: ordered selector array + timestamp comments (xiaohongshu/publish.ts)
  — first-match wins, comment records UI version and observed attribute values
- Pattern 4: nullish-coalescing field multi-path (xiaohongshu/user-helpers.ts)
  — covers camelCase/snake_case variants without assuming fixed key name

* docs(explorer): split SKILL.md into reference sub-documents

- Shrink main SKILL.md from 994 to 270 lines — core workflow only
- Extract all TS templates (Tier 1~4, pagination) to references/adapter-templates.md
- Add error handling standard: { error, remedy } pattern (remedy > hint)
- Add Tier 2.5 localStorage Bearer template with multi-tenant X-Server-Id example
- Extract cascading requests, tap debug, verbose mode, anti-change patterns to references/advanced-patterns.md
- Extract record workflow to references/record-workflow.md

* docs(skills): fix verify command — split by dev scenario

browser verify only reads ~/.opencli/clis/, not repo's clis/.
Split all verify instructions:
- Repo 贡献: npm run build + opencli <site> <cmd>
- 私人 adapter: opencli browser verify <site>/<name>

Fixes blocker in explorer:L209, L224 and oneshot:L286, L298

* docs(adapter-templates): add utils.ts extraction pattern for same-site adapters

* docs(skills): add decision matrix, stop conditions, sync comments

explorer: add path decision matrix before core workflow
oneshot: add explicit stop/switch conditions (when to escalate to explorer)
both: add keep-in-sync comment on the two-scenario verify block
2026-04-10 01:47:37 +08:00
jakevin f6f7f04ff6 docs(skills): unify browser tool names to opencli browser commands (#920)
* docs(skills): unify browser tool names to `opencli browser` commands

Replace abstract MCP tool names (browser_navigate, browser_snapshot,
browser_network_requests, browser_click, browser_evaluate) with
concrete `opencli browser` CLI commands in explorer and oneshot skills.

This aligns all three browser-related skills into a clear hierarchy:
- opencli-browser: atomic command reference
- opencli-oneshot: 4-step quick generation workflow
- opencli-explorer: full site exploration workflow

* docs(skills): address review — demote explore, fix eval placeholder

1. Demote `opencli explore` from "recommended" to "supplementary helper"
   and make `opencli browser` the explicit primary path for API discovery.
2. Fix `url` undefined variable in eval example — use `<API URL>` placeholder.
2026-04-10 00:39:07 +08:00
jakevin 43f87d2ede fix: restore cross-platform entries in package-lock.json (#919)
Build Chrome Extension / build (push) Has been cancelled
Release / release (push) Has been cancelled
The lockfile generated on Node 25/darwin dropped optional+peer deps
(@emnapi/core, @emnapi/runtime) needed by CI on linux/x64, causing
npm ci to fail.
2026-04-09 23:07:31 +08:00
jakevin b87bbc7107 chore: bump version to 1.7.0 (#917)
Bump CLI, extension package.json, and extension manifest.json to 1.7.0.
Update package-lock.json.
2026-04-09 21:52:40 +08:00
jakevin da86659566 fix(jianyu): avoid early api bucket cutoff (#916) 2026-04-09 21:39:03 +08:00
GanFanNewOrder 606bc59f7b fix(jianyu): stabilize search and add detail extraction contract (#912)
* fix(jianyu): stabilize search and add detail extraction contract

* fix(jianyu): require query evidence for search results

---------

Co-authored-by: 泽加武 <zejiawu@zejiawudeMac-mini.local>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-09 21:28:20 +08:00
jakevin 2ddf571445 feat: auto-close adapter windows, add OPENCLI_WINDOW_FOCUSED, document config (#915)
* feat: auto-close adapter windows, add OPENCLI_WINDOW_FOCUSED, document config

1. Adapter commands now close the automation window immediately after
   completion instead of waiting for the 30s idle timeout.

2. OPENCLI_WINDOW_FOCUSED=1 opens automation windows in the foreground
   (useful for debugging). Default remains background.

3. Add Configuration section to README (EN/ZH) and opencli-usage skill
   listing all stable user-facing environment variables.

* Fix OPENCLI_WINDOW_FOCUSED to be per-request, not frozen at daemon startup

Move env var read from daemon (startup-time constant) to CLI side
(sendCommandRaw), so it works correctly with the persistent daemon model.
Each request now reads the env var fresh and includes windowFocused in
the command payload.
2026-04-09 21:27:02 +08:00
Clearner1 7f31df2912 fix(xiaoe): resolve missing episodes for long courses via auto-scroll (#904)
* fix(xiaoe): resolve missing episodes for long courses by handling lazy load

* fix(xiaoe): keep lazy-load scroll until inner list stabilizes

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-09 21:21:29 +08:00
jakevin b0c9966774 Remove daemon status/restart references from docs and READMEs (#914)
These commands were removed in the persistent daemon refactor.
Only `daemon stop` remains as a user-facing command.
2026-04-09 20:47:07 +08:00
1497 changed files with 86595 additions and 67970 deletions
+7 -8
View File
@@ -3,7 +3,7 @@ name: Build Chrome Extension
on:
push:
branches: [ "main" ]
tags: [ "v*.*.*" ]
tags: [ "ext-v*" ]
paths:
- 'extension/**'
- '.github/workflows/build-extension.yml'
@@ -26,7 +26,7 @@ jobs:
- name: Setup Node.js
uses: actions/setup-node@v6
with:
node-version: 20
node-version: 22
cache: 'npm'
cache-dependency-path: extension/package-lock.json
@@ -44,23 +44,22 @@ jobs:
- name: Create Extension ZIP
run: |
EXT_VERSION=$(node -p "require('./extension/package.json').version")
cd extension-package
zip -r ../opencli-extension.zip .
zip -r ../opencli-extension-v${EXT_VERSION}.zip .
- name: Upload Artifacts (Action Run)
uses: actions/upload-artifact@v7
with:
name: opencli-extension-build
path: |
opencli-extension.zip
path: opencli-extension-v*.zip
retention-days: 7
- name: Attach to GitHub Release
if: startsWith(github.ref, 'refs/tags/')
uses: softprops/action-gh-release@v2.6.1
uses: softprops/action-gh-release@v3.0.0
with:
files: |
opencli-extension.zip
files: opencli-extension-v*.zip
draft: false
prerelease: false
env:
+2 -6
View File
@@ -47,7 +47,7 @@ jobs:
fail-fast: false
matrix:
os: ${{ (github.event_name == 'push' || github.event_name == 'schedule' || github.event_name == 'workflow_dispatch') && fromJSON('["ubuntu-latest","macos-latest","windows-latest"]') || fromJSON('["ubuntu-latest"]') }}
node-version: ${{ (github.event_name == 'push' || github.event_name == 'schedule' || github.event_name == 'workflow_dispatch') && fromJSON('["20","22"]') || fromJSON('["22"]') }}
node-version: ${{ (github.event_name == 'push' || github.event_name == 'schedule' || github.event_name == 'workflow_dispatch') && fromJSON('["22"]') || fromJSON('["22"]') }}
shard: [1, 2]
steps:
- uses: actions/checkout@v6
@@ -61,7 +61,7 @@ jobs:
run: npm ci
- name: Run unit tests (Node ${{ matrix.node-version }}, shard ${{ matrix.shard }}/2)
run: npm test -- --reporter=verbose --shard=${{ matrix.shard }}/2
run: npx vitest run --project unit --project extension --reporter=verbose --shard=${{ matrix.shard }}/2
# ── Bun compatibility check ──
bun-test:
@@ -136,12 +136,8 @@ jobs:
run: |
xvfb-run --auto-servernum --server-args="-screen 0 1280x720x24" \
npx vitest run tests/smoke/ --reporter=verbose
env:
OPENCLI_BROWSER_EXECUTABLE_PATH: ${{ steps.setup-chrome.outputs.chrome-path }}
- name: Run smoke tests (macOS / Windows)
if: runner.os != 'Linux'
run: npx vitest run tests/smoke/ --reporter=verbose
env:
OPENCLI_BROWSER_EXECUTABLE_PATH: ${{ steps.setup-chrome.outputs.chrome-path }}
timeout-minutes: 15
-4
View File
@@ -64,11 +64,7 @@ jobs:
run: |
xvfb-run --auto-servernum --server-args="-screen 0 1280x720x24" \
npx vitest run tests/e2e/ --reporter=verbose
env:
OPENCLI_BROWSER_EXECUTABLE_PATH: ${{ steps.setup-chrome.outputs.chrome-path }}
- name: Run E2E tests (macOS / Windows)
if: runner.os != 'Linux'
run: npx vitest run tests/e2e/ --reporter=verbose
env:
OPENCLI_BROWSER_EXECUTABLE_PATH: ${{ steps.setup-chrome.outputs.chrome-path }}
+21 -1
View File
@@ -26,10 +26,30 @@ jobs:
- name: Type check
run: npx tsc --noEmit
- name: Install extension dependencies
run: npm ci
working-directory: extension
- name: Build extension
run: npm run build
working-directory: extension
- name: Package extension
run: npm run package:release -- --out ../extension-package
working-directory: extension
- name: Create extension ZIP
run: |
EXT_VERSION=$(node -p "require('./extension/package.json').version")
cd extension-package
zip -r ../opencli-extension-v${EXT_VERSION}.zip .
- name: Create GitHub Release
uses: softprops/action-gh-release@v2.6.1
uses: softprops/action-gh-release@v3.0.0
with:
generate_release_notes: true
files: |
opencli-extension-v*.zip
- name: Publish to npm
run: npm publish --provenance --access public
+1
View File
@@ -3,6 +3,7 @@ dist/
!extension/dist/
*.tsbuildinfo
.opencli/
.worktrees/
.mcp.json
*.log
.DS_Store
+88
View File
@@ -1,5 +1,93 @@
# Changelog
## [1.7.0](https://github.com/jackwener/opencli/compare/v1.6.1...v1.7.0) (2026-04-11)
This is a major release with significant internal architecture changes.
Adapter code, validation, and error handling have been modernized.
### ⚠ BREAKING CHANGES
* **Node.js >= 21 required** — `import.meta.dirname` is used in core modules; Node 20 and below will fail at startup.
* **YAML adapters deprecated** — YAML-based `.yaml` adapters are no longer loaded. Existing YAML adapters must be converted to JS via `cli()` API. A deprecation warning is emitted if `.yaml` files are detected.
* **`.ts` adapters no longer loaded at runtime** — The runtime only discovers `.js` files. If you have `.ts` adapters in `~/.opencli/clis/`, compile them to `.js` or rewrite using plain JS. A warning is printed when `.ts` files without a matching `.js` are found.
* **Error output format changed** — All errors are now emitted as a structured YAML envelope to stderr. Scripts parsing stdout for `[{error, help}]` must switch to stderr / exit code. ([#923](https://github.com/jackwener/opencli/issues/923))
* **`tabId` replaced by `targetId`** — Cross-layer page identity now uses `targetId`. Extensions and plugins referencing `tabId` must update. ([#899](https://github.com/jackwener/opencli/issues/899))
* **`operate` renamed to `browser`** — All `opencli operate` commands are now `opencli browser`. ([#883](https://github.com/jackwener/opencli/issues/883))
### Features
* **auto-close adapter windows** — Browser tabs opened by adapters are automatically closed after execution; configurable via `OPENCLI_WINDOW_FOCUSED`. ([#915](https://github.com/jackwener/opencli/issues/915))
* **Self-Repair protocol** — Automatic adapter fixing when commands fail. ([#866](https://github.com/jackwener/opencli/issues/866))
* **EarlyHint callback** — Cost gating channel for generate pipeline. ([#882](https://github.com/jackwener/opencli/issues/882))
* **verified generate pipeline** — Structured contract for AI-driven adapter generation. ([#878](https://github.com/jackwener/opencli/issues/878))
* **structured diagnostic output** — AI-driven adapter repair gets structured diagnostics. ([#802](https://github.com/jackwener/opencli/issues/802))
* **auto-downgrade to YAML in non-TTY** — Machine-readable output when piped. ([#737](https://github.com/jackwener/opencli/issues/737))
* **Browser Use improvements** — Better click/type/state handling for browser automation. ([#707](https://github.com/jackwener/opencli/issues/707))
* **CDP session-level network capture** — Full network capture support for CDPPage. ([#815](https://github.com/jackwener/opencli/issues/815), [#816](https://github.com/jackwener/opencli/issues/816))
* **AutoResearch framework** — V2EX/Zhihu test suites (194 tasks). ([#731](https://github.com/jackwener/opencli/issues/731), [#717](https://github.com/jackwener/opencli/issues/717), [#741](https://github.com/jackwener/opencli/issues/741))
* **new adapters:** Gitee ([#845](https://github.com/jackwener/opencli/issues/845)), 闲鱼 ([#696](https://github.com/jackwener/opencli/issues/696)), 1688 ([#650](https://github.com/jackwener/opencli/issues/650), [#820](https://github.com/jackwener/opencli/issues/820)), LessWrong ([#773](https://github.com/jackwener/opencli/issues/773)), 虎扑 ([#751](https://github.com/jackwener/opencli/issues/751)), 小鹅通 ([#617](https://github.com/jackwener/opencli/issues/617)), 元宝 ([#693](https://github.com/jackwener/opencli/issues/693)), 即梦 ([#897](https://github.com/jackwener/opencli/issues/897), [#895](https://github.com/jackwener/opencli/issues/895)), Quark Drive ([#858](https://github.com/jackwener/opencli/issues/858)), GitHub Trending/Binance/Weather ([#214](https://github.com/jackwener/opencli/issues/214))
* **adapter enhancements:** Instagram post/reel/story/note ([#671](https://github.com/jackwener/opencli/issues/671)), Twitter image posts/replies ([#666](https://github.com/jackwener/opencli/issues/666), [#756](https://github.com/jackwener/opencli/issues/756)), 知乎 interactions ([#868](https://github.com/jackwener/opencli/issues/868)), Bilibili b23.tv short URL ([#740](https://github.com/jackwener/opencli/issues/740)), 雪球 kline/groups ([#809](https://github.com/jackwener/opencli/issues/809)), Amazon unified ranking ([#724](https://github.com/jackwener/opencli/issues/724)), Gemini deep-research ([#778](https://github.com/jackwener/opencli/issues/778)), 新浪财经热搜 ([#736](https://github.com/jackwener/opencli/issues/736)), linux-do topic split ([#821](https://github.com/jackwener/opencli/issues/821)), JD/淘宝/CNKI revived ([#248](https://github.com/jackwener/opencli/issues/248))
### Bug Fixes
* **security:** escape codegen strings and redact diagnostic body ([#930](https://github.com/jackwener/opencli/issues/930))
* **bilibili:** add missing domain for following cli ([#947](https://github.com/jackwener/opencli/issues/947))
* clean up stale `.ts` adapter files during upgrade ([#948](https://github.com/jackwener/opencli/issues/948))
* clean up legacy shim files and stale tmp files on upgrade ([#934](https://github.com/jackwener/opencli/issues/934))
* address deep review findings (security, correctness, consistency) ([#935](https://github.com/jackwener/opencli/issues/935))
* batch quality improvements — dedupe completion, unify logging, fix docs ([#945](https://github.com/jackwener/opencli/issues/945))
* graceful fallback when extension lacks network-capture support ([#865](https://github.com/jackwener/opencli/issues/865))
* handle missing electron executable gracefully ([#747](https://github.com/jackwener/opencli/issues/747))
* recover drifted tabs instead of abandoning them ([#715](https://github.com/jackwener/opencli/issues/715))
* retry on "No window with id" CDP error ([#892](https://github.com/jackwener/opencli/issues/892))
* **launcher:** graceful degradation and manual CDP override for Windows ([#744](https://github.com/jackwener/opencli/issues/744))
* **xiaohongshu:** scope note interaction selectors, replace blind retry with MutationObserver ([#839](https://github.com/jackwener/opencli/issues/839), [#730](https://github.com/jackwener/opencli/issues/730))
* **twitter:** relax reply composer timeout, use composer for text replies ([#862](https://github.com/jackwener/opencli/issues/862), [#860](https://github.com/jackwener/opencli/issues/860))
* **doubao:** preserve image URLs, connect to correct CDP target ([#708](https://github.com/jackwener/opencli/issues/708), [#674](https://github.com/jackwener/opencli/issues/674))
* **gemini:** stabilize ask reply state handling ([#735](https://github.com/jackwener/opencli/issues/735))
* **douban:** fix marks pagination and improve subject data extraction ([#752](https://github.com/jackwener/opencli/issues/752))
* **jianyu:** avoid early API bucket cutoff, stabilize search ([#916](https://github.com/jackwener/opencli/issues/916), [#912](https://github.com/jackwener/opencli/issues/912))
* **xiaoe:** resolve missing episodes for long courses via auto-scroll ([#904](https://github.com/jackwener/opencli/issues/904))
### Refactoring
* **adapters:** convert adapter layer from TypeScript to JavaScript ([#928](https://github.com/jackwener/opencli/issues/928))
* **adapters:** migrate all CLI adapters from YAML to TypeScript, then to JS ([#887](https://github.com/jackwener/opencli/issues/887), [#922](https://github.com/jackwener/opencli/issues/922))
* **validate:** switch from YAML-file scanning to registry-based validation ([#943](https://github.com/jackwener/opencli/issues/943))
* **strategy:** normalize strategy into runtime fields at registration time ([#941](https://github.com/jackwener/opencli/issues/941))
* **errors:** unify error output as YAML envelope to stderr ([#923](https://github.com/jackwener/opencli/issues/923))
* **daemon:** make daemon persistent, remove idle timeout ([#913](https://github.com/jackwener/opencli/issues/913))
* **browser:** unify browser error classification and deduplicate retry logic ([#908](https://github.com/jackwener/opencli/issues/908))
* **monorepo:** adapter separation — `clis/` at root ([#782](https://github.com/jackwener/opencli/issues/782))
* rename `operate` to `browser` ([#883](https://github.com/jackwener/opencli/issues/883))
* eliminate `any` types in core files ([#886](https://github.com/jackwener/opencli/issues/886))
* migrate adapter imports to package exports ([#795](https://github.com/jackwener/opencli/issues/795))
### Performance
* **P0 optimizations** — faster startup, reduced overhead ([#944](https://github.com/jackwener/opencli/issues/944))
* fast-path completion/version/shell-scripts to bypass full discovery ([#898](https://github.com/jackwener/opencli/issues/898))
* optimize browser pipeline — tab query dedup, parallel stealth, incremental snapshots ([#713](https://github.com/jackwener/opencli/issues/713))
* reduce round-trips in browser command hot path ([#712](https://github.com/jackwener/opencli/issues/712))
* skip blank page on first browser command ([#710](https://github.com/jackwener/opencli/issues/710))
### Documentation
* restructure README narrative ([#885](https://github.com/jackwener/opencli/issues/885))
* add Android Chrome usage guide ([#687](https://github.com/jackwener/opencli/issues/687))
* add Electron app CLI quickstart guide
* fix stale `.ts` references across skills and docs ([#954](https://github.com/jackwener/opencli/issues/954))
* unify skill command references and merge opencli-generate into opencli-explorer ([#891](https://github.com/jackwener/opencli/issues/891), [#894](https://github.com/jackwener/opencli/issues/894))
### Upgrade Guide
1. **Update Node.js** to v21 or later (v22 LTS recommended).
2. **Run `npm install -g @jackwener/opencli@latest`** — the preuninstall hook gracefully stops the old daemon; the first browser command after upgrade auto-restarts it.
3. **If you have custom `.ts` adapters** in `~/.opencli/clis/`, rename or compile them to `.js`. A warning will be printed on startup if stale `.ts` files are detected.
4. **If you have custom `.yaml` adapters**, convert them to JS using the `cli()` API (see `skills/opencli-explorer/references/adapter-templates.md`).
5. **If you parse error output from stdout**, switch to stderr. Errors are now structured YAML envelopes with typed exit codes.
## [1.6.1](https://github.com/jackwener/opencli/compare/v1.6.0...v1.6.1) (2026-04-02)
+7 -8
View File
@@ -18,7 +18,6 @@ npm run build
# 4. Run a few checks
npx tsc --noEmit
npm test
npm run test:adapter
# 5. Link globally (optional, for testing `opencli` command)
npm link
@@ -30,7 +29,7 @@ All adapters use TypeScript. Use the pipeline API for data-fetching commands, an
### Pipeline Adapter (Recommended for data-fetching commands)
Create a file like `clis/<site>/<command>.ts`:
Create a file like `clis/<site>/<command>.js`:
```typescript
import { cli, Strategy } from '@jackwener/opencli/registry';
@@ -60,11 +59,11 @@ cli({
});
```
See [`hackernews/top.ts`](clis/hackernews/top.ts) for a real example.
See [`hackernews/top.js`](clis/hackernews/top.js) for a real example.
### func() Adapter (For complex browser interactions)
Create a file like `clis/<site>/<command>.ts`:
Create a file like `clis/<site>/<command>.js`:
```typescript
import { cli, Strategy } from '@jackwener/opencli/registry';
@@ -152,8 +151,8 @@ args: [
See [TESTING.md](./TESTING.md) for the full guide and exact test locations.
```bash
npm test # Core unit tests (non-adapter)
npm run test:adapter # Focused adapter tests: zhihu/twitter/reddit/bilibili
npm test # Default local gate: unit + extension + adapter tests
npm run test:adapter # Adapter-only project (useful while iterating on adapters)
npx vitest run tests/e2e/ # E2E tests
npx vitest run # All tests
```
@@ -186,8 +185,8 @@ Common scopes: site name (`twitter`, `reddit`) or module name (`browser`, `pipel
3. Run the checks that apply:
```bash
npx tsc --noEmit # Type check
npm test # Core unit tests
npm run test:adapter # Focused adapter tests (if you touched adapter logic)
npm test # Default local gate: unit + extension + adapter
npm run test:adapter # Adapter-only project (optional while iterating on adapters)
opencli validate # Adapter validation
```
4. Commit using conventional commit format
+43 -27
View File
@@ -16,24 +16,16 @@ OpenCLI gives you one surface for three different kinds of automation:
It also works as a **CLI hub** for local tools such as `gh`, `docker`, and other binaries you register yourself, plus **desktop app adapters** for Electron apps like Cursor, Codex, Antigravity, ChatGPT, and Notion.
## Why OpenCLI
---
## Highlights
- **CLI All Electron** — CLI-ify apps like Antigravity Ultra! Now AI can control itself natively.
- **Browser Automation** — `browser` gives AI agents direct browser control: click, type, extract, screenshot — any interaction, fully scriptable.
- **Website → CLI** — Turn any website into a deterministic CLI: 70+ pre-built adapters, or crystallize your own with `opencli record`.
- **Desktop App Control** — Drive Electron apps (Cursor, Codex, ChatGPT, Notion, etc.) directly from the terminal via CDP.
- **Browser Automation** — `browser` gives AI agents direct browser control: click, type, extract, screenshot — fully scriptable.
- **Website → CLI** — Turn any website into a deterministic CLI: 87+ pre-built adapters, or generate your own with `opencli generate`.
- **Account-safe** — Reuses Chrome/Chromium logged-in state; your credentials never leave the browser.
- **Anti-detection built-in** — Patches `navigator.webdriver`, stubs `window.chrome`, fakes plugin lists, cleans ChromeDriver/Playwright globals, and strips CDP frames from Error stack traces. Extensive anti-fingerprinting and risk-control evasion measures baked in at every layer.
- **AI Agent ready** — `explore` discovers APIs, `synthesize` generates adapters, `cascade` finds auth strategies, `browser` controls the browser directly.
- **External CLI Hub** — Discover, auto-install, and passthrough commands to any external CLI (gh, obsidian, docker, etc). Zero setup.
- **Self-healing setup** — `opencli doctor` diagnoses and auto-starts the daemon, extension, and live browser connectivity.
- **Dynamic Loader** — Simply drop `.ts` adapters into the `clis/` folder for auto-registration.
- **CLI Hub** — Discover, auto-install, and passthrough commands to any external CLI (gh, docker, obsidian, etc).
- **Zero LLM cost** — No tokens consumed at runtime. Run 10,000 times and pay nothing.
- **Deterministic** — Same command, same output schema, every time. Pipeable, scriptable, CI-friendly.
- **Broad coverage** — 79+ sites across global and Chinese platforms (Bilibili, Zhihu, Xiaohongshu, Reddit, HackerNews, and more), plus desktop Electron apps via CDP.
---
@@ -49,7 +41,7 @@ npm install -g @jackwener/opencli
OpenCLI connects to Chrome/Chromium through a lightweight Browser Bridge extension plus a small local daemon. The daemon auto-starts when needed.
1. Download the latest `opencli-extension.zip` from the GitHub [Releases page](https://github.com/jackwener/opencli/releases).
1. Download the latest `opencli-extension-v{version}.zip` from the GitHub [Releases page](https://github.com/jackwener/opencli/releases).
2. Unzip it, open `chrome://extensions`, and enable **Developer mode**.
3. Click **Load unpacked** and select the unzipped folder.
@@ -57,7 +49,6 @@ OpenCLI connects to Chrome/Chromium through a lightweight Browser Bridge extensi
```bash
opencli doctor
opencli daemon status
```
### 4. Run your first commands
@@ -75,7 +66,7 @@ Use OpenCLI directly when you want a reliable command instead of a live browser
- `opencli list` shows every registered command.
- `opencli <site> <command>` runs a built-in or generated adapter.
- `opencli register mycli` exposes a local CLI through the same discovery surface.
- `opencli daemon status` and `opencli doctor` help diagnose browser connectivity.
- `opencli doctor` helps diagnose browser connectivity.
## For AI Agents
@@ -121,7 +112,7 @@ Use site-specific commands such as `opencli hackernews top` or `opencli reddit h
Use these commands when the site you need is not covered yet:
- `explore` inspects the page, network activity, and capability surface.
- `synthesize` turns exploration artifacts into evaluate-based YAML adapters.
- `synthesize` turns exploration artifacts into evaluate-based JS adapters.
- `generate` runs the verified generation path and returns either a usable command or a structured explanation of why completion was blocked or needs human review.
### `cascade`: auth strategy discovery
@@ -137,11 +128,26 @@ OpenCLI is not only for websites. It can also:
## Prerequisites
- **Node.js**: >= 20.0.0 (or **Bun** >= 1.0)
- **Node.js**: >= 21.0.0 (or **Bun** >= 1.0)
- **Chrome or Chromium** running and logged into the target site for browser-backed commands
> **Important**: Browser-backed commands reuse your Chrome/Chromium login session. If you get empty data or permission-like failures, first confirm the site is already open and authenticated in Chrome/Chromium.
## Configuration
| Variable | Default | Description |
|----------|---------|-------------|
| `OPENCLI_DAEMON_PORT` | `19825` | HTTP port for the daemon-extension bridge |
| `OPENCLI_WINDOW_FOCUSED` | `false` | Set to `1` to open automation windows in the foreground (useful for debugging) |
| `OPENCLI_BROWSER_CONNECT_TIMEOUT` | `30` | Seconds to wait for browser connection |
| `OPENCLI_BROWSER_COMMAND_TIMEOUT` | `60` | Seconds to wait for a single browser command |
| `OPENCLI_BROWSER_EXPLORE_TIMEOUT` | `120` | Seconds to wait for explore/record operations |
| `OPENCLI_CDP_ENDPOINT` | — | Chrome DevTools Protocol endpoint for remote browser or Electron apps |
| `OPENCLI_CDP_TARGET` | — | Filter CDP targets by URL substring (e.g. `detail.1688.com`) |
| `OPENCLI_VERBOSE` | `false` | Enable verbose logging (`-v` flag also works) |
| `OPENCLI_DIAGNOSTIC` | `false` | Set to `1` to capture structured diagnostic context on failures |
| `DEBUG_SNAPSHOT` | — | Set to `1` for DOM snapshot debug output |
## Update
```bash
@@ -185,7 +191,7 @@ To load the source Browser Bridge extension:
| **bilibili** | `hot` `search` `history` `feed` `ranking` `download` `comments` `dynamic` `favorite` `following` `me` `subtitle` `user-videos` |
| **tieba** | `hot` `posts` `search` `read` |
| **hupu** | `hot` `search` `detail` `mentions` `reply` `like` `unlike` |
| **twitter** | `trending` `search` `timeline` `bookmarks` `post` `download` `profile` `article` `like` `likes` `notifications` `reply` `reply-dm` `thread` `follow` `unfollow` `followers` `following` `block` `unblock` `bookmark` `unbookmark` `delete` `hide-reply` `accept` |
| **twitter** | `trending` `search` `timeline` `lists` `bookmarks` `post` `download` `profile` `article` `like` `likes` `notifications` `reply` `reply-dm` `thread` `follow` `unfollow` `followers` `following` `block` `unblock` `bookmark` `unbookmark` `delete` `hide-reply` `accept` |
| **reddit** | `hot` `frontpage` `popular` `search` `subreddit` `read` `user` `user-posts` `user-comments` `upvote` `upvoted` `save` `saved` `comment` `subscribe` |
| **zhihu** | `hot` `search` `question` `download` `follow` `like` `favorite` `comment` `answer` |
| **amazon** | `bestsellers` `search` `product` `offer` `discussion` `movers-shakers` `new-releases` |
@@ -198,8 +204,12 @@ To load the source Browser Bridge extension:
| **xianyu** | `search` `item` `chat` |
| **xiaoe** | `courses` `detail` `catalog` `play-url` `content` |
| **quark** | `ls` `mkdir` `mv` `rename` `rm` `save` `share-tree` |
| **uiverse** | `code` `preview` |
| **xiaoyuzhou** | `podcast` `podcast-episodes` `episode` `download` `transcript*` |
79+ adapters in total — **[→ see all supported sites & commands](./docs/adapters/index.md)**
87+ adapters in total — **[→ see all supported sites & commands](./docs/adapters/index.md)**
`*` `opencli xiaoyuzhou transcript` requires local Xiaoyuzhou credentials in `~/.opencli/xiaoyuzhou.json`.
## CLI Hub
@@ -230,7 +240,7 @@ Control Electron desktop apps directly from the terminal. Each adapter has its o
| **Cursor** | Control Cursor IDE — Composer, chat, code extraction | [Doc](./docs/adapters/desktop/cursor.md) |
| **Codex** | Drive OpenAI Codex CLI agent headlessly | [Doc](./docs/adapters/desktop/codex.md) |
| **Antigravity** | Control Antigravity Ultra from terminal | [Doc](./docs/adapters/desktop/antigravity.md) |
| **ChatGPT** | Automate ChatGPT macOS desktop app | [Doc](./docs/adapters/desktop/chatgpt.md) |
| **ChatGPT App** | Automate ChatGPT macOS desktop app | [Doc](./docs/adapters/desktop/chatgpt-app.md) |
| **ChatWise** | Multi-LLM client (GPT-4, Claude, Gemini) | [Doc](./docs/adapters/desktop/chatwise.md) |
| **Notion** | Search, read, write Notion pages | [Doc](./docs/adapters/desktop/notion.md) |
| **Discord** | Discord Desktop — messages, channels, servers | [Doc](./docs/adapters/desktop/discord.md) |
@@ -250,18 +260,24 @@ OpenCLI supports downloading images, videos, and articles from supported platfor
| **douban** | Images | Poster / still image lists |
| **pixiv** | Images | Original-quality illustrations, multi-page |
| **1688** | Images, Videos | Downloads page-visible product media from item pages |
| **xiaoyuzhou** | Audio, Transcript | Downloads episode audio from public pages and transcript JSON/text with local credentials |
| **zhihu** | Articles (Markdown) | Exports with optional image download |
| **weixin** | Articles (Markdown) | WeChat Official Account articles |
For video downloads, install `yt-dlp` first: `brew install yt-dlp`
```bash
opencli xiaohongshu download abc123 --output ./xhs
opencli xiaohongshu download "https://www.xiaohongshu.com/search_result/<id>?xsec_token=..." --output ./xhs
opencli xiaohongshu download "https://xhslink.com/..." --output ./xhs
opencli bilibili download BV1xxx --output ./bilibili
opencli twitter download elonmusk --limit 20 --output ./twitter
opencli 1688 download 841141931191 --output ./1688-downloads
opencli xiaoyuzhou download 69b3b675772ac2295bfc01d0 --output ./xiaoyuzhou
opencli xiaoyuzhou transcript 69dd0c98e2c8be31551f6a33 --output ./xiaoyuzhou-transcripts
```
`opencli xiaoyuzhou transcript` requires local Xiaoyuzhou credentials in `~/.opencli/xiaoyuzhou.json`.
## Output Formats
All built-in commands support `--format` / `-f` with `table` (default), `json`, `yaml`, `md`, and `csv`.
@@ -307,10 +323,10 @@ opencli plugin uninstall my-tool
| Plugin | Type | Description |
|--------|------|-------------|
| [opencli-plugin-github-trending](https://github.com/ByteYue/opencli-plugin-github-trending) | TS | GitHub Trending repositories |
| [opencli-plugin-hot-digest](https://github.com/ByteYue/opencli-plugin-hot-digest) | TS | Multi-platform trending aggregator |
| [opencli-plugin-juejin](https://github.com/Astro-Han/opencli-plugin-juejin) | TS | 稀土掘金 (Juejin) hot articles |
| [opencli-plugin-vk](https://github.com/flobo3/opencli-plugin-vk) | TS | VK (VKontakte) wall, feed, and search |
| [opencli-plugin-github-trending](https://github.com/ByteYue/opencli-plugin-github-trending) | JS | GitHub Trending repositories |
| [opencli-plugin-hot-digest](https://github.com/ByteYue/opencli-plugin-hot-digest) | JS | Multi-platform trending aggregator |
| [opencli-plugin-juejin](https://github.com/Astro-Han/opencli-plugin-juejin) | JS | 稀土掘金 (Juejin) hot articles |
| [opencli-plugin-vk](https://github.com/flobo3/opencli-plugin-vk) | JS | VK (VKontakte) wall, feed, and search |
See [Plugins Guide](./docs/guide/plugins.md) for creating your own plugin.
@@ -322,7 +338,7 @@ See [Plugins Guide](./docs/guide/plugins.md) for creating your own plugin.
```bash
opencli explore https://example.com --site mysite # Discover APIs + capabilities
opencli synthesize mysite # Generate TS adapters
opencli synthesize mysite # Generate JS adapters
opencli generate https://example.com --goal "hot" # One-shot: explore → synthesize → register
opencli cascade https://api.example.com/data # Auto-probe: PUBLIC → COOKIE → HEADER
```
@@ -336,7 +352,7 @@ See **[TESTING.md](./TESTING.md)** for how to run and write tests.
- **"Extension not connected"** — Ensure the Browser Bridge extension is installed and **enabled** in `chrome://extensions` in Chrome or Chromium.
- **"attach failed: Cannot access a chrome-extension:// URL"** — Another extension may be interfering. Try disabling other extensions temporarily.
- **Empty data or 'Unauthorized' error** — Your Chrome/Chromium login session may have expired. Navigate to the target site and log in again.
- **Node API errors** — Ensure Node.js >= 20. Some dependencies require modern Node APIs.
- **Node API errors** — Ensure Node.js >= 21. Some features require `node:util` styleText (stable in Node 21+).
- **Daemon issues** — Check status: `curl localhost:19825/status` · View logs: `curl localhost:19825/logs`
## Star History
+55 -25
View File
@@ -10,20 +10,22 @@
OpenCLI 可以用同一套 CLI 做三类事情:
- **直接使用现成适配器**:B站、知乎、小红书、Twitter/X、Reddit、HackerNews 等 [79+ 站点](#内置命令) 开箱即用。
- **直接使用现成适配器**:B站、知乎、小红书、Twitter/X、Reddit、HackerNews 等 [87+ 站点](#内置命令) 开箱即用。
- **直接驱动浏览器**:用 `opencli browser` 让 AI Agent 实时点击、输入、提取、截图、检查页面状态。
- **把新网站生成成 CLI**:通过 `explore``synthesize``generate``cascade` 从真实页面行为推导出新的适配器。
除了网站能力,OpenCLI 还是一个 **CLI 枢纽**:你可以把 `gh``docker` 等本地工具统一注册到 `opencli` 下,也可以通过桌面端适配器控制 Cursor、Codex、Antigravity、ChatGPT、Notion 等 Electron 应用。
## 为什么是 OpenCLI
## 亮点
- **同一个心智模型**:网站、浏览器自动化、Electron 应用、本地 CLI 都走同一个入口
- **复用真实会话**:浏览器命令直接使用你已经登录的 Chrome/Chromium,而不是重新造一套认证
- **输出稳定**:适配器命令返回固定结构,适合 shell、脚本、CI 和 AI Agent 工具调用
- **面向 AI Agent**`browser` 负责实时操作,`explore` 负责探索接口,`synthesize` 负责生成适配器,`cascade` 负责探测认证路径
- **运行成本低**:已有命令运行时不消耗模型 token
- **天然可扩展**:既能用内置能力,也能注册本地 CLI,或直接往 `clis/``.ts` 适配器
- **桌面应用控制** — 通过 CDP 直接在终端驱动 Electron 应用(Cursor、Codex、ChatGPT、Notion 等)
- **浏览器自动化** — `browser` 让 AI Agent 直接控制浏览器:点击、输入、提取、截图,完全可编程
- **网站 → CLI** — 把任何网站变成确定性 CLI:87+ 内置适配器,或用 `opencli generate` 生成新的
- **账号安全** — 复用 Chrome/Chromium 登录态,凭证永远不会离开浏览器
- **面向 AI Agent** — `explore` 发现 API`synthesize` 生成适配器,`cascade` 探测认证策略,`browser` 直接控制浏览器
- **CLI 枢纽** — 统一发现、自动安装、纯透传任何外部 CLI(gh、docker、obsidian 等)
- **零 LLM 成本** — 运行时不消耗模型 token,跑 10,000 次也不花一分钱。
- **确定性输出** — 相同命令,相同输出结构,每次一致。可管道、可脚本、CI 友好。
## 快速开始
@@ -37,7 +39,7 @@ npm install -g @jackwener/opencli
OpenCLI 通过轻量 Browser Bridge 扩展和本地微型 daemon 与 Chrome/Chromium 通信。daemon 会按需自动启动。
1. 到 GitHub [Releases 页面](https://github.com/jackwener/opencli/releases) 下载最新的 `opencli-extension.zip`
1. 到 GitHub [Releases 页面](https://github.com/jackwener/opencli/releases) 下载最新的 `opencli-extension-v{version}.zip`
2. 解压后打开 `chrome://extensions`,启用 **开发者模式**
3. 点击 **加载已解压的扩展程序**,选择解压后的目录。
@@ -45,7 +47,6 @@ OpenCLI 通过轻量 Browser Bridge 扩展和本地微型 daemon 与 Chrome/Chro
```bash
opencli doctor
opencli daemon status
```
### 4. 跑第一个命令
@@ -63,7 +64,7 @@ opencli bilibili hot --limit 5
- `opencli list` 查看当前所有命令
- `opencli <site> <command>` 调用内置或生成好的适配器
- `opencli register mycli` 把本地 CLI 接入同一发现入口
- `opencli doctor` / `opencli daemon status` 处理浏览器连通性问题
- `opencli doctor` 处理浏览器连通性问题
## 给 AI Agent
@@ -125,11 +126,26 @@ OpenCLI 不只是网站 CLI,还可以:
## 前置要求
- **Node.js**: >= 20.0.0
- **Node.js**: >= 21.0.0
- 浏览器型命令需要 Chrome 或 Chromium 处于运行中,并已登录目标网站
> **重要**:浏览器型命令直接复用你的 Chrome/Chromium 登录态。如果拿到空数据或出现权限类失败,先确认目标站点已经在浏览器里打开并完成登录。
## 配置
| 变量 | 默认值 | 说明 |
|------|--------|------|
| `OPENCLI_DAEMON_PORT` | `19825` | daemon-extension 通信端口 |
| `OPENCLI_WINDOW_FOCUSED` | `false` | 设为 `1` 时 automation 窗口在前台打开(适合调试) |
| `OPENCLI_BROWSER_CONNECT_TIMEOUT` | `30` | 浏览器连接超时(秒) |
| `OPENCLI_BROWSER_COMMAND_TIMEOUT` | `60` | 单个浏览器命令超时(秒) |
| `OPENCLI_BROWSER_EXPLORE_TIMEOUT` | `120` | explore/record 操作超时(秒) |
| `OPENCLI_CDP_ENDPOINT` | — | Chrome DevTools Protocol 端点,用于远程浏览器或 Electron 应用 |
| `OPENCLI_CDP_TARGET` | — | 按 URL 子串过滤 CDP target(如 `detail.1688.com` |
| `OPENCLI_VERBOSE` | `false` | 启用详细日志(`-v` 也可以) |
| `OPENCLI_DIAGNOSTIC` | `false` | 设为 `1` 时在失败时输出结构化诊断上下文 |
| `DEBUG_SNAPSHOT` | — | 设为 `1` 输出 DOM 快照调试信息 |
## 更新
```bash
@@ -171,7 +187,7 @@ npm link
| 站点 | 命令 | 模式 |
|------|------|------|
| **twitter** | `trending` `bookmarks` `profile` `search` `timeline` `thread` `following` `followers` `notifications` `post` `reply` `delete` `like` `article` `follow` `unfollow` `bookmark` `unbookmark` `download` `accept` `reply-dm` `block` `unblock` `hide-reply` | 浏览器 |
| **twitter** | `trending` `search` `timeline` `lists` `bookmarks` `profile` `thread` `following` `followers` `notifications` `post` `reply` `delete` `like` `likes` `article` `follow` `unfollow` `bookmark` `unbookmark` `download` `accept` `reply-dm` `block` `unblock` `hide-reply` | 浏览器 |
| **reddit** | `hot` `frontpage` `popular` `search` `subreddit` `read` `user` `user-posts` `user-comments` `upvote` `save` `comment` `subscribe` `saved` `upvoted` | 浏览器 |
| **tieba** | `hot` `posts` `search` `read` | 浏览器 |
| **hupu** | `hot` `search` `detail` `mentions` `reply` `like` `unlike` | 浏览器 |
@@ -186,15 +202,16 @@ npm link
| **v2ex** | `hot` `latest` `topic` `node` `user` `member` `replies` `nodes` `daily` `me` `notifications` | 公开 / 浏览器 |
| **xueqiu** | `feed` `hot-stock` `hot` `search` `stock` `comments` `watchlist` `earnings-date` `fund-holdings` `fund-snapshot` | 浏览器 |
| **antigravity** | `status` `send` `read` `new` `dump` `extract-code` `model` `watch` | 桌面端 |
| **chatgpt** | `status` `new` `send` `read` `ask` `model` | 桌面端 |
| **chatgpt-app** | `status` `new` `send` `read` `ask` `model` | 桌面端 |
| **xiaohongshu** | `search` `notifications` `feed` `user` `download` `publish` `creator-notes` `creator-note-detail` `creator-notes-summary` `creator-profile` `creator-stats` | 浏览器 |
| **xiaoe** | `courses` `detail` `catalog` `play-url` `content` | 浏览器 |
| **quark** | `ls` `mkdir` `mv` `rename` `rm` `save` `share-tree` | 浏览器 |
| **uiverse** | `code` `preview` | 浏览器 |
| **apple-podcasts** | `search` `episodes` `top` | 公开 |
| **xiaoyuzhou** | `podcast` `podcast-episodes` `episode` | 公开 |
| **xiaoyuzhou** | `podcast` `podcast-episodes` `episode` `download` `transcript*` | 公开 |
| **zhihu** | `hot` `search` `question` `download` `follow` `like` `favorite` `comment` `answer` | 浏览器 |
| **weixin** | `download` | 浏览器 |
| **youtube** | `search` `video` `transcript` | 浏览器 |
| **youtube** | `search` `video` `transcript` `comments` `channel` `playlist` `feed` `history` `watch-later` `subscriptions` `like` `unlike` `subscribe` `unsubscribe` | 浏览器 |
| **boss** | `search` `detail` `recommend` `joblist` `greet` `batchgreet` `send` `chatlist` `chatmsg` `invite` `mark` `exchange` `resume` `stats` | 浏览器 |
| **coupang** | `search` `add-to-cart` | 浏览器 |
| **bbc** | `news` | 公共 API |
@@ -216,7 +233,7 @@ npm link
| **sinafinance** | `news` | 🌐 公开 |
| **barchart** | `quote` `options` `greeks` `flow` | 浏览器 |
| **chaoxing** | `assignments` `exams` | 浏览器 |
| **grok** | `ask` | 浏览器 |
| **grok** | `ask` `image` | 浏览器 |
| **hf** | `top` | 公开 |
| **jike** | `feed` `search` `create` `like` `comment` `repost` `notifications` `post` `topic` `user` | 浏览器 |
| **jimeng** | `generate` `history` | 浏览器 |
@@ -249,7 +266,9 @@ npm link
| **douyin** | `videos` `publish` `drafts` `draft` `delete` `stats` `profile` `update` `hashtag` `location` `activities` `collections` | 浏览器 |
| **yuanbao** | `new` `ask` | 浏览器 |
79+ 适配器 — **[→ 查看完整命令列表](./docs/adapters/index.md)**
87+ 适配器 — **[→ 查看完整命令列表](./docs/adapters/index.md)**
`*` `opencli xiaoyuzhou transcript` 需要本地小宇宙凭证:`~/.opencli/xiaoyuzhou.json`
### 外部 CLI 枢纽
@@ -284,7 +303,7 @@ opencli register mycli
| **Cursor** | 控制 Cursor IDE — Composer、对话、代码提取等 | [Doc](./docs/adapters/desktop/cursor.md) |
| **Codex** | 在后台(无头)驱动 OpenAI Codex CLI Agent | [Doc](./docs/adapters/desktop/codex.md) |
| **Antigravity** | 在终端直接控制 Antigravity Ultra | [Doc](./docs/adapters/desktop/antigravity.md) |
| **ChatGPT** | 自动化操作 ChatGPT macOS 桌面客户端 | [Doc](./docs/adapters/desktop/chatgpt.md) |
| **ChatGPT App** | 自动化操作 ChatGPT macOS 桌面客户端 | [Doc](./docs/adapters/desktop/chatgpt-app.md) |
| **ChatWise** | 多 LLM 客户端(GPT-4、Claude、Gemini | [Doc](./docs/adapters/desktop/chatwise.md) |
| **Notion** | 搜索、读取、写入 Notion 页面 | [Doc](./docs/adapters/desktop/notion.md) |
| **Discord** | Discord 桌面版 — 消息、频道、服务器 | [Doc](./docs/adapters/desktop/discord.md) |
@@ -303,6 +322,7 @@ OpenCLI 支持从各平台下载图片、视频和文章。
| **Twitter/X** | 图片、视频 | 从用户媒体页或单条推文下载 |
| **Pixiv** | 图片 | 下载原始画质插画,支持多页作品 |
| **1688** | 图片、视频 | 下载商品页中可见的商品素材 |
| **小宇宙** | 音频、转录 | 从公开单集数据下载音频,并使用本地凭证下载转录 JSON / 文本 |
| **知乎** | 文章(Markdown) | 导出文章,可选下载图片到本地 |
| **微信公众号** | 文章(Markdown | 导出微信公众号文章为 Markdown |
| **豆瓣** | 图片 | 下载电影条目的海报 / 剧照图片 |
@@ -322,7 +342,8 @@ brew install yt-dlp
```bash
# 下载小红书笔记中的图片/视频
opencli xiaohongshu download abc123 --output ./xhs
opencli xiaohongshu download "https://www.xiaohongshu.com/search_result/<id>?xsec_token=..." --output ./xhs
opencli xiaohongshu download "https://xhslink.com/..." --output ./xhs
# 下载B站视频(需要 yt-dlp
opencli bilibili download BV1xxx --output ./bilibili
@@ -340,6 +361,12 @@ opencli douban download 30382501 --output ./douban
# 下载 1688 商品页中的图片 / 视频素材
opencli 1688 download 841141931191 --output ./1688-downloads
# 下载小宇宙单集音频
opencli xiaoyuzhou download 69b3b675772ac2295bfc01d0 --output ./xiaoyuzhou
# 下载小宇宙单集转录
opencli xiaoyuzhou transcript 69dd0c98e2c8be31551f6a33 --output ./xiaoyuzhou-transcripts
# 导出知乎文章为 Markdown
opencli zhihu download "https://zhuanlan.zhihu.com/p/xxx" --output ./zhihu
@@ -350,6 +377,8 @@ opencli zhihu download "https://zhuanlan.zhihu.com/p/xxx" --download-images
opencli weixin download --url "https://mp.weixin.qq.com/s/xxx" --output ./weixin
```
`opencli xiaoyuzhou transcript` 需要本地小宇宙凭证:`~/.opencli/xiaoyuzhou.json`
## 输出格式
@@ -394,7 +423,7 @@ esac
## 插件
通过社区贡献的插件扩展 OpenCLI。插件使用与内置命令相同的 YAML/TS 格式,启动时自动发现。
通过社区贡献的插件扩展 OpenCLI。插件使用与内置命令相同的 JS 格式,启动时自动发现。
```bash
opencli plugin install github:user/opencli-plugin-my-tool # 安装
@@ -408,9 +437,10 @@ opencli plugin uninstall my-tool # 卸载
| 插件 | 类型 | 描述 |
|------|------|------|
| [opencli-plugin-github-trending](https://github.com/ByteYue/opencli-plugin-github-trending) | YAML | GitHub Trending 仓库 |
| [opencli-plugin-hot-digest](https://github.com/ByteYue/opencli-plugin-hot-digest) | TS | 多平台热榜聚合 |
| [opencli-plugin-juejin](https://github.com/Astro-Han/opencli-plugin-juejin) | YAML | 稀土掘金热门文章 |
| [opencli-plugin-github-trending](https://github.com/ByteYue/opencli-plugin-github-trending) | JS | GitHub Trending 仓库 |
| [opencli-plugin-hot-digest](https://github.com/ByteYue/opencli-plugin-hot-digest) | JS | 多平台热榜聚合 |
| [opencli-plugin-juejin](https://github.com/Astro-Han/opencli-plugin-juejin) | JS | 稀土掘金热门文章 |
| [opencli-plugin-vk](https://github.com/flobo3/opencli-plugin-vk) | JS | VK (VKontakte) 动态、信息流和搜索 |
详见 [插件指南](./docs/zh/guide/plugins.md) 了解如何创建自己的插件。
@@ -447,7 +477,7 @@ opencli cascade https://api.example.com/data
- **返回空数据,或者报错 "Unauthorized"**
- Chrome/Chromium 里的登录态可能已经过期。请打开当前页面,在新标签页重新手工登录或刷新该页面。
- **Node API 错误 (如 parseArgs, fs 等)**
- 确保 Node.js 版本 `>= 20`
- 确保 Node.js 版本 `>= 21``node:util``styleText` 需要 Node 21+
- **Daemon 问题**
- 检查 daemon 状态:`curl localhost:19825/status`
- 查看扩展日志:`curl localhost:19825/logs`
+18 -17
View File
@@ -30,12 +30,15 @@ tests/
├── smoke/
│ └── api-health.test.ts # 外部 API、adapter 定义、命令注册健康检查
src/
── **/*.test.ts # 单元测试(当前 32 个文件
── **/*.test.ts # 单元测试(unit project
clis/
└── **/*.test.{ts,js} # adapter 测试(adapter project
```
| 层 | 位置 | 当前文件数 | 运行方式 | 用途 |
|---|---|---:|---|---|
| 单元测试 | `src/**/*.test.ts` | 32 | `npx vitest run src/` | 内部模块、pipeline、adapter 工具函数 |
| 单元测试 | `src/**/*.test.ts` | 32 | `npm test` | 内部模块、pipeline、runtime |
| Adapter 测试 | `clis/**/*.test.{ts,js}` | - | `npm test` / `npm run test:adapter` | adapter 命令与数据归一化 |
| E2E 测试 | `tests/e2e/*.test.ts` | 5 | `npx vitest run tests/e2e/` | 真实 CLI 命令执行 |
| 烟雾测试 | `tests/smoke/*.test.ts` | 1 | `npx vitest run tests/smoke/` | 外部 API 与注册完整性 |
@@ -43,7 +46,7 @@ src/
## 当前覆盖范围
### 单元测试32 个文件)
### 单元测试与 Adapter 测试
| 领域 | 文件 |
|---|---|
@@ -100,8 +103,11 @@ npm run build # 编译(E2E / smoke 测试需要 dist/src/main.js
### 运行命令
```bash
# 全部单元测试
npx vitest run src/
# 默认本地测试口径(unit + extension + adapter
npm test
# 只跑 adapter project
npm run test:adapter
# 全部 E2E 测试(会真实调用外部 API / 浏览器)
npx vitest run tests/e2e/
@@ -110,7 +116,7 @@ npx vitest run tests/e2e/
npx vitest run tests/smoke/
# 单个测试文件
npx vitest run clis/apple-podcasts/commands.test.ts
npm test -- --run clis/apple-podcasts/commands.test.ts
npx vitest run tests/e2e/management.test.ts
# 全部测试
@@ -192,7 +198,8 @@ it('producthunt me fails gracefully without login', async () => {
| Job | 触发条件 | 内容 |
|---|---|---|
| `build` | push/PR 到 `main`,`dev` | `tsc --noEmit` + `npm run build` |
| `unit-test` | push/PR 到 `main`,`dev` | Node `20``22` 双版本运行 `src/` 单元测试,按 `2` shard 并行 |
| `unit-test` | push/PR 到 `main`,`dev` | Node `22` 运行 `unit + extension`,按 `2` shard 并行 |
| `adapter-test` | push/PR 到 `main`,`dev` | Node `22` 单独运行 `adapter` project |
| `smoke-test` | `schedule``workflow_dispatch` | 安装真实 Chrome`xvfb-run` 执行 `tests/smoke/` |
### `e2e-headed.yml`
@@ -201,19 +208,18 @@ it('producthunt me fails gracefully without login', async () => {
|---|---|---|
| `e2e-headed` | push/PR 到 `main`,`dev`,或手动触发 | 安装真实 Chrome,`xvfb-run` 执行 `tests/e2e/` |
E2E 与 smoke 都使用 `./.github/actions/setup-chrome` 准备真实 Chrome,并通过 `OPENCLI_BROWSER_EXECUTABLE_PATH` 注入浏览器路径
E2E 与 smoke 都使用 `./.github/actions/setup-chrome` 准备真实 Chrome。
### Sharding
单元测试使用 vitest 内置 shard并在 Node `20` / `22` 两个版本上运行
CI 里的 `unit-test` job 使用 vitest shard只切 `unit + extension`,避免和独立的 `adapter-test` job 重复
```yaml
strategy:
matrix:
node-version: ['20', '22']
shard: [1, 2]
steps:
- run: npx vitest run src/ --reporter=verbose --shard=${{ matrix.shard }}/2
- run: npx vitest run --project unit --project extension --reporter=verbose --shard=${{ matrix.shard }}/2
```
---
@@ -227,12 +233,7 @@ opencli 通过 Browser Bridge 扩展连接浏览器:
| 扩展已安装 / 已连接 | Extension 模式 | 本地用户,连接已登录的 Chrome |
| 无扩展 token | CLI 自行拉起浏览器 | CI、无登录态或纯自动化场景 |
CI 中使用 `OPENCLI_BROWSER_EXECUTABLE_PATH` 指定真实 Chrome 路径:
```yaml
env:
OPENCLI_BROWSER_EXECUTABLE_PATH: ${{ steps.setup-chrome.outputs.chrome-path }}
```
CI 通过 `./.github/actions/setup-chrome` 准备真实 Chrome,再直接执行测试。
---
+19066
View File
File diff suppressed because it is too large Load Diff
+68 -121
View File
@@ -1,52 +1,7 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
assertAuthenticatedState,
buildDetailUrl,
buildProvenance,
cleanText,
extractOfferId,
gotoAndReadState,
type MediaSource,
uniqueMediaSources,
} from './shared.js';
interface AssetBrowserPayload {
href?: string;
title?: string;
offerTitle?: string;
offerId?: string | number;
gallery?: {
mainImage?: string[];
offerImgList?: string[];
wlImageInfos?: Array<{ fullPathImageURI?: string }>;
[key: string]: unknown;
};
scannedAssets?: MediaSource[];
}
export interface Normalized1688Assets {
offer_id: string | null;
title: string | null;
item_url: string;
main_images: string[];
sku_images: string[];
detail_images: string[];
videos: string[];
other_images: string[];
raw_assets: MediaSource[];
source: string[];
main_count: number;
sku_count: number;
detail_count: number;
video_count: number;
source_url: string;
fetched_at: string;
strategy: string;
}
function scriptToReadAssets(): string {
return `
import { assertAuthenticatedState, buildDetailUrl, buildProvenance, cleanText, extractOfferId, gotoAndReadState, uniqueMediaSources, } from './shared.js';
function scriptToReadAssets() {
return `
(() => {
const root = window.context ?? {};
const model = root.result?.global?.globalData?.model ?? null;
@@ -174,84 +129,76 @@ function scriptToReadAssets(): string {
})()
`;
}
function normalizeAssets(payload: AssetBrowserPayload): Normalized1688Assets {
const offerId = cleanText(String(payload.offerId ?? '')) || extractOfferId(cleanText(payload.href)) || null;
const itemUrl = offerId ? buildDetailUrl(offerId) : cleanText(payload.href);
const seededAssets: MediaSource[] = [
...((payload.gallery?.mainImage ?? []).map((url) => ({ type: 'image' as const, group: 'main' as const, url, source: 'page_state:mainImage' }))),
...((payload.gallery?.offerImgList ?? []).map((url) => ({ type: 'image' as const, group: 'main' as const, url, source: 'page_state:offerImgList' }))),
...((payload.gallery?.wlImageInfos ?? []).map((item) => ({
type: 'image' as const,
group: 'main' as const,
url: item?.fullPathImageURI ?? '',
source: 'page_state:wlImageInfos',
}))),
];
const assets = uniqueMediaSources([...seededAssets, ...(payload.scannedAssets ?? [])]);
const mainImages = assets.filter((item) => item.type === 'image' && item.group === 'main').map((item) => item.url);
const skuImages = assets.filter((item) => item.type === 'image' && item.group === 'sku').map((item) => item.url);
const detailImages = assets.filter((item) => item.type === 'image' && item.group === 'detail').map((item) => item.url);
const videos = assets.filter((item) => item.type === 'video').map((item) => item.url);
const otherImages = assets
.filter((item) => item.type === 'image' && !['main', 'sku', 'detail'].includes(item.group))
.map((item) => item.url);
return {
offer_id: offerId,
title: cleanText(payload.offerTitle) || cleanText(payload.title) || null,
item_url: itemUrl,
main_images: mainImages,
sku_images: skuImages,
detail_images: detailImages,
videos,
other_images: otherImages,
raw_assets: assets,
source: [...new Set(assets.map((item) => cleanText(item.source)).filter(Boolean))],
main_count: mainImages.length,
sku_count: skuImages.length,
detail_count: detailImages.length,
video_count: videos.length,
...buildProvenance(cleanText(payload.href) || itemUrl),
};
function normalizeAssets(payload) {
const offerId = cleanText(String(payload.offerId ?? '')) || extractOfferId(cleanText(payload.href)) || null;
const itemUrl = offerId ? buildDetailUrl(offerId) : cleanText(payload.href);
const seededAssets = [
...((payload.gallery?.mainImage ?? []).map((url) => ({ type: 'image', group: 'main', url, source: 'page_state:mainImage' }))),
...((payload.gallery?.offerImgList ?? []).map((url) => ({ type: 'image', group: 'main', url, source: 'page_state:offerImgList' }))),
...((payload.gallery?.wlImageInfos ?? []).map((item) => ({
type: 'image',
group: 'main',
url: item?.fullPathImageURI ?? '',
source: 'page_state:wlImageInfos',
}))),
];
const assets = uniqueMediaSources([...seededAssets, ...(payload.scannedAssets ?? [])]);
const mainImages = assets.filter((item) => item.type === 'image' && item.group === 'main').map((item) => item.url);
const skuImages = assets.filter((item) => item.type === 'image' && item.group === 'sku').map((item) => item.url);
const detailImages = assets.filter((item) => item.type === 'image' && item.group === 'detail').map((item) => item.url);
const videos = assets.filter((item) => item.type === 'video').map((item) => item.url);
const otherImages = assets
.filter((item) => item.type === 'image' && !['main', 'sku', 'detail'].includes(item.group))
.map((item) => item.url);
return {
offer_id: offerId,
title: cleanText(payload.offerTitle) || cleanText(payload.title) || null,
item_url: itemUrl,
main_images: mainImages,
sku_images: skuImages,
detail_images: detailImages,
videos,
other_images: otherImages,
raw_assets: assets,
source: [...new Set(assets.map((item) => cleanText(item.source)).filter(Boolean))],
main_count: mainImages.length,
sku_count: skuImages.length,
detail_count: detailImages.length,
video_count: videos.length,
...buildProvenance(cleanText(payload.href) || itemUrl),
};
}
async function readAssetsPayload(page: IPage, itemUrl: string): Promise<AssetBrowserPayload> {
const state = await gotoAndReadState(page, itemUrl, 2500, 'assets');
assertAuthenticatedState(state, 'assets');
await page.autoScroll({ times: 3, delayMs: 400 });
await page.wait(1);
return await page.evaluate(scriptToReadAssets()) as AssetBrowserPayload;
async function readAssetsPayload(page, itemUrl) {
const state = await gotoAndReadState(page, itemUrl, 2500, 'assets');
assertAuthenticatedState(state, 'assets');
await page.autoScroll({ times: 3, delayMs: 400 });
await page.wait(1);
return await page.evaluate(scriptToReadAssets());
}
export async function extractAssetsForInput(page: IPage, input: string): Promise<Normalized1688Assets> {
const itemUrl = buildDetailUrl(String(input ?? ''));
const payload = await readAssetsPayload(page, itemUrl);
return normalizeAssets(payload);
export async function extractAssetsForInput(page, input) {
const itemUrl = buildDetailUrl(String(input ?? ''));
const payload = await readAssetsPayload(page, itemUrl);
return normalizeAssets(payload);
}
cli({
site: '1688',
name: 'assets',
description: '列出 1688 商品页可提取的图片/视频素材',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 商品 URL 或 offer ID(如 887904326744',
site: '1688',
name: 'assets',
description: '列出 1688 商品页可提取的图片/视频素材',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 商品 URL 或 offer ID(如 887904326744',
},
],
columns: ['offer_id', 'title', 'main_count', 'sku_count', 'detail_count', 'video_count'],
func: async (page, kwargs) => {
return [await extractAssetsForInput(page, String(kwargs.input ?? ''))];
},
],
columns: ['offer_id', 'title', 'main_count', 'sku_count', 'detail_count', 'video_count'],
func: async (page, kwargs) => {
return [await extractAssetsForInput(page, String(kwargs.input ?? ''))];
},
});
export const __test__ = {
normalizeAssets,
normalizeAssets,
};
+39
View File
@@ -0,0 +1,39 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './assets.js';
import { __test__ as sharedTest } from './shared.js';
describe('1688 assets normalization', () => {
it('normalizes gallery and scanned assets into grouped media lists', () => {
const result = __test__.normalizeAssets({
href: 'https://detail.1688.com/offer/887904326744.html',
title: '测试商品 - 阿里巴巴',
offerTitle: '测试商品',
offerId: 887904326744,
gallery: {
mainImage: ['//img.example.com/main-1.jpg'],
offerImgList: ['https://img.example.com/main-2.jpg'],
wlImageInfos: [{ fullPathImageURI: 'https://img.example.com/main-3.jpg' }],
},
scannedAssets: [
{ type: 'image', group: 'sku', url: 'https://img.example.com/sku-1.png', source: 'dom:.sku' },
{ type: 'image', group: 'detail', url: 'https://img.example.com/detail-1.jpg', source: 'dom:.detail' },
{ type: 'video', group: 'video', url: 'https://video.example.com/demo.mp4', source: 'script' },
{ type: 'image', group: 'detail', url: 'blob:https://detail.1688.com/1', source: 'ignore' },
],
});
expect(result.offer_id).toBe('887904326744');
expect(result.main_images).toEqual([
'https://img.example.com/main-1.jpg',
'https://img.example.com/main-2.jpg',
'https://img.example.com/main-3.jpg',
]);
expect(result.sku_images).toEqual(['https://img.example.com/sku-1.png']);
expect(result.detail_images).toEqual(['https://img.example.com/detail-1.jpg']);
expect(result.videos).toEqual(['https://video.example.com/demo.mp4']);
expect(result.main_count).toBe(3);
expect(result.video_count).toBe(1);
});
it('normalizes media urls from style syntax and protocol-relative URLs', () => {
expect(sharedTest.normalizeMediaUrl('url("//img.example.com/1.jpg")')).toBe('https://img.example.com/1.jpg');
expect(sharedTest.normalizeMediaUrl('blob:https://detail.1688.com/1')).toBe('');
});
});
-42
View File
@@ -1,42 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './assets.js';
import { __test__ as sharedTest } from './shared.js';
describe('1688 assets normalization', () => {
it('normalizes gallery and scanned assets into grouped media lists', () => {
const result = __test__.normalizeAssets({
href: 'https://detail.1688.com/offer/887904326744.html',
title: '测试商品 - 阿里巴巴',
offerTitle: '测试商品',
offerId: 887904326744,
gallery: {
mainImage: ['//img.example.com/main-1.jpg'],
offerImgList: ['https://img.example.com/main-2.jpg'],
wlImageInfos: [{ fullPathImageURI: 'https://img.example.com/main-3.jpg' }],
},
scannedAssets: [
{ type: 'image', group: 'sku', url: 'https://img.example.com/sku-1.png', source: 'dom:.sku' },
{ type: 'image', group: 'detail', url: 'https://img.example.com/detail-1.jpg', source: 'dom:.detail' },
{ type: 'video', group: 'video', url: 'https://video.example.com/demo.mp4', source: 'script' },
{ type: 'image', group: 'detail', url: 'blob:https://detail.1688.com/1', source: 'ignore' },
],
});
expect(result.offer_id).toBe('887904326744');
expect(result.main_images).toEqual([
'https://img.example.com/main-1.jpg',
'https://img.example.com/main-2.jpg',
'https://img.example.com/main-3.jpg',
]);
expect(result.sku_images).toEqual(['https://img.example.com/sku-1.png']);
expect(result.detail_images).toEqual(['https://img.example.com/detail-1.jpg']);
expect(result.videos).toEqual(['https://video.example.com/demo.mp4']);
expect(result.main_count).toBe(3);
expect(result.video_count).toBe(1);
});
it('normalizes media urls from style syntax and protocol-relative URLs', () => {
expect(sharedTest.normalizeMediaUrl('url("//img.example.com/1.jpg")')).toBe('https://img.example.com/1.jpg');
expect(sharedTest.normalizeMediaUrl('blob:https://detail.1688.com/1')).toBe('');
});
});
+76
View File
@@ -0,0 +1,76 @@
import * as path from 'node:path';
import { formatCookieHeader } from '@jackwener/opencli/download';
import { downloadMedia } from '@jackwener/opencli/download/media-download';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { cleanText } from './shared.js';
import { extractAssetsForInput } from './assets.js';
function extFromUrl(url, fallback) {
try {
const ext = path.extname(new URL(url).pathname).toLowerCase();
if (ext && ext.length <= 8)
return ext;
}
catch {
// ignore
}
return fallback;
}
function toDownloadItems(offerId, assets) {
const items = [];
const pushImages = (urls, prefix) => {
urls.forEach((url, index) => {
items.push({
type: 'image',
url,
filename: `${offerId}_${prefix}_${String(index + 1).padStart(2, '0')}${extFromUrl(url, '.jpg')}`,
});
});
};
pushImages(assets.main_images, 'main');
pushImages(assets.sku_images, 'sku');
pushImages(assets.detail_images, 'detail');
pushImages(assets.other_images, 'other');
assets.videos.forEach((url, index) => {
items.push({
type: 'video',
url,
filename: `${offerId}_video_${String(index + 1).padStart(2, '0')}${extFromUrl(url, '.mp4')}`,
});
});
return items;
}
cli({
site: '1688',
name: 'download',
description: '批量下载 1688 商品页可提取的图片和视频素材',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 商品 URL 或 offer ID(如 887904326744',
},
{ name: 'output', default: './1688-downloads', help: '输出目录' },
],
columns: ['index', 'type', 'status', 'size'],
func: async (page, kwargs) => {
const assets = await extractAssetsForInput(page, String(kwargs.input ?? ''));
const offerId = cleanText(assets.offer_id) || '1688';
const items = toDownloadItems(offerId, assets);
const browserCookies = await page.getCookies({ domain: '1688.com' });
return downloadMedia(items, {
output: String(kwargs.output || './1688-downloads'),
subdir: offerId,
cookies: formatCookieHeader(browserCookies),
browserCookies,
filenamePrefix: offerId,
timeout: 60000,
});
},
});
export const __test__ = {
extFromUrl,
toDownloadItems,
};
+31
View File
@@ -0,0 +1,31 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './download.js';
describe('1688 download helpers', () => {
it('builds stable filenames for grouped assets', () => {
const items = __test__.toDownloadItems('887904326744', {
offer_id: '887904326744',
title: '测试商品',
item_url: 'https://detail.1688.com/offer/887904326744.html',
main_images: ['https://img.example.com/a.jpg'],
sku_images: ['https://img.example.com/b.png'],
detail_images: ['https://img.example.com/c.webp'],
videos: ['https://video.example.com/d.mp4'],
other_images: [],
raw_assets: [],
source: [],
main_count: 1,
sku_count: 1,
detail_count: 1,
video_count: 1,
source_url: 'https://detail.1688.com/offer/887904326744.html',
fetched_at: new Date().toISOString(),
strategy: 'cookie',
});
expect(items.map((item) => item.filename)).toEqual([
'887904326744_main_01.jpg',
'887904326744_sku_01.png',
'887904326744_detail_01.webp',
'887904326744_video_01.mp4',
]);
});
});
-33
View File
@@ -1,33 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './download.js';
describe('1688 download helpers', () => {
it('builds stable filenames for grouped assets', () => {
const items = __test__.toDownloadItems('887904326744', {
offer_id: '887904326744',
title: '测试商品',
item_url: 'https://detail.1688.com/offer/887904326744.html',
main_images: ['https://img.example.com/a.jpg'],
sku_images: ['https://img.example.com/b.png'],
detail_images: ['https://img.example.com/c.webp'],
videos: ['https://video.example.com/d.mp4'],
other_images: [],
raw_assets: [],
source: [],
main_count: 1,
sku_count: 1,
detail_count: 1,
video_count: 1,
source_url: 'https://detail.1688.com/offer/887904326744.html',
fetched_at: new Date().toISOString(),
strategy: 'cookie',
});
expect(items.map((item) => item.filename)).toEqual([
'887904326744_main_01.jpg',
'887904326744_sku_01.png',
'887904326744_detail_01.webp',
'887904326744_video_01.mp4',
]);
});
});
-83
View File
@@ -1,83 +0,0 @@
import * as path from 'node:path';
import { formatCookieHeader } from '@jackwener/opencli/download';
import { downloadMedia, type MediaItem } from '@jackwener/opencli/download/media-download';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { cleanText } from './shared.js';
import { extractAssetsForInput } from './assets.js';
function extFromUrl(url: string, fallback: string): string {
try {
const ext = path.extname(new URL(url).pathname).toLowerCase();
if (ext && ext.length <= 8) return ext;
} catch {
// ignore
}
return fallback;
}
function toDownloadItems(offerId: string, assets: Awaited<ReturnType<typeof extractAssetsForInput>>): MediaItem[] {
const items: MediaItem[] = [];
const pushImages = (urls: string[], prefix: string) => {
urls.forEach((url, index) => {
items.push({
type: 'image',
url,
filename: `${offerId}_${prefix}_${String(index + 1).padStart(2, '0')}${extFromUrl(url, '.jpg')}`,
});
});
};
pushImages(assets.main_images, 'main');
pushImages(assets.sku_images, 'sku');
pushImages(assets.detail_images, 'detail');
pushImages(assets.other_images, 'other');
assets.videos.forEach((url, index) => {
items.push({
type: 'video',
url,
filename: `${offerId}_video_${String(index + 1).padStart(2, '0')}${extFromUrl(url, '.mp4')}`,
});
});
return items;
}
cli({
site: '1688',
name: 'download',
description: '批量下载 1688 商品页可提取的图片和视频素材',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 商品 URL 或 offer ID(如 887904326744',
},
{ name: 'output', default: './1688-downloads', help: '输出目录' },
],
columns: ['index', 'type', 'status', 'size'],
func: async (page, kwargs) => {
const assets = await extractAssetsForInput(page, String(kwargs.input ?? ''));
const offerId = cleanText(assets.offer_id) || '1688';
const items = toDownloadItems(offerId, assets);
const browserCookies = await page.getCookies({ domain: '1688.com' });
return downloadMedia(items, {
output: String(kwargs.output || './1688-downloads'),
subdir: offerId,
cookies: formatCookieHeader(browserCookies),
browserCookies,
filenamePrefix: offerId,
timeout: 60000,
});
},
});
export const __test__ = {
extFromUrl,
toDownloadItems,
};
+187
View File
@@ -0,0 +1,187 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { isRecord } from '@jackwener/opencli/utils';
import { assertAuthenticatedState, buildDetailUrl, buildProvenance, canonicalizeSellerUrl, cleanMultilineText, cleanText, extractLocation, extractMemberId, extractOfferId, extractShopId, gotoAndReadState, normalizePriceTiers, parseMoqText, parsePriceText, toNumber, uniqueNonEmpty, } from './shared.js';
function normalizeItemPayload(payload) {
const href = cleanText(payload.href);
const bodyText = cleanMultilineText(payload.bodyText);
const sellerName = cleanText(payload.seller?.companyName);
const sellerUrlRaw = cleanText(payload.seller?.winportUrl
?? payload.seller?.sellerWinportUrlMap?.defaultUrl
?? payload.seller?.sellerWinportUrlMap?.indexUrl);
const sellerUrl = canonicalizeSellerUrl(sellerUrlRaw);
const offerId = cleanText(String(payload.offerId ?? '')) || extractOfferId(href) || null;
const memberId = cleanText(payload.seller?.memberId) || extractMemberId(sellerUrlRaw || href) || null;
const shopId = extractShopId(sellerUrl ?? href);
const unit = cleanText(payload.trade?.unit);
const priceDisplay = cleanText(payload.trade?.priceDisplay);
const priceRange = parsePriceText(priceDisplay ? `¥${priceDisplay}` : bodyText);
const moqText = extractMoqText(bodyText, payload.trade?.beginAmount, unit);
const moq = parseMoqText(moqText);
const services = uniqueServices(payload);
const serviceBadges = uniqueNonEmpty(services.map((service) => cleanText(service.serviceName)));
const attributes = normalizeVisibleAttributes(payload.trade?.offerIDatacenterSellInfo);
const priceTiers = normalizePriceTiers(payload.trade?.offerPriceModel?.currentPrices ?? [], unit || null);
const images = uniqueNonEmpty([
...(payload.gallery?.mainImage ?? []),
...(payload.gallery?.offerImgList ?? []),
...((payload.gallery?.wlImageInfos ?? []).map((item) => item.fullPathImageURI ?? '')),
]);
const detailUrl = offerId ? buildDetailUrl(offerId) : href;
const provenance = buildProvenance(href || detailUrl);
return {
offer_id: offerId,
member_id: memberId,
shop_id: shopId,
title: cleanText(payload.offerTitle) || stripAlibabaSuffix(payload.title) || firstNonEmptyLine(bodyText) || null,
item_url: detailUrl,
main_images: images,
price_text: priceRange.price_text || null,
price_tiers: priceTiers,
currency: priceRange.currency,
moq_text: moq.moq_text || null,
moq_value: moq.moq_value,
seller_name: sellerName || null,
seller_url: sellerUrl,
shop_name: sellerName || null,
origin_place: extractLocation(bodyText),
delivery_days_text: extractDeliveryDaysText(bodyText, services, payload.shipping),
customization_text: extractKeywordLine(bodyText, ['来样定制', '来图定制', '支持定制', '可定制', '定制']),
private_label_text: extractKeywordLine(bodyText, ['贴牌', '贴标', '定制logo', '打logo', 'OEM', 'ODM']),
visible_attributes: attributes,
sales_text: extractSalesText(bodyText),
service_badges: serviceBadges,
stock_quantity: extractStockQuantity(bodyText),
...provenance,
};
}
function normalizeVisibleAttributes(raw) {
if (!isRecord(raw))
return [];
return Object.entries(raw)
.filter(([key, value]) => key !== 'sellPointModel' && cleanText(key) && cleanText(String(value)))
.map(([key, value]) => ({ key: cleanText(key), value: cleanText(String(value)) }));
}
function uniqueServices(payload) {
const combined = [
...(Array.isArray(payload.services) ? payload.services : []),
...(Array.isArray(payload.shipping?.protectionInfos) ? payload.shipping.protectionInfos : []),
...(Array.isArray(payload.shipping?.buyerProtectionModel) ? payload.shipping.buyerProtectionModel : []),
];
const seen = new Set();
const result = [];
for (const service of combined) {
const key = cleanText(service.serviceName);
if (!key || seen.has(key))
continue;
seen.add(key);
result.push(service);
}
return result;
}
function stripAlibabaSuffix(title) {
return cleanText(title).replace(/\s*-\s*阿里巴巴$/, '').trim();
}
function firstNonEmptyLine(text) {
return text.split('\n').map((line) => cleanText(line)).find(Boolean) ?? '';
}
function extractMoqText(bodyText, beginAmount, unit) {
const lineMatch = bodyText.match(/\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)\s*起批/);
if (lineMatch)
return lineMatch[0];
const moqValue = toNumber(beginAmount);
if (moqValue !== null) {
return `${moqValue}${unit || ''}起批`;
}
return '';
}
function extractDeliveryDaysText(bodyText, services, shipping) {
const shippingText = cleanText(shipping?.deliveryLimitText) || cleanText(shipping?.logisticsText);
if (shippingText)
return shippingText;
const textMatch = bodyText.match(/\d+\s*(?:小时|天)(?:内)?发货/);
if (textMatch)
return textMatch[0];
const hourMatch = services.find((service) => typeof service.agreeDeliveryHours === 'number');
if (hourMatch && typeof hourMatch.agreeDeliveryHours === 'number') {
return `${hourMatch.agreeDeliveryHours}小时内发货`;
}
return null;
}
function extractKeywordLine(bodyText, keywords) {
const lines = bodyText.split('\n').map((line) => cleanText(line)).filter(Boolean);
for (const line of lines) {
if (keywords.some((keyword) => line.includes(keyword))) {
return line;
}
}
return null;
}
function extractSalesText(bodyText) {
const match = bodyText.match(/(?:全网销量|已售)\s*\d+(?:\.\d+)?\+?\s*[件套个单]?/);
return match ? cleanText(match[0]) : null;
}
function extractStockQuantity(bodyText) {
const match = bodyText.match(/库存\s*(\d+)/);
return match ? Number.parseInt(match[1], 10) : null;
}
async function readItemPayload(page, itemUrl) {
const state = await gotoAndReadState(page, itemUrl, 2500, 'item');
assertAuthenticatedState(state, 'item');
const payload = await page.evaluate(`
(() => {
const root = window.context ?? {};
const model = root.result?.global?.globalData?.model ?? null;
const toJson = (value) => JSON.parse(JSON.stringify(value ?? null));
return {
href: window.location.href,
title: document.title || '',
bodyText: document.body ? document.body.innerText || '' : '',
offerTitle: model?.offerTitleModel?.subject ?? '',
offerId: model?.tradeModel?.offerId ?? '',
seller: toJson(model?.sellerModel),
trade: toJson(model?.tradeModel),
gallery: toJson(root.result?.data?.gallery?.fields ?? null),
shipping: toJson(root.result?.data?.shippingServices?.fields ?? null),
services: toJson(root.result?.data?.shippingServices?.fields?.protectionInfos ?? []),
};
})()
`);
const resolvedOfferId = cleanText(String(payload.offerId ?? '')) || extractOfferId(cleanText(payload.href));
if (!resolvedOfferId) {
throw new CommandExecutionError('1688 item page did not expose product context', '当前 tab 非商品详情上下文,请切到 detail.1688.com 商品页并重试');
}
return payload;
}
cli({
site: '1688',
name: 'item',
description: '1688 商品详情(公开商品字段、价格阶梯、卖家基础信息)',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 商品 URL 或 offer ID(如 887904326744',
},
],
columns: ['offer_id', 'title', 'price_text', 'moq_text', 'seller_name', 'origin_place'],
func: async (page, kwargs) => {
const itemUrl = buildDetailUrl(String(kwargs.input ?? ''));
const payload = await readItemPayload(page, itemUrl);
return [normalizeItemPayload(payload)];
},
});
export const __test__ = {
normalizeItemPayload,
normalizeVisibleAttributes,
stripAlibabaSuffix,
extractMoqText,
extractDeliveryDaysText,
extractKeywordLine,
extractSalesText,
extractStockQuantity,
};
+67
View File
@@ -0,0 +1,67 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './item.js';
describe('1688 item normalization', () => {
it('normalizes public item payload into contract fields', () => {
const result = __test__.normalizeItemPayload({
href: 'https://detail.1688.com/offer/887904326744.html',
title: '法式春季长袖开衫连衣裙女新款大码女装碎花吊带裙套装142077 - 阿里巴巴',
bodyText: `
青岛沁澜衣品服装有限公司
入驻13年
主营:大码女装
店铺回头率
87%
山东青岛
3套起批
已售1600+套
支持定制logo
`,
offerTitle: '法式春季长袖开衫连衣裙女新款大码女装碎花吊带裙套装142077',
offerId: 887904326744,
seller: {
companyName: '青岛沁澜衣品服装有限公司',
memberId: 'b2b-1641351767',
winportUrl: 'https://yinuoweierfushi.1688.com/page/index.html?spm=a1',
},
trade: {
beginAmount: 3,
priceDisplay: '96.00-98.00',
unit: '套',
saleCount: 1655,
offerIDatacenterSellInfo: {
面料名称: '莫代尔',
主面料成分: '莫代尔纤维',
sellPointModel: '{"ignore":true}',
},
offerPriceModel: {
currentPrices: [
{ beginAmount: 3, price: '98.00' },
{ beginAmount: 50, price: '97.00' },
],
},
},
gallery: {
mainImage: ['https://example.com/1.jpg'],
offerImgList: ['https://example.com/2.jpg'],
wlImageInfos: [{ fullPathImageURI: 'https://example.com/3.jpg' }],
},
services: [
{ serviceName: '延期必赔', agreeDeliveryHours: 360 },
{ serviceName: '品质保障' },
],
});
expect(result.offer_id).toBe('887904326744');
expect(result.member_id).toBe('b2b-1641351767');
expect(result.shop_id).toBe('yinuoweierfushi');
expect(result.seller_url).toBe('https://yinuoweierfushi.1688.com');
expect(result.price_text).toBe('¥96.00-98.00');
expect(result.moq_text).toBe('3套起批');
expect(result.origin_place).toBe('山东青岛');
expect(result.delivery_days_text).toBe('360小时内发货');
expect(result.private_label_text).toBe('支持定制logo');
expect(result.visible_attributes).toEqual([
{ key: '面料名称', value: '莫代尔' },
{ key: '主面料成分', value: '莫代尔纤维' },
]);
});
});
-69
View File
@@ -1,69 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './item.js';
describe('1688 item normalization', () => {
it('normalizes public item payload into contract fields', () => {
const result = __test__.normalizeItemPayload({
href: 'https://detail.1688.com/offer/887904326744.html',
title: '法式春季长袖开衫连衣裙女新款大码女装碎花吊带裙套装142077 - 阿里巴巴',
bodyText: `
青岛沁澜衣品服装有限公司
入驻13年
主营:大码女装
店铺回头率
87%
山东青岛
3套起批
已售1600+套
支持定制logo
`,
offerTitle: '法式春季长袖开衫连衣裙女新款大码女装碎花吊带裙套装142077',
offerId: 887904326744,
seller: {
companyName: '青岛沁澜衣品服装有限公司',
memberId: 'b2b-1641351767',
winportUrl: 'https://yinuoweierfushi.1688.com/page/index.html?spm=a1',
},
trade: {
beginAmount: 3,
priceDisplay: '96.00-98.00',
unit: '套',
saleCount: 1655,
offerIDatacenterSellInfo: {
: '莫代尔',
: '莫代尔纤维',
sellPointModel: '{"ignore":true}',
},
offerPriceModel: {
currentPrices: [
{ beginAmount: 3, price: '98.00' },
{ beginAmount: 50, price: '97.00' },
],
},
},
gallery: {
mainImage: ['https://example.com/1.jpg'],
offerImgList: ['https://example.com/2.jpg'],
wlImageInfos: [{ fullPathImageURI: 'https://example.com/3.jpg' }],
},
services: [
{ serviceName: '延期必赔', agreeDeliveryHours: 360 },
{ serviceName: '品质保障' },
],
});
expect(result.offer_id).toBe('887904326744');
expect(result.member_id).toBe('b2b-1641351767');
expect(result.shop_id).toBe('yinuoweierfushi');
expect(result.seller_url).toBe('https://yinuoweierfushi.1688.com');
expect(result.price_text).toBe('¥96.00-98.00');
expect(result.moq_text).toBe('3套起批');
expect(result.origin_place).toBe('山东青岛');
expect(result.delivery_days_text).toBe('360小时内发货');
expect(result.private_label_text).toBe('支持定制logo');
expect(result.visible_attributes).toEqual([
{ key: '面料名称', value: '莫代尔' },
{ key: '主面料成分', value: '莫代尔纤维' },
]);
});
});
-282
View File
@@ -1,282 +0,0 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import { isRecord } from '@jackwener/opencli/utils';
import {
assertAuthenticatedState,
buildDetailUrl,
buildProvenance,
canonicalizeSellerUrl,
cleanMultilineText,
cleanText,
extractLocation,
extractMemberId,
extractOfferId,
extractShopId,
gotoAndReadState,
normalizePriceTiers,
parseMoqText,
parsePriceText,
toNumber,
uniqueNonEmpty,
} from './shared.js';
interface BuyerProtectionModel {
serviceName?: string;
shortBuyerDesc?: string;
packageBuyerDesc?: string;
textDesc?: string;
agreeDeliveryHours?: number;
}
interface ItemBrowserPayload {
href?: string;
title?: string;
bodyText?: string;
offerTitle?: string;
offerId?: string | number;
seller?: {
companyName?: string;
memberId?: string;
winportUrl?: string;
sellerWinportUrlMap?: Record<string, string>;
};
trade?: {
beginAmount?: string | number;
priceDisplay?: string;
unit?: string;
saleCount?: string | number;
offerIDatacenterSellInfo?: Record<string, unknown>;
offerPriceModel?: {
currentPrices?: Array<{ beginAmount?: string | number; price?: string | number }>;
};
};
gallery?: {
mainImage?: string[];
offerImgList?: string[];
wlImageInfos?: Array<{ fullPathImageURI?: string }>;
};
shipping?: {
deliveryLimitText?: string;
logisticsText?: string;
protectionInfos?: BuyerProtectionModel[];
buyerProtectionModel?: BuyerProtectionModel[];
};
services?: BuyerProtectionModel[];
}
interface VisibleAttribute {
key: string;
value: string;
}
function normalizeItemPayload(payload: ItemBrowserPayload): Record<string, unknown> {
const href = cleanText(payload.href);
const bodyText = cleanMultilineText(payload.bodyText);
const sellerName = cleanText(payload.seller?.companyName);
const sellerUrlRaw = cleanText(
payload.seller?.winportUrl
?? payload.seller?.sellerWinportUrlMap?.defaultUrl
?? payload.seller?.sellerWinportUrlMap?.indexUrl,
);
const sellerUrl = canonicalizeSellerUrl(sellerUrlRaw);
const offerId = cleanText(String(payload.offerId ?? '')) || extractOfferId(href) || null;
const memberId = cleanText(payload.seller?.memberId) || extractMemberId(sellerUrlRaw || href) || null;
const shopId = extractShopId(sellerUrl ?? href);
const unit = cleanText(payload.trade?.unit);
const priceDisplay = cleanText(payload.trade?.priceDisplay);
const priceRange = parsePriceText(priceDisplay ? `¥${priceDisplay}` : bodyText);
const moqText = extractMoqText(bodyText, payload.trade?.beginAmount, unit);
const moq = parseMoqText(moqText);
const services = uniqueServices(payload);
const serviceBadges = uniqueNonEmpty(services.map((service) => cleanText(service.serviceName)));
const attributes = normalizeVisibleAttributes(payload.trade?.offerIDatacenterSellInfo);
const priceTiers = normalizePriceTiers(payload.trade?.offerPriceModel?.currentPrices ?? [], unit || null);
const images = uniqueNonEmpty([
...(payload.gallery?.mainImage ?? []),
...(payload.gallery?.offerImgList ?? []),
...((payload.gallery?.wlImageInfos ?? []).map((item) => item.fullPathImageURI ?? '')),
]);
const detailUrl = offerId ? buildDetailUrl(offerId) : href;
const provenance = buildProvenance(href || detailUrl);
return {
offer_id: offerId,
member_id: memberId,
shop_id: shopId,
title: cleanText(payload.offerTitle) || stripAlibabaSuffix(payload.title) || firstNonEmptyLine(bodyText) || null,
item_url: detailUrl,
main_images: images,
price_text: priceRange.price_text || null,
price_tiers: priceTiers,
currency: priceRange.currency,
moq_text: moq.moq_text || null,
moq_value: moq.moq_value,
seller_name: sellerName || null,
seller_url: sellerUrl,
shop_name: sellerName || null,
origin_place: extractLocation(bodyText),
delivery_days_text: extractDeliveryDaysText(bodyText, services, payload.shipping),
customization_text: extractKeywordLine(bodyText, ['来样定制', '来图定制', '支持定制', '可定制', '定制']),
private_label_text: extractKeywordLine(bodyText, ['贴牌', '贴标', '定制logo', '打logo', 'OEM', 'ODM']),
visible_attributes: attributes,
sales_text: extractSalesText(bodyText),
service_badges: serviceBadges,
stock_quantity: extractStockQuantity(bodyText),
...provenance,
};
}
function normalizeVisibleAttributes(raw: unknown): VisibleAttribute[] {
if (!isRecord(raw)) return [];
return Object.entries(raw)
.filter(([key, value]) => key !== 'sellPointModel' && cleanText(key) && cleanText(String(value)))
.map(([key, value]) => ({ key: cleanText(key), value: cleanText(String(value)) }));
}
function uniqueServices(payload: ItemBrowserPayload): BuyerProtectionModel[] {
const combined = [
...(Array.isArray(payload.services) ? payload.services : []),
...(Array.isArray(payload.shipping?.protectionInfos) ? payload.shipping.protectionInfos : []),
...(Array.isArray(payload.shipping?.buyerProtectionModel) ? payload.shipping.buyerProtectionModel : []),
];
const seen = new Set<string>();
const result: BuyerProtectionModel[] = [];
for (const service of combined) {
const key = cleanText(service.serviceName);
if (!key || seen.has(key)) continue;
seen.add(key);
result.push(service);
}
return result;
}
function stripAlibabaSuffix(title: string | undefined): string {
return cleanText(title).replace(/\s*-\s*阿里巴巴$/, '').trim();
}
function firstNonEmptyLine(text: string): string {
return text.split('\n').map((line) => cleanText(line)).find(Boolean) ?? '';
}
function extractMoqText(bodyText: string, beginAmount: string | number | undefined, unit: string): string {
const lineMatch = bodyText.match(/\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)\s*起批/);
if (lineMatch) return lineMatch[0];
const moqValue = toNumber(beginAmount);
if (moqValue !== null) {
return `${moqValue}${unit || ''}起批`;
}
return '';
}
function extractDeliveryDaysText(
bodyText: string,
services: BuyerProtectionModel[],
shipping: ItemBrowserPayload['shipping'],
): string | null {
const shippingText = cleanText(shipping?.deliveryLimitText) || cleanText(shipping?.logisticsText);
if (shippingText) return shippingText;
const textMatch = bodyText.match(/\d+\s*(?:小时|天)(?:内)?发货/);
if (textMatch) return textMatch[0];
const hourMatch = services.find((service) => typeof service.agreeDeliveryHours === 'number');
if (hourMatch && typeof hourMatch.agreeDeliveryHours === 'number') {
return `${hourMatch.agreeDeliveryHours}小时内发货`;
}
return null;
}
function extractKeywordLine(bodyText: string, keywords: string[]): string | null {
const lines = bodyText.split('\n').map((line) => cleanText(line)).filter(Boolean);
for (const line of lines) {
if (keywords.some((keyword) => line.includes(keyword))) {
return line;
}
}
return null;
}
function extractSalesText(bodyText: string): string | null {
const match = bodyText.match(/(?:全网销量|已售)\s*\d+(?:\.\d+)?\+?\s*[件套个单]?/);
return match ? cleanText(match[0]) : null;
}
function extractStockQuantity(bodyText: string): number | null {
const match = bodyText.match(/库存\s*(\d+)/);
return match ? Number.parseInt(match[1], 10) : null;
}
async function readItemPayload(page: IPage, itemUrl: string): Promise<ItemBrowserPayload> {
const state = await gotoAndReadState(page, itemUrl, 2500, 'item');
assertAuthenticatedState(state, 'item');
const payload = await page.evaluate(`
(() => {
const root = window.context ?? {};
const model = root.result?.global?.globalData?.model ?? null;
const toJson = (value) => JSON.parse(JSON.stringify(value ?? null));
return {
href: window.location.href,
title: document.title || '',
bodyText: document.body ? document.body.innerText || '' : '',
offerTitle: model?.offerTitleModel?.subject ?? '',
offerId: model?.tradeModel?.offerId ?? '',
seller: toJson(model?.sellerModel),
trade: toJson(model?.tradeModel),
gallery: toJson(root.result?.data?.gallery?.fields ?? null),
shipping: toJson(root.result?.data?.shippingServices?.fields ?? null),
services: toJson(root.result?.data?.shippingServices?.fields?.protectionInfos ?? []),
};
})()
`) as ItemBrowserPayload;
const resolvedOfferId = cleanText(String(payload.offerId ?? '')) || extractOfferId(cleanText(payload.href));
if (!resolvedOfferId) {
throw new CommandExecutionError(
'1688 item page did not expose product context',
'当前 tab 非商品详情上下文,请切到 detail.1688.com 商品页并重试',
);
}
return payload;
}
cli({
site: '1688',
name: 'item',
description: '1688 商品详情(公开商品字段、价格阶梯、卖家基础信息)',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 商品 URL 或 offer ID(如 887904326744',
},
],
columns: ['offer_id', 'title', 'price_text', 'moq_text', 'seller_name', 'origin_place'],
func: async (page, kwargs) => {
const itemUrl = buildDetailUrl(String(kwargs.input ?? ''));
const payload = await readItemPayload(page, itemUrl);
return [normalizeItemPayload(payload)];
},
});
export const __test__ = {
normalizeItemPayload,
normalizeVisibleAttributes,
stripAlibabaSuffix,
extractMoqText,
extractDeliveryDaysText,
extractKeywordLine,
extractSalesText,
extractStockQuantity,
};
+309
View File
@@ -0,0 +1,309 @@
import { CommandExecutionError, EmptyResultError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { FACTORY_BADGE_PATTERNS, SERVICE_BADGE_PATTERNS, assertAuthenticatedState, buildProvenance, buildSearchUrl, canonicalizeItemUrl, canonicalizeSellerUrl, cleanText, extractBadges, extractLocation, extractMemberId, extractOfferId, extractShopId, gotoAndReadState, parseMoqText, parsePriceText, SEARCH_LIMIT_DEFAULT, SEARCH_LIMIT_MAX, parseSearchLimit, uniqueNonEmpty, } from './shared.js';
const SEARCH_ITEM_URL_PATTERNS = [
'detail.1688.com/offer/',
'detail.m.1688.com/page/index.html?offerId=',
];
const MAX_SEARCH_PAGES = 12;
function normalizeSearchCandidate(candidate, sourceUrl) {
const canonicalItemUrl = canonicalizeItemUrl(cleanText(candidate.item_url));
const containerText = cleanText(candidate.container_text);
const priceText = firstNonEmpty([
normalizeInlineText(candidate.price_text),
normalizeInlineText(extractPriceText(candidate.hover_price_text)),
]);
const priceRange = parsePriceText(priceText || containerText);
const moq = parseMoqText(firstNonEmpty([
normalizeInlineText(candidate.moq_text),
normalizeInlineText(extractMoqText(containerText)),
]));
const canonicalSellerUrl = canonicalizeSellerUrl(cleanText(candidate.seller_url));
const evidenceText = uniqueNonEmpty([
containerText,
...(candidate.desc_rows ?? []),
...(candidate.tag_items ?? []),
...(candidate.hover_items ?? []),
]).join('\n');
const badges = extractBadges(evidenceText, [...FACTORY_BADGE_PATTERNS, ...SERVICE_BADGE_PATTERNS]);
const salesText = firstNonEmpty([
extractSalesText(candidate.sales_text),
extractSalesText(containerText),
]);
const returnRateText = extractReturnRateText([...(candidate.tag_items ?? []), ...(candidate.hover_items ?? [])]);
const provenance = buildProvenance(sourceUrl);
return {
rank: 0,
offer_id: extractOfferId(canonicalItemUrl ?? '') ?? null,
member_id: extractMemberId(canonicalSellerUrl ?? '') ?? null,
shop_id: extractShopId(canonicalSellerUrl ?? '') ?? null,
title: cleanText(candidate.title) || firstWord(containerText) || null,
item_url: canonicalItemUrl,
seller_name: cleanText(candidate.seller_name) || null,
seller_url: canonicalSellerUrl,
price_text: priceRange.price_text || null,
price_min: priceRange.price_min,
price_max: priceRange.price_max,
currency: priceRange.currency,
moq_text: moq.moq_text || null,
moq_value: moq.moq_value,
location: extractLocation(containerText),
badges,
sales_text: salesText || null,
return_rate_text: returnRateText,
source_url: provenance.source_url,
fetched_at: provenance.fetched_at,
strategy: provenance.strategy,
};
}
function extractMoqText(text) {
const normalized = normalizeInlineText(text);
return normalized.match(/\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)\s*起批/i)?.[0]
?? normalized.match(/≥\s*\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)?/i)?.[0]
?? normalized.match(/\d+(?:\.\d+)?\s*(?:~|-|至|到)\s*\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)/i)?.[0]
?? '';
}
function extractPriceText(text) {
const normalized = normalizeInlineText(text);
return normalized.match(/[¥$€]\s*\d+(?:\.\d+)?/)?.[0] ?? '';
}
function extractSalesText(text) {
const normalized = normalizeInlineText(text);
if (!normalized)
return '';
if (/^\d+(?:\.\d+)?\+?\s*(件|套|个|单)$/.test(normalized)) {
return normalized;
}
const match = normalized.match(/(?:已售|销量|售)\s*\d+(?:\.\d+)?\+?\s*(件|套|个|单)?/);
return match ? cleanText(match[0]) : '';
}
function firstWord(text) {
return text.split(/\s+/).find(Boolean) ?? '';
}
function firstNonEmpty(values) {
return values.map((value) => cleanText(value)).find(Boolean) ?? '';
}
function normalizeInlineText(text) {
return cleanText(text)
.replace(/([¥$€])\s+(?=\d)/g, '$1')
.replace(/(\d)\s*\.\s*(\d)/g, '$1.$2')
.replace(/\s*([~-])\s*/g, '$1')
.trim();
}
function extractReturnRateText(values) {
return uniqueNonEmpty(values.map((value) => normalizeInlineText(value)))
.find((value) => /^回头率\s*\d+(?:\.\d+)?%$/.test(value))
?? null;
}
function buildDedupeKey(row) {
if (row.offer_id)
return `offer:${row.offer_id}`;
if (row.item_url)
return `url:${row.item_url}`;
return null;
}
async function readSearchPayload(page, url) {
const state = await gotoAndReadState(page, url, 2500, 'search');
assertAuthenticatedState(state, 'search');
const payload = await page.evaluate(`
(() => {
const normalizeText = (value) => (value || '').replace(/\\s+/g, ' ').trim();
const normalizeUrl = (href) => {
if (!href) return '';
try {
return new URL(href, window.location.href).toString();
} catch {
return '';
}
};
const isItemHref = (href) => ${JSON.stringify(SEARCH_ITEM_URL_PATTERNS)}
.some((pattern) => (href || '').includes(pattern));
const uniqueTexts = (values) => [...new Set(values.map((value) => normalizeText(value)).filter(Boolean))];
const collectTexts = (root, selector) => uniqueTexts(
Array.from(root.querySelectorAll(selector)).map((node) => node.innerText || node.textContent || ''),
);
const firstText = (root, selectors) => {
for (const selector of selectors) {
const node = root.querySelector(selector);
const value = normalizeText(node ? node.innerText || node.textContent || '' : '');
if (value) return value;
}
return '';
};
const findMoqText = (values, priceText) => {
const moqPattern = /(≥\\s*\\d+(?:\\.\\d+)?\\s*(件|个|套|箱|包|双|台|把|只)?)|(\\d+(?:\\.\\d+)?\\s*(?:~|-|至|到)\\s*\\d+(?:\\.\\d+)?\\s*(件|个|套|箱|包|双|台|把|只))|(\\d+(?:\\.\\d+)?\\s*(件|个|套|箱|包|双|台|把|只)\\s*起批)/i;
return values.find((value) => moqPattern.test(value))
|| normalizeText(priceText).match(moqPattern)?.[0]
|| '';
};
const isSellerHref = (href) => {
if (!href) return false;
try {
const url = new URL(href, window.location.href);
const host = url.hostname || '';
if (!host.endsWith('.1688.com')) return false;
if (
host === 's.1688.com'
|| host === 'r.1688.com'
|| host === 'air.1688.com'
|| host === 'detail.1688.com'
|| host === 'detail.m.1688.com'
|| host === 'dj.1688.com'
) {
return false;
}
return true;
} catch {
return false;
}
};
const pickContainer = (anchor) => {
let node = anchor;
while (node && node !== document.body) {
const text = normalizeText(node.innerText || node.textContent || '');
if (text.length >= 40 && text.length <= 2000) {
return node;
}
node = node.parentElement;
}
return anchor;
};
const collectCandidates = () => {
const anchors = Array.from(document.querySelectorAll('a')).filter((anchor) => isItemHref(anchor.href || ''));
const seen = new Set();
const items = [];
for (const anchor of anchors) {
const href = anchor.href || '';
if (!href || seen.has(href)) continue;
seen.add(href);
const container = pickContainer(anchor);
const tagItems = collectTexts(container, '.offer-tag-row .offer-desc-item');
const hoverItems = collectTexts(container, '.offer-hover-wrapper .offer-desc-item');
const sellerAnchor = Array.from(container.querySelectorAll('a'))
.find((link) => isSellerHref(link.href || ''));
const hoverPriceText = firstText(container, [
'.offer-hover-wrapper .hover-price-item',
'.offer-hover-wrapper .price-item',
]);
items.push({
item_url: href,
title: firstText(container, ['.offer-title-row .title-text', '.offer-title-row'])
|| normalizeText(anchor.innerText || anchor.textContent || ''),
container_text: normalizeText(container.innerText || container.textContent || ''),
desc_rows: collectTexts(container, '.offer-desc-row'),
price_text: firstText(container, ['.offer-price-row .price-item']),
sales_text: firstText(container, ['.offer-price-row .col-desc_after', '.offer-desc-row .col-desc_after']),
hover_price_text: hoverPriceText,
moq_text: findMoqText(hoverItems, hoverPriceText),
tag_items: tagItems,
hover_items: hoverItems,
seller_name: sellerAnchor ? normalizeText(sellerAnchor.innerText || sellerAnchor.textContent || '') : null,
seller_url: sellerAnchor ? sellerAnchor.href : null,
});
}
return items;
};
const findNextUrl = () => {
const selectors = [
'a.fui-next:not(.disabled)',
'a.next-pagination-item:not(.disabled)',
'a[rel="next"]:not(.disabled)',
'a[data-role="next"]:not(.disabled)',
];
for (const selector of selectors) {
const node = document.querySelector(selector);
if (!node) continue;
const href = normalizeUrl(node.getAttribute('href') || node.href || '');
if (href) return href;
}
const textBased = Array.from(document.querySelectorAll('a'))
.find((node) => /下一页|next/i.test(normalizeText(node.textContent || '')));
if (!textBased) return '';
return normalizeUrl(textBased.getAttribute('href') || textBased.href || '');
};
return {
href: window.location.href,
title: document.title || '',
bodyText: document.body ? document.body.innerText || '' : '',
next_url: findNextUrl(),
candidates: collectCandidates(),
};
})()
`);
if (!payload || typeof payload !== 'object') {
throw new CommandExecutionError('1688 search page did not return a readable payload', 'Open the same query in Chrome and verify the page is fully loaded before retrying.');
}
return payload;
}
async function collectSearchRows(page, query, limit) {
const rowsByKey = new Map();
const seenPages = new Set();
let nextUrl = buildSearchUrl(query);
let pageCount = 0;
while (nextUrl && rowsByKey.size < limit && pageCount < MAX_SEARCH_PAGES) {
if (seenPages.has(nextUrl))
break;
seenPages.add(nextUrl);
pageCount += 1;
const payload = await readSearchPayload(page, nextUrl);
const sourceUrl = cleanText(payload.href) || nextUrl;
const candidates = Array.isArray(payload.candidates) ? payload.candidates : [];
for (const candidate of candidates) {
const row = normalizeSearchCandidate(candidate, sourceUrl);
const dedupeKey = buildDedupeKey(row);
if (!dedupeKey || rowsByKey.has(dedupeKey))
continue;
rowsByKey.set(dedupeKey, row);
if (rowsByKey.size >= limit)
break;
}
const candidateNextUrl = cleanText(payload.next_url);
if (!candidateNextUrl || candidateNextUrl === sourceUrl)
break;
nextUrl = candidateNextUrl;
}
if (rowsByKey.size === 0) {
throw new EmptyResultError('1688 search', 'No visible results were extracted. Retry with a different query or open the same search page in Chrome first.');
}
return [...rowsByKey.values()]
.slice(0, limit)
.map((row, index) => ({ ...row, rank: index + 1 }));
}
cli({
site: '1688',
name: 'search',
description: '1688 商品搜索(结果候选、卖家链接、价格/MOQ/销量文本)',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'query',
required: true,
positional: true,
help: '搜索关键词,如 "置物架"',
},
{
name: 'limit',
type: 'int',
default: SEARCH_LIMIT_DEFAULT,
help: `结果数量上限(默认 ${SEARCH_LIMIT_DEFAULT},最大 ${SEARCH_LIMIT_MAX}`,
},
],
columns: ['rank', 'title', 'price_text', 'moq_text', 'seller_name', 'location'],
func: async (page, kwargs) => {
const query = String(kwargs.query ?? '');
const limit = parseSearchLimit(kwargs.limit);
return collectSearchRows(page, query, limit);
},
});
export const __test__ = {
normalizeSearchCandidate,
extractMoqText,
extractSalesText,
firstWord,
buildDedupeKey,
};
+75
View File
@@ -0,0 +1,75 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './search.js';
describe('1688 search normalization', () => {
it('normalizes search candidates into structured result rows', () => {
const result = __test__.normalizeSearchCandidate({
item_url: 'https://detail.1688.com/offer/887904326744.html',
title: '宿舍置物架桌面加高架',
container_text: '宿舍置物架桌面加高架 ¥56.00 2套起批 山东青岛 已售300+套',
price_text: '¥ 56 .00',
sales_text: '300+套',
moq_text: '2套起批',
tag_items: ['退货包运费', '回头率52%'],
hover_items: ['验厂报告'],
seller_name: '青岛沁澜衣品服装有限公司',
seller_url: 'https://yinuoweierfushi.1688.com/page/index.html?spm=a123',
}, 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=置物架');
expect(result.rank).toBe(0);
expect(result.offer_id).toBe('887904326744');
expect(result.shop_id).toBe('yinuoweierfushi');
expect(result.item_url).toBe('https://detail.1688.com/offer/887904326744.html');
expect(result.seller_url).toBe('https://yinuoweierfushi.1688.com');
expect(result.price_text).toBe('¥56.00');
expect(result.price_min).toBe(56);
expect(result.price_max).toBe(56);
expect(result.moq_value).toBe(2);
expect(result.location).toBe('山东青岛');
expect(result.sales_text).toBe('300+套');
expect(result.badges).toEqual(expect.arrayContaining(['退货包运费', '验厂报告']));
expect(result.return_rate_text).toBe('回头率52%');
});
it('does not use hover_price_text as MOQ source', () => {
const result = __test__.normalizeSearchCandidate({
item_url: 'https://detail.1688.com/offer/887904326744.html',
title: 'test',
container_text: 'test ¥56.00',
price_text: '¥ 56 .00',
hover_price_text: '¥56.00 3件起批',
moq_text: null,
}, 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=test');
// hover_price_text should not be used for MOQ extraction
expect(result.moq_text).toBeNull();
expect(result.moq_value).toBeNull();
});
it('extracts offer id from mobile detail search links', () => {
const result = __test__.normalizeSearchCandidate({
item_url: 'http://detail.m.1688.com/page/index.html?offerId=910933345396&sortType=&pageId=',
title: '',
container_text: '桌面书桌办公室工位收纳展示新中式博古架多层茶具厨房摆放置物架 ¥24.3 已售20+件',
price_text: '¥ 14 .28',
sales_text: '1500+件',
moq_text: '≥2个',
seller_name: '泰商国际贸易(宁阳)有限公司',
seller_url: 'http://tsgjmy.1688.com/',
}, 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=桌面置物架');
expect(result.offer_id).toBe('910933345396');
expect(result.shop_id).toBe('tsgjmy');
expect(result.item_url).toBe('https://detail.1688.com/offer/910933345396.html');
expect(result.title).toContain('桌面书桌办公室工位收纳展示');
expect(result.price_text).toBe('¥14.28');
expect(result.sales_text).toBe('1500+件');
expect(result.moq_text).toBe('≥2个');
expect(result.moq_value).toBe(2);
});
it('prefers offer id and falls back to item url for dedupe key', () => {
expect(__test__.buildDedupeKey({
offer_id: '123456',
item_url: 'https://detail.1688.com/offer/123456.html',
})).toBe('offer:123456');
expect(__test__.buildDedupeKey({
offer_id: null,
item_url: 'https://detail.1688.com/offer/123456.html',
})).toBe('url:https://detail.1688.com/offer/123456.html');
expect(__test__.buildDedupeKey({ offer_id: null, item_url: null })).toBeNull();
});
});
-81
View File
@@ -1,81 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './search.js';
describe('1688 search normalization', () => {
it('normalizes search candidates into structured result rows', () => {
const result = __test__.normalizeSearchCandidate({
item_url: 'https://detail.1688.com/offer/887904326744.html',
title: '宿舍置物架桌面加高架',
container_text: '宿舍置物架桌面加高架 ¥56.00 2套起批 山东青岛 已售300+套',
price_text: '¥ 56 .00',
sales_text: '300+套',
moq_text: '2套起批',
tag_items: ['退货包运费', '回头率52%'],
hover_items: ['验厂报告'],
seller_name: '青岛沁澜衣品服装有限公司',
seller_url: 'https://yinuoweierfushi.1688.com/page/index.html?spm=a123',
}, 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=置物架');
expect(result.rank).toBe(0);
expect(result.offer_id).toBe('887904326744');
expect(result.shop_id).toBe('yinuoweierfushi');
expect(result.item_url).toBe('https://detail.1688.com/offer/887904326744.html');
expect(result.seller_url).toBe('https://yinuoweierfushi.1688.com');
expect(result.price_text).toBe('¥56.00');
expect(result.price_min).toBe(56);
expect(result.price_max).toBe(56);
expect(result.moq_value).toBe(2);
expect(result.location).toBe('山东青岛');
expect(result.sales_text).toBe('300+套');
expect(result.badges).toEqual(expect.arrayContaining(['退货包运费', '验厂报告']));
expect(result.return_rate_text).toBe('回头率52%');
});
it('does not use hover_price_text as MOQ source', () => {
const result = __test__.normalizeSearchCandidate({
item_url: 'https://detail.1688.com/offer/887904326744.html',
title: 'test',
container_text: 'test ¥56.00',
price_text: '¥ 56 .00',
hover_price_text: '¥56.00 3件起批',
moq_text: null,
}, 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=test');
// hover_price_text should not be used for MOQ extraction
expect(result.moq_text).toBeNull();
expect(result.moq_value).toBeNull();
});
it('extracts offer id from mobile detail search links', () => {
const result = __test__.normalizeSearchCandidate({
item_url: 'http://detail.m.1688.com/page/index.html?offerId=910933345396&sortType=&pageId=',
title: '',
container_text: '桌面书桌办公室工位收纳展示新中式博古架多层茶具厨房摆放置物架 ¥24.3 已售20+件',
price_text: '¥ 14 .28',
sales_text: '1500+件',
moq_text: '≥2个',
seller_name: '泰商国际贸易(宁阳)有限公司',
seller_url: 'http://tsgjmy.1688.com/',
}, 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=桌面置物架');
expect(result.offer_id).toBe('910933345396');
expect(result.shop_id).toBe('tsgjmy');
expect(result.item_url).toBe('https://detail.1688.com/offer/910933345396.html');
expect(result.title).toContain('桌面书桌办公室工位收纳展示');
expect(result.price_text).toBe('¥14.28');
expect(result.sales_text).toBe('1500+件');
expect(result.moq_text).toBe('≥2个');
expect(result.moq_value).toBe(2);
});
it('prefers offer id and falls back to item url for dedupe key', () => {
expect(__test__.buildDedupeKey({
offer_id: '123456',
item_url: 'https://detail.1688.com/offer/123456.html',
})).toBe('offer:123456');
expect(__test__.buildDedupeKey({
offer_id: null,
item_url: 'https://detail.1688.com/offer/123456.html',
})).toBe('url:https://detail.1688.com/offer/123456.html');
expect(__test__.buildDedupeKey({ offer_id: null, item_url: null })).toBeNull();
});
});
-402
View File
@@ -1,402 +0,0 @@
import { CommandExecutionError, EmptyResultError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
FACTORY_BADGE_PATTERNS,
SERVICE_BADGE_PATTERNS,
assertAuthenticatedState,
buildProvenance,
buildSearchUrl,
canonicalizeItemUrl,
canonicalizeSellerUrl,
cleanText,
extractBadges,
extractLocation,
extractMemberId,
extractOfferId,
extractShopId,
gotoAndReadState,
parseMoqText,
parsePriceText,
SEARCH_LIMIT_DEFAULT,
SEARCH_LIMIT_MAX,
parseSearchLimit,
uniqueNonEmpty,
} from './shared.js';
interface SearchPayload {
href?: string;
title?: string;
bodyText?: string;
next_url?: string;
candidates?: Array<{
item_url?: string;
title?: string;
container_text?: string;
desc_rows?: string[];
price_text?: string | null;
sales_text?: string | null;
hover_price_text?: string | null;
moq_text?: string | null;
tag_items?: string[];
hover_items?: string[];
seller_name?: string | null;
seller_url?: string | null;
}>;
}
interface SearchRow {
rank: number;
offer_id: string | null;
member_id: string | null;
shop_id: string | null;
title: string | null;
item_url: string | null;
seller_name: string | null;
seller_url: string | null;
price_text: string | null;
price_min: number | null;
price_max: number | null;
currency: string | null;
moq_text: string | null;
moq_value: number | null;
location: string | null;
badges: string[];
sales_text: string | null;
return_rate_text: string | null;
source_url: string;
fetched_at: string;
strategy: string;
}
const SEARCH_ITEM_URL_PATTERNS = [
'detail.1688.com/offer/',
'detail.m.1688.com/page/index.html?offerId=',
];
const MAX_SEARCH_PAGES = 12;
function normalizeSearchCandidate(
candidate: NonNullable<SearchPayload['candidates']>[number],
sourceUrl: string,
): SearchRow {
const canonicalItemUrl = canonicalizeItemUrl(cleanText(candidate.item_url));
const containerText = cleanText(candidate.container_text);
const priceText = firstNonEmpty([
normalizeInlineText(candidate.price_text),
normalizeInlineText(extractPriceText(candidate.hover_price_text)),
]);
const priceRange = parsePriceText(priceText || containerText);
const moq = parseMoqText(firstNonEmpty([
normalizeInlineText(candidate.moq_text),
normalizeInlineText(extractMoqText(containerText)),
]));
const canonicalSellerUrl = canonicalizeSellerUrl(cleanText(candidate.seller_url));
const evidenceText = uniqueNonEmpty([
containerText,
...(candidate.desc_rows ?? []),
...(candidate.tag_items ?? []),
...(candidate.hover_items ?? []),
]).join('\n');
const badges = extractBadges(evidenceText, [...FACTORY_BADGE_PATTERNS, ...SERVICE_BADGE_PATTERNS]);
const salesText = firstNonEmpty([
extractSalesText(candidate.sales_text),
extractSalesText(containerText),
]);
const returnRateText = extractReturnRateText([...(candidate.tag_items ?? []), ...(candidate.hover_items ?? [])]);
const provenance = buildProvenance(sourceUrl);
return {
rank: 0,
offer_id: extractOfferId(canonicalItemUrl ?? '') ?? null,
member_id: extractMemberId(canonicalSellerUrl ?? '') ?? null,
shop_id: extractShopId(canonicalSellerUrl ?? '') ?? null,
title: cleanText(candidate.title) || firstWord(containerText) || null,
item_url: canonicalItemUrl,
seller_name: cleanText(candidate.seller_name) || null,
seller_url: canonicalSellerUrl,
price_text: priceRange.price_text || null,
price_min: priceRange.price_min,
price_max: priceRange.price_max,
currency: priceRange.currency,
moq_text: moq.moq_text || null,
moq_value: moq.moq_value,
location: extractLocation(containerText),
badges,
sales_text: salesText || null,
return_rate_text: returnRateText,
source_url: provenance.source_url,
fetched_at: provenance.fetched_at,
strategy: provenance.strategy,
};
}
function extractMoqText(text: string | null | undefined): string {
const normalized = normalizeInlineText(text);
return normalized.match(/\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)\s*起批/i)?.[0]
?? normalized.match(/≥\s*\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)?/i)?.[0]
?? normalized.match(/\d+(?:\.\d+)?\s*(?:~|-|至|到)\s*\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只)/i)?.[0]
?? '';
}
function extractPriceText(text: string | null | undefined): string {
const normalized = normalizeInlineText(text);
return normalized.match(/[¥$€]\s*\d+(?:\.\d+)?/)?.[0] ?? '';
}
function extractSalesText(text: string | null | undefined): string {
const normalized = normalizeInlineText(text);
if (!normalized) return '';
if (/^\d+(?:\.\d+)?\+?\s*(件|套|个|单)$/.test(normalized)) {
return normalized;
}
const match = normalized.match(/(?:已售|销量|售)\s*\d+(?:\.\d+)?\+?\s*(件|套|个|单)?/);
return match ? cleanText(match[0]) : '';
}
function firstWord(text: string): string {
return text.split(/\s+/).find(Boolean) ?? '';
}
function firstNonEmpty(values: Array<string | null | undefined>): string {
return values.map((value) => cleanText(value)).find(Boolean) ?? '';
}
function normalizeInlineText(text: string | null | undefined): string {
return cleanText(text)
.replace(/([¥$€])\s+(?=\d)/g, '$1')
.replace(/(\d)\s*\.\s*(\d)/g, '$1.$2')
.replace(/\s*([~-])\s*/g, '$1')
.trim();
}
function extractReturnRateText(values: string[]): string | null {
return uniqueNonEmpty(values.map((value) => normalizeInlineText(value)))
.find((value) => /^回头率\s*\d+(?:\.\d+)?%$/.test(value))
?? null;
}
function buildDedupeKey(row: Pick<SearchRow, 'offer_id' | 'item_url'>): string | null {
if (row.offer_id) return `offer:${row.offer_id}`;
if (row.item_url) return `url:${row.item_url}`;
return null;
}
async function readSearchPayload(page: IPage, url: string): Promise<SearchPayload> {
const state = await gotoAndReadState(page, url, 2500, 'search');
assertAuthenticatedState(state, 'search');
const payload = await page.evaluate(`
(() => {
const normalizeText = (value) => (value || '').replace(/\\s+/g, ' ').trim();
const normalizeUrl = (href) => {
if (!href) return '';
try {
return new URL(href, window.location.href).toString();
} catch {
return '';
}
};
const isItemHref = (href) => ${JSON.stringify(SEARCH_ITEM_URL_PATTERNS)}
.some((pattern) => (href || '').includes(pattern));
const uniqueTexts = (values) => [...new Set(values.map((value) => normalizeText(value)).filter(Boolean))];
const collectTexts = (root, selector) => uniqueTexts(
Array.from(root.querySelectorAll(selector)).map((node) => node.innerText || node.textContent || ''),
);
const firstText = (root, selectors) => {
for (const selector of selectors) {
const node = root.querySelector(selector);
const value = normalizeText(node ? node.innerText || node.textContent || '' : '');
if (value) return value;
}
return '';
};
const findMoqText = (values, priceText) => {
const moqPattern = /(≥\\s*\\d+(?:\\.\\d+)?\\s*(件|个|套|箱|包|双|台|把|只)?)|(\\d+(?:\\.\\d+)?\\s*(?:~|-|至|到)\\s*\\d+(?:\\.\\d+)?\\s*(件|个|套|箱|包|双|台|把|只))|(\\d+(?:\\.\\d+)?\\s*(件|个|套|箱|包|双|台|把|只)\\s*起批)/i;
return values.find((value) => moqPattern.test(value))
|| normalizeText(priceText).match(moqPattern)?.[0]
|| '';
};
const isSellerHref = (href) => {
if (!href) return false;
try {
const url = new URL(href, window.location.href);
const host = url.hostname || '';
if (!host.endsWith('.1688.com')) return false;
if (
host === 's.1688.com'
|| host === 'r.1688.com'
|| host === 'air.1688.com'
|| host === 'detail.1688.com'
|| host === 'detail.m.1688.com'
|| host === 'dj.1688.com'
) {
return false;
}
return true;
} catch {
return false;
}
};
const pickContainer = (anchor) => {
let node = anchor;
while (node && node !== document.body) {
const text = normalizeText(node.innerText || node.textContent || '');
if (text.length >= 40 && text.length <= 2000) {
return node;
}
node = node.parentElement;
}
return anchor;
};
const collectCandidates = () => {
const anchors = Array.from(document.querySelectorAll('a')).filter((anchor) => isItemHref(anchor.href || ''));
const seen = new Set();
const items = [];
for (const anchor of anchors) {
const href = anchor.href || '';
if (!href || seen.has(href)) continue;
seen.add(href);
const container = pickContainer(anchor);
const tagItems = collectTexts(container, '.offer-tag-row .offer-desc-item');
const hoverItems = collectTexts(container, '.offer-hover-wrapper .offer-desc-item');
const sellerAnchor = Array.from(container.querySelectorAll('a'))
.find((link) => isSellerHref(link.href || ''));
const hoverPriceText = firstText(container, [
'.offer-hover-wrapper .hover-price-item',
'.offer-hover-wrapper .price-item',
]);
items.push({
item_url: href,
title: firstText(container, ['.offer-title-row .title-text', '.offer-title-row'])
|| normalizeText(anchor.innerText || anchor.textContent || ''),
container_text: normalizeText(container.innerText || container.textContent || ''),
desc_rows: collectTexts(container, '.offer-desc-row'),
price_text: firstText(container, ['.offer-price-row .price-item']),
sales_text: firstText(container, ['.offer-price-row .col-desc_after', '.offer-desc-row .col-desc_after']),
hover_price_text: hoverPriceText,
moq_text: findMoqText(hoverItems, hoverPriceText),
tag_items: tagItems,
hover_items: hoverItems,
seller_name: sellerAnchor ? normalizeText(sellerAnchor.innerText || sellerAnchor.textContent || '') : null,
seller_url: sellerAnchor ? sellerAnchor.href : null,
});
}
return items;
};
const findNextUrl = () => {
const selectors = [
'a.fui-next:not(.disabled)',
'a.next-pagination-item:not(.disabled)',
'a[rel="next"]:not(.disabled)',
'a[data-role="next"]:not(.disabled)',
];
for (const selector of selectors) {
const node = document.querySelector(selector);
if (!node) continue;
const href = normalizeUrl(node.getAttribute('href') || node.href || '');
if (href) return href;
}
const textBased = Array.from(document.querySelectorAll('a'))
.find((node) => /下一页|next/i.test(normalizeText(node.textContent || '')));
if (!textBased) return '';
return normalizeUrl(textBased.getAttribute('href') || textBased.href || '');
};
return {
href: window.location.href,
title: document.title || '',
bodyText: document.body ? document.body.innerText || '' : '',
next_url: findNextUrl(),
candidates: collectCandidates(),
};
})()
`) as SearchPayload;
if (!payload || typeof payload !== 'object') {
throw new CommandExecutionError(
'1688 search page did not return a readable payload',
'Open the same query in Chrome and verify the page is fully loaded before retrying.',
);
}
return payload;
}
async function collectSearchRows(page: IPage, query: string, limit: number): Promise<SearchRow[]> {
const rowsByKey = new Map<string, SearchRow>();
const seenPages = new Set<string>();
let nextUrl = buildSearchUrl(query);
let pageCount = 0;
while (nextUrl && rowsByKey.size < limit && pageCount < MAX_SEARCH_PAGES) {
if (seenPages.has(nextUrl)) break;
seenPages.add(nextUrl);
pageCount += 1;
const payload = await readSearchPayload(page, nextUrl);
const sourceUrl = cleanText(payload.href) || nextUrl;
const candidates = Array.isArray(payload.candidates) ? payload.candidates : [];
for (const candidate of candidates) {
const row = normalizeSearchCandidate(candidate, sourceUrl);
const dedupeKey = buildDedupeKey(row);
if (!dedupeKey || rowsByKey.has(dedupeKey)) continue;
rowsByKey.set(dedupeKey, row);
if (rowsByKey.size >= limit) break;
}
const candidateNextUrl = cleanText(payload.next_url);
if (!candidateNextUrl || candidateNextUrl === sourceUrl) break;
nextUrl = candidateNextUrl;
}
if (rowsByKey.size === 0) {
throw new EmptyResultError(
'1688 search',
'No visible results were extracted. Retry with a different query or open the same search page in Chrome first.',
);
}
return [...rowsByKey.values()]
.slice(0, limit)
.map((row, index) => ({ ...row, rank: index + 1 }));
}
cli({
site: '1688',
name: 'search',
description: '1688 商品搜索(结果候选、卖家链接、价格/MOQ/销量文本)',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'query',
required: true,
positional: true,
help: '搜索关键词,如 "置物架"',
},
{
name: 'limit',
type: 'int',
default: SEARCH_LIMIT_DEFAULT,
help: `结果数量上限(默认 ${SEARCH_LIMIT_DEFAULT},最大 ${SEARCH_LIMIT_MAX}`,
},
],
columns: ['rank', 'title', 'price_text', 'moq_text', 'seller_name', 'location'],
func: async (page, kwargs) => {
const query = String(kwargs.query ?? '');
const limit = parseSearchLimit(kwargs.limit);
return collectSearchRows(page, query, limit);
},
});
export const __test__ = {
normalizeSearchCandidate,
extractMoqText,
extractSalesText,
firstWord,
buildDedupeKey,
};
+557
View File
@@ -0,0 +1,557 @@
import { ArgumentError, AuthRequiredError, CommandExecutionError } from '@jackwener/opencli/errors';
export const SITE = '1688';
export const HOME_URL = 'https://www.1688.com/';
export const SEARCH_URL_PREFIX = 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=';
export const DETAIL_URL_PREFIX = 'https://detail.1688.com/offer/';
export const STORE_MOBILE_URL_PREFIX = 'https://winport.m.1688.com/page/index.html?memberId=';
export const STRATEGY = 'cookie';
export const SEARCH_LIMIT_DEFAULT = 20;
export const SEARCH_LIMIT_MAX = 100;
const STORE_GENERIC_HOSTS = new Set(['www', 'detail', 's', 'winport', 'work', 'air', 'dj']);
const TRACKING_QUERY_KEYS = new Set([
'spm',
'tracelog',
'clickid',
'source',
'scene',
'from',
'src',
'ns',
'cna',
'pvid',
]);
const CAPTCHA_URL_MARKER = '/_____tmd_____/punish';
const CAPTCHA_TEXT_PATTERNS = [
'请拖动下方滑块完成验证',
'请按住滑块,拖动到最右边',
'通过验证以确保正常访问',
'验证码拦截',
'访问验证',
'滑动验证',
];
const LOGIN_TEXT_PATTERNS = [
'请登录',
'登录后',
'账号登录',
'手机登录',
'立即登录',
'扫码登录',
'请先完成登录',
'请先登录后查看',
];
const LOGIN_URL_PATTERNS = ['/member/login', 'passport', 'login.taobao.com', 'account.1688.com'];
export const FACTORY_BADGE_PATTERNS = [
'源头工厂',
'深度验厂',
'实力工厂',
'工厂档案',
'加工专区',
'验厂报告',
'厂家直销',
'生产厂家',
'工厂直供',
];
export const SERVICE_BADGE_PATTERNS = [
'延期必赔',
'品质保障',
'破损包赔',
'退货包运费',
'晚发必赔',
'7*24小时响应',
'48小时发货',
'72小时发货',
'后天达',
'包邮',
'闪电拿样',
];
const CHINA_LOCATIONS = [
'北京',
'天津',
'上海',
'重庆',
'河北',
'山西',
'辽宁',
'吉林',
'黑龙江',
'江苏',
'浙江',
'安徽',
'福建',
'江西',
'山东',
'河南',
'湖北',
'湖南',
'广东',
'海南',
'四川',
'贵州',
'云南',
'陕西',
'甘肃',
'青海',
'台湾',
'内蒙古',
'广西',
'西藏',
'宁夏',
'新疆',
'香港',
'澳门',
];
export function cleanText(value) {
return typeof value === 'string'
? value.replace(/\u00a0/g, ' ').replace(/\s+/g, ' ').trim()
: '';
}
export function cleanMultilineText(value) {
return typeof value === 'string'
? value
.replace(/\u00a0/g, ' ')
.split('\n')
.map((line) => line.replace(/\s+/g, ' ').trim())
.filter(Boolean)
.join('\n')
: '';
}
export function uniqueNonEmpty(values) {
return [...new Set(values.map((value) => cleanText(value)).filter(Boolean))];
}
export function parseSearchLimit(input) {
const parsed = Number.parseInt(String(input ?? SEARCH_LIMIT_DEFAULT), 10);
if (!Number.isFinite(parsed) || parsed < 1) {
throw new ArgumentError('1688 search --limit must be a positive integer', 'Example: opencli 1688 search "桌面置物架" --limit 20');
}
return Math.min(SEARCH_LIMIT_MAX, parsed);
}
export function buildSearchUrl(query) {
const normalized = cleanText(query);
if (!normalized) {
throw new ArgumentError('1688 search query cannot be empty', 'Example: opencli 1688 search "桌面置物架" --limit 20');
}
return `${SEARCH_URL_PREFIX}${encodeURIComponent(normalized)}`;
}
export function buildDetailUrl(input) {
const offerId = extractOfferId(input);
if (!offerId) {
throw new ArgumentError('1688 item expects an offer URL or offer ID', 'Example: opencli 1688 item 887904326744');
}
return `${DETAIL_URL_PREFIX}${offerId}.html`;
}
export function resolveStoreUrl(input) {
const normalized = cleanText(input);
if (!normalized) {
throw new ArgumentError('1688 store expects a store URL or member ID', 'Example: opencli 1688 store https://yinuoweierfushi.1688.com/');
}
const memberId = extractMemberId(normalized);
if (memberId) {
return `${STORE_MOBILE_URL_PREFIX}${memberId}`;
}
if (/^https?:\/\//i.test(normalized)) {
return canonicalizeStoreUrl(normalized);
}
if (normalized.endsWith('.1688.com')) {
return canonicalizeStoreUrl(`https://${normalized}`);
}
if (/^[a-z0-9-]+$/i.test(normalized)) {
return canonicalizeStoreUrl(`https://${normalized}.1688.com`);
}
throw new ArgumentError('1688 store expects a store URL or member ID', 'Example: opencli 1688 store b2b-22154705262941f196');
}
export function canonicalizeStoreUrl(input) {
const url = parse1688Url(input);
const memberId = extractMemberId(url.toString());
if (memberId) {
return `${STORE_MOBILE_URL_PREFIX}${memberId}`;
}
const host = normalizeStoreHost(url.hostname);
if (!host) {
throw new ArgumentError('Invalid 1688 store URL', 'Example: opencli 1688 store https://yinuoweierfushi.1688.com/');
}
return `https://${host}`;
}
export function canonicalizeItemUrl(input) {
const offerId = extractOfferId(input);
if (offerId) {
return `${DETAIL_URL_PREFIX}${offerId}.html`;
}
const url = parse1688UrlOrNull(input);
if (!url)
return null;
stripTrackingParams(url);
url.hash = '';
return url.toString();
}
export function canonicalizeSellerUrl(input) {
const memberId = extractMemberId(input);
if (memberId) {
return `${STORE_MOBILE_URL_PREFIX}${memberId}`;
}
const url = parse1688UrlOrNull(input);
if (!url)
return null;
const host = normalizeStoreHost(url.hostname);
if (!host)
return null;
return `https://${host}`;
}
export function extractOfferId(input) {
const normalized = cleanText(input);
if (!normalized)
return null;
const directId = normalized.match(/^\d{6,}$/)?.[0];
if (directId)
return directId;
const detailMatch = normalized.match(/\/offer\/(\d{6,})\.html/i);
if (detailMatch)
return detailMatch[1];
const queryMatch = normalized.match(/[?&]offerId=(\d{6,})/i);
if (queryMatch)
return queryMatch[1];
return null;
}
export function extractMemberId(input) {
const normalized = cleanText(input);
if (!normalized)
return null;
const direct = normalized.match(/\bb2b-[a-z0-9]+\b/i)?.[0];
if (direct)
return direct;
const queryMatch = normalized.match(/[?&]memberId=(b2b-[a-z0-9]+)/i);
if (queryMatch)
return queryMatch[1];
const mobileMatch = normalized.match(/\/winport\/(b2b-[a-z0-9]+)\.html/i);
if (mobileMatch)
return mobileMatch[1];
return null;
}
export function extractShopId(input) {
const normalized = cleanText(input);
if (!normalized)
return null;
try {
const url = new URL(/^https?:\/\//i.test(normalized) ? normalized : `https://${normalized}`);
const host = normalizeStoreHost(url.hostname);
if (!host)
return null;
return host.split('.')[0] ?? null;
}
catch {
return /^[a-z0-9-]+$/i.test(normalized) ? normalized : null;
}
}
export function buildProvenance(sourceUrl) {
return {
source_url: sourceUrl,
fetched_at: new Date().toISOString(),
strategy: STRATEGY,
};
}
export function parsePriceText(text) {
const normalized = normalizeNumericText(cleanText(text));
const matches = normalized.match(/\d+(?:,\d{3})*(?:\.\d+)?/g) ?? [];
const values = matches
.map((value) => Number.parseFloat(value.replace(/,/g, '')))
.filter((value) => Number.isFinite(value));
if (values.length === 0) {
return {
price_text: normalized,
price_min: null,
price_max: null,
currency: null,
};
}
return {
price_text: normalized,
price_min: values[0] ?? null,
price_max: values[values.length - 1] ?? values[0] ?? null,
currency: normalized.includes('¥') || normalized.includes('元') ? 'CNY' : null,
};
}
export function normalizePriceTiers(rawTiers, unit) {
return rawTiers
.map((tier) => {
const quantityMin = toNumber(tier.beginAmount);
const priceText = cleanText(tier.price);
const price = toNumber(tier.price);
return {
quantity_text: quantityMin !== null ? `${quantityMin}${unit ?? ''}` : '',
quantity_min: quantityMin,
price_text: priceText,
price,
currency: priceText ? 'CNY' : null,
};
})
.filter((tier) => tier.price_text);
}
export function parseMoqText(text) {
const normalized = normalizeNumericText(cleanText(text));
const match = normalized.match(/(\d+(?:\.\d+)?)\s*(件|个|套|箱|包|双|台|把|只|pcs|piece|pieces)?\s*起批/i)
?? normalized.match(/≥\s*(\d+(?:\.\d+)?)/);
const rangeMatch = normalized.match(/(\d+(?:\.\d+)?)\s*(?:~|-|至|到)\s*\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只|pcs|piece|pieces)/i);
if (!match && !rangeMatch) {
return {
moq_text: normalized,
moq_value: null,
};
}
return {
moq_text: normalized,
moq_value: Number.parseFloat((match ?? rangeMatch)[1]),
};
}
export function extractLocation(text) {
const normalized = cleanMultilineText(text);
const primaryRegion = normalized.split(/送至|发往/)[0] ?? normalized;
const lines = primaryRegion.split('\n');
for (const line of lines) {
const compact = cleanText(line);
if (!compact || compact.length > 16)
continue;
if (CHINA_LOCATIONS.some((location) => compact.startsWith(location))) {
return compact;
}
}
const locationPattern = new RegExp(`(${CHINA_LOCATIONS.join('|')})[\\u4e00-\\u9fa5]{0,8}`);
return primaryRegion.match(locationPattern)?.[0] ?? null;
}
export function extractAddress(text) {
const normalized = cleanMultilineText(text);
const lineMatch = normalized.match(/地址[:]\s*([^\n]+)/);
if (lineMatch)
return cleanText(lineMatch[1]);
return normalized
.split('\n')
.map((line) => cleanText(line))
.find((line) => line.includes('省') || line.includes('市') || line.includes('区') || line.includes('县'))
?? null;
}
export function extractMetric(text, label) {
const normalized = cleanMultilineText(text);
const direct = normalized.match(new RegExp(`(?:^|\\n)\\s*${escapeForRegex(label)}[:]?\\s*([^\\n]+)`));
if (direct)
return cleanText(direct[1]);
const lineBased = normalized.match(new RegExp(`(?:^|\\n)\\s*${escapeForRegex(label)}\\n([^\\n]+)`));
return lineBased ? cleanText(lineBased[1]) : null;
}
export function extractYearsOnPlatform(text) {
return text.match(/入驻\d+年/)?.[0] ?? null;
}
export function extractMainBusiness(text) {
const value = extractMetric(text, '主营');
return value ? value.replace(/^/, '').trim() : null;
}
export function extractBadges(text, candidates) {
return uniqueNonEmpty(candidates.filter((candidate) => cleanMultilineText(text).includes(candidate)));
}
export function guessTopCategories(text) {
const mainBusiness = extractMainBusiness(text);
if (!mainBusiness)
return [];
return uniqueNonEmpty(mainBusiness.split(/[、,/|]/).map((value) => value.trim()));
}
export function isCaptchaState(state) {
const href = cleanText(state.href).toLowerCase();
const title = cleanText(state.title);
const bodyText = cleanMultilineText(state.body_text);
if (href.includes(CAPTCHA_URL_MARKER))
return true;
return CAPTCHA_TEXT_PATTERNS.some((pattern) => title.includes(pattern) || bodyText.includes(pattern));
}
export function isLoginState(state) {
const href = cleanText(state.href).toLowerCase();
const title = cleanText(state.title);
const bodyText = cleanMultilineText(state.body_text);
if (LOGIN_URL_PATTERNS.some((pattern) => href.includes(pattern)))
return true;
return LOGIN_TEXT_PATTERNS.some((pattern) => title.includes(pattern) || bodyText.includes(pattern));
}
export function buildCaptchaHint(action) {
return [
`Open a clean 1688 ${action} page in the shared Chrome profile and finish any slider challenge first.`,
'If you run opencli via CDP, set OPENCLI_CDP_TARGET=1688.com or a more specific 1688 host before retrying.',
].join(' ');
}
export async function readPageState(page) {
const result = await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
body_text: document.body ? document.body.innerText || '' : '',
}))()
`);
return {
href: cleanText(result.href),
title: cleanText(result.title),
body_text: cleanMultilineText(result.body_text),
};
}
export async function gotoAndReadState(page, url, settleMs = 2500, action = 'page') {
try {
await page.goto(url, { settleMs });
await page.wait(1.5);
return readPageState(page);
}
catch (error) {
const message = error instanceof Error ? error.message : String(error);
if (message.includes('Inspected target navigated or closed')
|| message.includes('Cannot find context with specified id')
|| message.includes('Target closed')) {
throw new CommandExecutionError(`1688 ${action} navigation lost the current browser target`, `${buildCaptchaHint(action)} If CDP is attached to a stale or blocked tab, open a fresh 1688 tab and point OPENCLI_CDP_TARGET at that tab.`);
}
throw error;
}
}
export async function ensure1688Session(page) {
const state = await gotoAndReadState(page, HOME_URL, 1500, 'homepage');
assertAuthenticatedState(state, 'homepage');
}
export function assertAuthenticatedState(state, action) {
if (!isCaptchaState(state) && !isLoginState(state))
return;
throw new AuthRequiredError('1688.com', `请先在共享 Chrome 完成 1688 登录/验证,再重试(${action}`);
}
export function assertNotCaptcha(state, action) {
assertAuthenticatedState(state, action);
}
export function toNumber(value) {
if (typeof value === 'number' && Number.isFinite(value)) {
return value;
}
if (typeof value === 'string') {
const normalized = value.replace(/,/g, '').trim();
if (!normalized)
return null;
const parsed = Number.parseFloat(normalized);
return Number.isFinite(parsed) ? parsed : null;
}
return null;
}
export function limitCandidates(values, limit) {
const normalizedLimit = Math.max(1, Math.trunc(limit) || 1);
return values.slice(0, normalizedLimit);
}
export function normalizeMediaUrl(input) {
const raw = cleanText(input);
if (!raw)
return '';
let value = raw
.replace(/^url\((.*)\)$/i, '$1')
.replace(/^['"]|['"]$/g, '')
.replace(/\\u002F/g, '/')
.replace(/&amp;/g, '&')
.trim();
if (!value || value.startsWith('data:') || value.startsWith('blob:'))
return '';
if (value.startsWith('//'))
value = `https:${value}`;
try {
const url = new URL(value);
return url.toString();
}
catch {
return '';
}
}
export function uniqueMediaSources(values) {
const seen = new Set();
const result = [];
for (const value of values) {
const url = normalizeMediaUrl(value.url);
if (!url)
continue;
const key = `${value.type}:${url}`;
if (seen.has(key))
continue;
seen.add(key);
result.push({
...value,
url,
source: cleanText(value.source) || undefined,
});
}
return result;
}
function normalizeNumericText(value) {
return value
.replace(/([¥$€])\s+(?=\d)/g, '$1')
.replace(/(\d)\s*\.\s*(\d)/g, '$1.$2')
.replace(/\s*([~-])\s*/g, '$1')
.trim();
}
function escapeForRegex(value) {
return value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
}
function parse1688Url(input) {
const normalized = cleanText(input);
try {
const url = new URL(normalized);
if (!url.hostname.endsWith('.1688.com') && url.hostname !== '1688.com' && url.hostname !== 'www.1688.com') {
throw new Error('invalid-host');
}
stripTrackingParams(url);
url.hash = '';
return url;
}
catch {
throw new ArgumentError('Invalid 1688 URL', 'Use a URL under 1688.com (for example: https://detail.1688.com/offer/887904326744.html)');
}
}
function parse1688UrlOrNull(input) {
try {
return parse1688Url(input);
}
catch {
return null;
}
}
function normalizeStoreHost(hostname) {
const lower = cleanText(hostname).toLowerCase();
if (!lower.endsWith('.1688.com'))
return null;
const [subdomain] = lower.split('.');
if (!subdomain || STORE_GENERIC_HOSTS.has(subdomain))
return null;
return lower;
}
function stripTrackingParams(url) {
const keys = [...url.searchParams.keys()];
for (const key of keys) {
if (TRACKING_QUERY_KEYS.has(key) || key.toLowerCase().startsWith('utm_')) {
url.searchParams.delete(key);
}
}
}
export const __test__ = {
SEARCH_LIMIT_DEFAULT,
SEARCH_LIMIT_MAX,
parseSearchLimit,
buildSearchUrl,
buildDetailUrl,
resolveStoreUrl,
canonicalizeStoreUrl,
canonicalizeItemUrl,
canonicalizeSellerUrl,
extractOfferId,
extractMemberId,
extractShopId,
parsePriceText,
normalizePriceTiers,
parseMoqText,
extractLocation,
extractAddress,
extractMetric,
extractYearsOnPlatform,
extractMainBusiness,
extractBadges,
guessTopCategories,
isCaptchaState,
isLoginState,
cleanText,
cleanMultilineText,
uniqueNonEmpty,
normalizeMediaUrl,
uniqueMediaSources,
limitCandidates,
};
+57
View File
@@ -0,0 +1,57 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './shared.js';
describe('1688 shared helpers', () => {
it('builds encoded search URLs and validates limit', () => {
expect(__test__.buildSearchUrl('置物架')).toBe('https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=%E7%BD%AE%E7%89%A9%E6%9E%B6');
expect(() => __test__.buildSearchUrl(' ')).toThrowError(/cannot be empty/i);
expect(__test__.parseSearchLimit(3)).toBe(3);
expect(__test__.parseSearchLimit('1000')).toBe(__test__.SEARCH_LIMIT_MAX);
expect(() => __test__.parseSearchLimit('0')).toThrowError(/positive integer/i);
});
it('extracts IDs and canonicalizes urls', () => {
expect(__test__.extractOfferId('887904326744')).toBe('887904326744');
expect(__test__.extractOfferId('https://detail.1688.com/offer/887904326744.html')).toBe('887904326744');
expect(__test__.extractMemberId('https://winport.m.1688.com/page/index.html?memberId=b2b-1641351767')).toBe('b2b-1641351767');
expect(__test__.extractMemberId('b2b-22154705262941f196')).toBe('b2b-22154705262941f196');
expect(__test__.resolveStoreUrl('b2b-22154705262941f196')).toBe('https://winport.m.1688.com/page/index.html?memberId=b2b-22154705262941f196');
expect(__test__.canonicalizeStoreUrl('https://yinuoweierfushi.1688.com/page/index.html?spm=foo')).toBe('https://yinuoweierfushi.1688.com');
expect(__test__.canonicalizeItemUrl('http://detail.m.1688.com/page/index.html?offerId=910933345396&spm=x')).toBe('https://detail.1688.com/offer/910933345396.html');
expect(__test__.canonicalizeSellerUrl('https://yinuoweierfushi.1688.com/page/contactinfo.html?tracelog=1')).toBe('https://yinuoweierfushi.1688.com');
expect(__test__.extractShopId('https://yinuoweierfushi.1688.com/page/index.html')).toBe('yinuoweierfushi');
});
it('parses price ranges and moq text', () => {
expect(__test__.parsePriceText('¥96.00-98.00')).toEqual({
price_text: '¥96.00-98.00',
price_min: 96,
price_max: 98,
currency: 'CNY',
});
expect(__test__.parsePriceText('¥ 14 .28')).toEqual({
price_text: '¥14.28',
price_min: 14.28,
price_max: 14.28,
currency: 'CNY',
});
expect(__test__.parseMoqText('3套起批')).toEqual({
moq_text: '3套起批',
moq_value: 3,
});
expect(__test__.parseMoqText('2~999个')).toEqual({
moq_text: '2~999个',
moq_value: 2,
});
});
it('detects captcha and login states', () => {
expect(__test__.extractLocation('山东青岛 送至 江苏苏州')).toBe('山东青岛');
expect(__test__.isCaptchaState({
href: 'https://s.1688.com/_____tmd_____/punish',
title: '验证码拦截',
body_text: '请拖动下方滑块完成验证',
})).toBe(true);
expect(__test__.isLoginState({
href: 'https://login.taobao.com/member/login.jhtml',
title: '账号登录',
body_text: '请登录后继续',
})).toBe(true);
});
});
-75
View File
@@ -1,75 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './shared.js';
describe('1688 shared helpers', () => {
it('builds encoded search URLs and validates limit', () => {
expect(__test__.buildSearchUrl('置物架')).toBe(
'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=%E7%BD%AE%E7%89%A9%E6%9E%B6',
);
expect(() => __test__.buildSearchUrl(' ')).toThrowError(/cannot be empty/i);
expect(__test__.parseSearchLimit(3)).toBe(3);
expect(__test__.parseSearchLimit('1000')).toBe(__test__.SEARCH_LIMIT_MAX);
expect(() => __test__.parseSearchLimit('0')).toThrowError(/positive integer/i);
});
it('extracts IDs and canonicalizes urls', () => {
expect(__test__.extractOfferId('887904326744')).toBe('887904326744');
expect(__test__.extractOfferId('https://detail.1688.com/offer/887904326744.html')).toBe('887904326744');
expect(__test__.extractMemberId('https://winport.m.1688.com/page/index.html?memberId=b2b-1641351767')).toBe('b2b-1641351767');
expect(__test__.extractMemberId('b2b-22154705262941f196')).toBe('b2b-22154705262941f196');
expect(__test__.resolveStoreUrl('b2b-22154705262941f196')).toBe(
'https://winport.m.1688.com/page/index.html?memberId=b2b-22154705262941f196',
);
expect(__test__.canonicalizeStoreUrl('https://yinuoweierfushi.1688.com/page/index.html?spm=foo')).toBe(
'https://yinuoweierfushi.1688.com',
);
expect(__test__.canonicalizeItemUrl('http://detail.m.1688.com/page/index.html?offerId=910933345396&spm=x')).toBe(
'https://detail.1688.com/offer/910933345396.html',
);
expect(__test__.canonicalizeSellerUrl('https://yinuoweierfushi.1688.com/page/contactinfo.html?tracelog=1')).toBe(
'https://yinuoweierfushi.1688.com',
);
expect(__test__.extractShopId('https://yinuoweierfushi.1688.com/page/index.html')).toBe('yinuoweierfushi');
});
it('parses price ranges and moq text', () => {
expect(__test__.parsePriceText('¥96.00-98.00')).toEqual({
price_text: '¥96.00-98.00',
price_min: 96,
price_max: 98,
currency: 'CNY',
});
expect(__test__.parsePriceText('¥ 14 .28')).toEqual({
price_text: '¥14.28',
price_min: 14.28,
price_max: 14.28,
currency: 'CNY',
});
expect(__test__.parseMoqText('3套起批')).toEqual({
moq_text: '3套起批',
moq_value: 3,
});
expect(__test__.parseMoqText('2~999个')).toEqual({
moq_text: '2~999个',
moq_value: 2,
});
});
it('detects captcha and login states', () => {
expect(__test__.extractLocation('山东青岛 送至 江苏苏州')).toBe('山东青岛');
expect(__test__.isCaptchaState({
href: 'https://s.1688.com/_____tmd_____/punish',
title: '验证码拦截',
body_text: '请拖动下方滑块完成验证',
})).toBe(true);
expect(__test__.isLoginState({
href: 'https://login.taobao.com/member/login.jhtml',
title: '账号登录',
body_text: '请登录后继续',
})).toBe(true);
});
});
-672
View File
@@ -1,672 +0,0 @@
import { ArgumentError, AuthRequiredError, CommandExecutionError } from '@jackwener/opencli/errors';
import type { IPage } from '@jackwener/opencli/types';
export const SITE = '1688';
export const HOME_URL = 'https://www.1688.com/';
export const SEARCH_URL_PREFIX = 'https://s.1688.com/selloffer/offer_search.htm?charset=utf8&keywords=';
export const DETAIL_URL_PREFIX = 'https://detail.1688.com/offer/';
export const STORE_MOBILE_URL_PREFIX = 'https://winport.m.1688.com/page/index.html?memberId=';
export const STRATEGY = 'cookie';
export const SEARCH_LIMIT_DEFAULT = 20;
export const SEARCH_LIMIT_MAX = 100;
const STORE_GENERIC_HOSTS = new Set(['www', 'detail', 's', 'winport', 'work', 'air', 'dj']);
const TRACKING_QUERY_KEYS = new Set([
'spm',
'tracelog',
'clickid',
'source',
'scene',
'from',
'src',
'ns',
'cna',
'pvid',
]);
const CAPTCHA_URL_MARKER = '/_____tmd_____/punish';
const CAPTCHA_TEXT_PATTERNS = [
'请拖动下方滑块完成验证',
'请按住滑块,拖动到最右边',
'通过验证以确保正常访问',
'验证码拦截',
'访问验证',
'滑动验证',
];
const LOGIN_TEXT_PATTERNS = [
'请登录',
'登录后',
'账号登录',
'手机登录',
'立即登录',
'扫码登录',
'请先完成登录',
'请先登录后查看',
];
const LOGIN_URL_PATTERNS = ['/member/login', 'passport', 'login.taobao.com', 'account.1688.com'];
export const FACTORY_BADGE_PATTERNS = [
'源头工厂',
'深度验厂',
'实力工厂',
'工厂档案',
'加工专区',
'验厂报告',
'厂家直销',
'生产厂家',
'工厂直供',
];
export const SERVICE_BADGE_PATTERNS = [
'延期必赔',
'品质保障',
'破损包赔',
'退货包运费',
'晚发必赔',
'7*24小时响应',
'48小时发货',
'72小时发货',
'后天达',
'包邮',
'闪电拿样',
];
const CHINA_LOCATIONS = [
'北京',
'天津',
'上海',
'重庆',
'河北',
'山西',
'辽宁',
'吉林',
'黑龙江',
'江苏',
'浙江',
'安徽',
'福建',
'江西',
'山东',
'河南',
'湖北',
'湖南',
'广东',
'海南',
'四川',
'贵州',
'云南',
'陕西',
'甘肃',
'青海',
'台湾',
'内蒙古',
'广西',
'西藏',
'宁夏',
'新疆',
'香港',
'澳门',
];
export interface ProvenanceFields {
source_url: string;
fetched_at: string;
strategy: string;
}
export interface PageState {
href: string;
title: string;
body_text: string;
}
export interface PriceRange {
price_text: string;
price_min: number | null;
price_max: number | null;
currency: string | null;
}
export interface MoqValue {
moq_text: string;
moq_value: number | null;
}
export interface PriceTier {
quantity_text: string;
quantity_min: number | null;
price_text: string;
price: number | null;
currency: string | null;
}
export interface SearchCandidate {
item_url: string;
title: string;
container_text: string;
seller_name: string | null;
seller_url: string | null;
}
export interface MediaSource {
type: 'image' | 'video';
group: 'main' | 'sku' | 'detail' | 'video' | 'unknown';
url: string;
source?: string;
}
export function cleanText(value: unknown): string {
return typeof value === 'string'
? value.replace(/\u00a0/g, ' ').replace(/\s+/g, ' ').trim()
: '';
}
export function cleanMultilineText(value: unknown): string {
return typeof value === 'string'
? value
.replace(/\u00a0/g, ' ')
.split('\n')
.map((line) => line.replace(/\s+/g, ' ').trim())
.filter(Boolean)
.join('\n')
: '';
}
export function uniqueNonEmpty(values: Array<string | null | undefined>): string[] {
return [...new Set(values.map((value) => cleanText(value)).filter(Boolean))];
}
export function parseSearchLimit(input: unknown): number {
const parsed = Number.parseInt(String(input ?? SEARCH_LIMIT_DEFAULT), 10);
if (!Number.isFinite(parsed) || parsed < 1) {
throw new ArgumentError(
'1688 search --limit must be a positive integer',
'Example: opencli 1688 search "桌面置物架" --limit 20',
);
}
return Math.min(SEARCH_LIMIT_MAX, parsed);
}
export function buildSearchUrl(query: string): string {
const normalized = cleanText(query);
if (!normalized) {
throw new ArgumentError(
'1688 search query cannot be empty',
'Example: opencli 1688 search "桌面置物架" --limit 20',
);
}
return `${SEARCH_URL_PREFIX}${encodeURIComponent(normalized)}`;
}
export function buildDetailUrl(input: string): string {
const offerId = extractOfferId(input);
if (!offerId) {
throw new ArgumentError(
'1688 item expects an offer URL or offer ID',
'Example: opencli 1688 item 887904326744',
);
}
return `${DETAIL_URL_PREFIX}${offerId}.html`;
}
export function resolveStoreUrl(input: string): string {
const normalized = cleanText(input);
if (!normalized) {
throw new ArgumentError(
'1688 store expects a store URL or member ID',
'Example: opencli 1688 store https://yinuoweierfushi.1688.com/',
);
}
const memberId = extractMemberId(normalized);
if (memberId) {
return `${STORE_MOBILE_URL_PREFIX}${memberId}`;
}
if (/^https?:\/\//i.test(normalized)) {
return canonicalizeStoreUrl(normalized);
}
if (normalized.endsWith('.1688.com')) {
return canonicalizeStoreUrl(`https://${normalized}`);
}
if (/^[a-z0-9-]+$/i.test(normalized)) {
return canonicalizeStoreUrl(`https://${normalized}.1688.com`);
}
throw new ArgumentError(
'1688 store expects a store URL or member ID',
'Example: opencli 1688 store b2b-22154705262941f196',
);
}
export function canonicalizeStoreUrl(input: string): string {
const url = parse1688Url(input);
const memberId = extractMemberId(url.toString());
if (memberId) {
return `${STORE_MOBILE_URL_PREFIX}${memberId}`;
}
const host = normalizeStoreHost(url.hostname);
if (!host) {
throw new ArgumentError(
'Invalid 1688 store URL',
'Example: opencli 1688 store https://yinuoweierfushi.1688.com/',
);
}
return `https://${host}`;
}
export function canonicalizeItemUrl(input: string): string | null {
const offerId = extractOfferId(input);
if (offerId) {
return `${DETAIL_URL_PREFIX}${offerId}.html`;
}
const url = parse1688UrlOrNull(input);
if (!url) return null;
stripTrackingParams(url);
url.hash = '';
return url.toString();
}
export function canonicalizeSellerUrl(input: string): string | null {
const memberId = extractMemberId(input);
if (memberId) {
return `${STORE_MOBILE_URL_PREFIX}${memberId}`;
}
const url = parse1688UrlOrNull(input);
if (!url) return null;
const host = normalizeStoreHost(url.hostname);
if (!host) return null;
return `https://${host}`;
}
export function extractOfferId(input: string): string | null {
const normalized = cleanText(input);
if (!normalized) return null;
const directId = normalized.match(/^\d{6,}$/)?.[0];
if (directId) return directId;
const detailMatch = normalized.match(/\/offer\/(\d{6,})\.html/i);
if (detailMatch) return detailMatch[1];
const queryMatch = normalized.match(/[?&]offerId=(\d{6,})/i);
if (queryMatch) return queryMatch[1];
return null;
}
export function extractMemberId(input: string): string | null {
const normalized = cleanText(input);
if (!normalized) return null;
const direct = normalized.match(/\bb2b-[a-z0-9]+\b/i)?.[0];
if (direct) return direct;
const queryMatch = normalized.match(/[?&]memberId=(b2b-[a-z0-9]+)/i);
if (queryMatch) return queryMatch[1];
const mobileMatch = normalized.match(/\/winport\/(b2b-[a-z0-9]+)\.html/i);
if (mobileMatch) return mobileMatch[1];
return null;
}
export function extractShopId(input: string): string | null {
const normalized = cleanText(input);
if (!normalized) return null;
try {
const url = new URL(/^https?:\/\//i.test(normalized) ? normalized : `https://${normalized}`);
const host = normalizeStoreHost(url.hostname);
if (!host) return null;
return host.split('.')[0] ?? null;
} catch {
return /^[a-z0-9-]+$/i.test(normalized) ? normalized : null;
}
}
export function buildProvenance(sourceUrl: string): ProvenanceFields {
return {
source_url: sourceUrl,
fetched_at: new Date().toISOString(),
strategy: STRATEGY,
};
}
export function parsePriceText(text: string): PriceRange {
const normalized = normalizeNumericText(cleanText(text));
const matches = normalized.match(/\d+(?:,\d{3})*(?:\.\d+)?/g) ?? [];
const values = matches
.map((value) => Number.parseFloat(value.replace(/,/g, '')))
.filter((value) => Number.isFinite(value));
if (values.length === 0) {
return {
price_text: normalized,
price_min: null,
price_max: null,
currency: null,
};
}
return {
price_text: normalized,
price_min: values[0] ?? null,
price_max: values[values.length - 1] ?? values[0] ?? null,
currency: normalized.includes('¥') || normalized.includes('元') ? 'CNY' : null,
};
}
export function normalizePriceTiers(
rawTiers: Array<{ beginAmount?: unknown; price?: unknown }>,
unit: string | null,
): PriceTier[] {
return rawTiers
.map((tier) => {
const quantityMin = toNumber(tier.beginAmount);
const priceText = cleanText(tier.price);
const price = toNumber(tier.price);
return {
quantity_text: quantityMin !== null ? `${quantityMin}${unit ?? ''}` : '',
quantity_min: quantityMin,
price_text: priceText,
price,
currency: priceText ? 'CNY' : null,
};
})
.filter((tier) => tier.price_text);
}
export function parseMoqText(text: string): MoqValue {
const normalized = normalizeNumericText(cleanText(text));
const match = normalized.match(/(\d+(?:\.\d+)?)\s*(件|个|套|箱|包|双|台|把|只|pcs|piece|pieces)?\s*起批/i)
?? normalized.match(/≥\s*(\d+(?:\.\d+)?)/);
const rangeMatch = normalized.match(
/(\d+(?:\.\d+)?)\s*(?:~|-|至|到)\s*\d+(?:\.\d+)?\s*(件|个|套|箱|包|双|台|把|只|pcs|piece|pieces)/i,
);
if (!match && !rangeMatch) {
return {
moq_text: normalized,
moq_value: null,
};
}
return {
moq_text: normalized,
moq_value: Number.parseFloat((match ?? rangeMatch)![1]),
};
}
export function extractLocation(text: string): string | null {
const normalized = cleanMultilineText(text);
const primaryRegion = normalized.split(/送至|发往/)[0] ?? normalized;
const lines = primaryRegion.split('\n');
for (const line of lines) {
const compact = cleanText(line);
if (!compact || compact.length > 16) continue;
if (CHINA_LOCATIONS.some((location) => compact.startsWith(location))) {
return compact;
}
}
const locationPattern = new RegExp(`(${CHINA_LOCATIONS.join('|')})[\\u4e00-\\u9fa5]{0,8}`);
return primaryRegion.match(locationPattern)?.[0] ?? null;
}
export function extractAddress(text: string): string | null {
const normalized = cleanMultilineText(text);
const lineMatch = normalized.match(/地址[:]\s*([^\n]+)/);
if (lineMatch) return cleanText(lineMatch[1]);
return normalized
.split('\n')
.map((line) => cleanText(line))
.find((line) => line.includes('省') || line.includes('市') || line.includes('区') || line.includes('县'))
?? null;
}
export function extractMetric(text: string, label: string): string | null {
const normalized = cleanMultilineText(text);
const direct = normalized.match(new RegExp(`(?:^|\\n)\\s*${escapeForRegex(label)}[:]?\\s*([^\\n]+)`));
if (direct) return cleanText(direct[1]);
const lineBased = normalized.match(new RegExp(`(?:^|\\n)\\s*${escapeForRegex(label)}\\n([^\\n]+)`));
return lineBased ? cleanText(lineBased[1]) : null;
}
export function extractYearsOnPlatform(text: string): string | null {
return text.match(/入驻\d+年/)?.[0] ?? null;
}
export function extractMainBusiness(text: string): string | null {
const value = extractMetric(text, '主营');
return value ? value.replace(/^/, '').trim() : null;
}
export function extractBadges(text: string, candidates: string[]): string[] {
return uniqueNonEmpty(candidates.filter((candidate) => cleanMultilineText(text).includes(candidate)));
}
export function guessTopCategories(text: string): string[] {
const mainBusiness = extractMainBusiness(text);
if (!mainBusiness) return [];
return uniqueNonEmpty(mainBusiness.split(/[、,/|]/).map((value) => value.trim()));
}
export function isCaptchaState(state: Partial<PageState>): boolean {
const href = cleanText(state.href).toLowerCase();
const title = cleanText(state.title);
const bodyText = cleanMultilineText(state.body_text);
if (href.includes(CAPTCHA_URL_MARKER)) return true;
return CAPTCHA_TEXT_PATTERNS.some((pattern) => title.includes(pattern) || bodyText.includes(pattern));
}
export function isLoginState(state: Partial<PageState>): boolean {
const href = cleanText(state.href).toLowerCase();
const title = cleanText(state.title);
const bodyText = cleanMultilineText(state.body_text);
if (LOGIN_URL_PATTERNS.some((pattern) => href.includes(pattern))) return true;
return LOGIN_TEXT_PATTERNS.some((pattern) => title.includes(pattern) || bodyText.includes(pattern));
}
export function buildCaptchaHint(action: string): string {
return [
`Open a clean 1688 ${action} page in the shared Chrome profile and finish any slider challenge first.`,
'If you run opencli via CDP, set OPENCLI_CDP_TARGET=1688.com or a more specific 1688 host before retrying.',
].join(' ');
}
export async function readPageState(page: IPage): Promise<PageState> {
const result = await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
body_text: document.body ? document.body.innerText || '' : '',
}))()
`) as Partial<PageState>;
return {
href: cleanText(result.href),
title: cleanText(result.title),
body_text: cleanMultilineText(result.body_text),
};
}
export async function gotoAndReadState(
page: IPage,
url: string,
settleMs: number = 2500,
action: string = 'page',
): Promise<PageState> {
try {
await page.goto(url, { settleMs });
await page.wait(1.5);
return readPageState(page);
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
if (
message.includes('Inspected target navigated or closed')
|| message.includes('Cannot find context with specified id')
|| message.includes('Target closed')
) {
throw new CommandExecutionError(
`1688 ${action} navigation lost the current browser target`,
`${buildCaptchaHint(action)} If CDP is attached to a stale or blocked tab, open a fresh 1688 tab and point OPENCLI_CDP_TARGET at that tab.`,
);
}
throw error;
}
}
export async function ensure1688Session(page: IPage): Promise<void> {
const state = await gotoAndReadState(page, HOME_URL, 1500, 'homepage');
assertAuthenticatedState(state, 'homepage');
}
export function assertAuthenticatedState(state: PageState, action: string): void {
if (!isCaptchaState(state) && !isLoginState(state)) return;
throw new AuthRequiredError('1688.com', `请先在共享 Chrome 完成 1688 登录/验证,再重试(${action}`);
}
export function assertNotCaptcha(state: PageState, action: string): void {
assertAuthenticatedState(state, action);
}
export function toNumber(value: unknown): number | null {
if (typeof value === 'number' && Number.isFinite(value)) {
return value;
}
if (typeof value === 'string') {
const normalized = value.replace(/,/g, '').trim();
if (!normalized) return null;
const parsed = Number.parseFloat(normalized);
return Number.isFinite(parsed) ? parsed : null;
}
return null;
}
export function limitCandidates<T>(values: T[], limit: number): T[] {
const normalizedLimit = Math.max(1, Math.trunc(limit) || 1);
return values.slice(0, normalizedLimit);
}
export function normalizeMediaUrl(input: unknown): string {
const raw = cleanText(input);
if (!raw) return '';
let value = raw
.replace(/^url\((.*)\)$/i, '$1')
.replace(/^['"]|['"]$/g, '')
.replace(/\\u002F/g, '/')
.replace(/&amp;/g, '&')
.trim();
if (!value || value.startsWith('data:') || value.startsWith('blob:')) return '';
if (value.startsWith('//')) value = `https:${value}`;
try {
const url = new URL(value);
return url.toString();
} catch {
return '';
}
}
export function uniqueMediaSources(values: MediaSource[]): MediaSource[] {
const seen = new Set<string>();
const result: MediaSource[] = [];
for (const value of values) {
const url = normalizeMediaUrl(value.url);
if (!url) continue;
const key = `${value.type}:${url}`;
if (seen.has(key)) continue;
seen.add(key);
result.push({
...value,
url,
source: cleanText(value.source) || undefined,
});
}
return result;
}
function normalizeNumericText(value: string): string {
return value
.replace(/([¥$€])\s+(?=\d)/g, '$1')
.replace(/(\d)\s*\.\s*(\d)/g, '$1.$2')
.replace(/\s*([~-])\s*/g, '$1')
.trim();
}
function escapeForRegex(value: string): string {
return value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
}
function parse1688Url(input: string): URL {
const normalized = cleanText(input);
try {
const url = new URL(normalized);
if (!url.hostname.endsWith('.1688.com') && url.hostname !== '1688.com' && url.hostname !== 'www.1688.com') {
throw new Error('invalid-host');
}
stripTrackingParams(url);
url.hash = '';
return url;
} catch {
throw new ArgumentError(
'Invalid 1688 URL',
'Use a URL under 1688.com (for example: https://detail.1688.com/offer/887904326744.html)',
);
}
}
function parse1688UrlOrNull(input: string): URL | null {
try {
return parse1688Url(input);
} catch {
return null;
}
}
function normalizeStoreHost(hostname: string): string | null {
const lower = cleanText(hostname).toLowerCase();
if (!lower.endsWith('.1688.com')) return null;
const [subdomain] = lower.split('.');
if (!subdomain || STORE_GENERIC_HOSTS.has(subdomain)) return null;
return lower;
}
function stripTrackingParams(url: URL): void {
const keys = [...url.searchParams.keys()];
for (const key of keys) {
if (TRACKING_QUERY_KEYS.has(key) || key.toLowerCase().startsWith('utm_')) {
url.searchParams.delete(key);
}
}
}
export const __test__ = {
SEARCH_LIMIT_DEFAULT,
SEARCH_LIMIT_MAX,
parseSearchLimit,
buildSearchUrl,
buildDetailUrl,
resolveStoreUrl,
canonicalizeStoreUrl,
canonicalizeItemUrl,
canonicalizeSellerUrl,
extractOfferId,
extractMemberId,
extractShopId,
parsePriceText,
normalizePriceTiers,
parseMoqText,
extractLocation,
extractAddress,
extractMetric,
extractYearsOnPlatform,
extractMainBusiness,
extractBadges,
guessTopCategories,
isCaptchaState,
isLoginState,
cleanText,
cleanMultilineText,
uniqueNonEmpty,
normalizeMediaUrl,
uniqueMediaSources,
limitCandidates,
};
+226
View File
@@ -0,0 +1,226 @@
import { CommandExecutionError, EmptyResultError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { FACTORY_BADGE_PATTERNS, SERVICE_BADGE_PATTERNS, assertAuthenticatedState, buildDetailUrl, buildProvenance, canonicalizeSellerUrl, canonicalizeStoreUrl, cleanMultilineText, cleanText, extractAddress, extractBadges, extractMemberId, extractMetric, extractOfferId, extractShopId, extractYearsOnPlatform, gotoAndReadState, guessTopCategories, resolveStoreUrl, uniqueNonEmpty, } from './shared.js';
function normalizeStorePayload(input) {
const storePayload = input.storePayload;
const contactPayload = input.contactPayload;
const seed = input.seed;
const contactText = cleanMultilineText(contactPayload?.bodyText);
const storeText = cleanMultilineText(storePayload?.bodyText);
const seedText = cleanMultilineText(seed?.bodyText);
const combinedText = [contactText, storeText, seedText].filter(Boolean).join('\n');
const sellerUrlRaw = cleanText(seed?.seller?.winportUrl
?? seed?.seller?.sellerWinportUrlMap?.defaultUrl
?? storePayload?.href
?? input.resolvedUrl);
const storeUrl = safeCanonicalStoreUrl(sellerUrlRaw || input.resolvedUrl) ?? input.resolvedUrl;
const sellerUrl = canonicalizeSellerUrl(sellerUrlRaw) ?? storeUrl;
const companyUrl = pickCompanyUrl(contactPayload?.href, storeUrl);
const memberId = cleanText(seed?.seller?.memberId)
|| input.explicitMemberId
|| extractMemberId(input.resolvedUrl)
|| extractMemberId(storePayload?.href ?? '')
|| null;
const shopId = extractShopId(sellerUrl) ?? extractShopId(storeUrl);
const companyName = cleanText(seed?.seller?.companyName)
|| firstNamedLine(contactText)
|| firstNamedLine(storeText)
|| null;
const serviceBadges = uniqueNonEmpty([
...extractBadges(combinedText, SERVICE_BADGE_PATTERNS),
...((seed?.services ?? []).map((service) => cleanText(service.serviceName))),
]);
const factoryBadges = extractBadges(combinedText, FACTORY_BADGE_PATTERNS);
return {
member_id: memberId,
shop_id: shopId,
store_name: companyName,
store_url: storeUrl,
company_name: companyName,
company_url: companyUrl,
business_model_text: firstMetric(combinedText, ['经营模式', '生产加工', '主营产品']),
years_on_platform_text: extractYearsOnPlatform(combinedText),
location: extractAddress(contactText) ?? extractAddress(storeText),
staff_size_text: firstMetric(combinedText, ['员工人数', '员工总数']),
factory_badges: factoryBadges,
service_badges: serviceBadges,
response_rate_text: firstMetric(combinedText, ['响应率', '回复率', '响应速度']),
return_rate_text: extractReturnRate(combinedText),
top_categories: guessTopCategories(combinedText),
phone_text: extractMetric(contactText, '电话'),
mobile_text: extractMetric(contactText, '手机'),
...buildProvenance(cleanText(contactPayload?.href) || cleanText(storePayload?.href) || input.resolvedUrl),
};
}
function safeCanonicalStoreUrl(url) {
try {
return canonicalizeStoreUrl(url);
}
catch {
return null;
}
}
function pickCompanyUrl(contactHref, storeUrl) {
const fromPage = cleanText(contactHref);
if (fromPage) {
const normalized = buildContactUrl(fromPage);
if (normalized)
return normalized;
}
return buildContactUrl(storeUrl);
}
function buildContactUrl(storeUrl) {
try {
const parsed = new URL(storeUrl);
if (!parsed.hostname.endsWith('.1688.com'))
return null;
return `${parsed.protocol}//${parsed.hostname}/page/contactinfo.html`;
}
catch {
return null;
}
}
function firstNamedLine(text) {
return text
.split('\n')
.map((line) => cleanText(line))
.find((line) => line.includes('有限公司') || line.includes('商行') || line.includes('工厂'))
?? null;
}
function firstMetric(text, labels) {
for (const label of labels) {
const value = extractMetric(text, label);
if (value)
return value;
}
return null;
}
function extractReturnRate(text) {
const inline = text.match(/回头率\s*([0-9.]+%)/);
if (inline)
return cleanText(inline[0]);
const multiline = text.match(/回头率\s*\n\s*([0-9.]+%)/);
if (!multiline)
return null;
return `回头率${cleanText(multiline[1])}`;
}
function firstOfferId(links) {
for (const link of links) {
const offerId = extractOfferId(link);
if (offerId)
return offerId;
}
return null;
}
function firstContactUrl(links) {
for (const link of links) {
const url = buildContactUrl(link);
if (url)
return url;
}
return null;
}
async function readStorePayload(page, url, action) {
const state = await gotoAndReadState(page, url, 2500, action);
assertAuthenticatedState(state, action);
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
bodyText: document.body ? document.body.innerText || '' : '',
offerLinks: Array.from(document.querySelectorAll('a[href*="detail.1688.com/offer/"], a[href*="offerId="]'))
.map((anchor) => anchor.href)
.filter(Boolean),
contactLinks: Array.from(document.querySelectorAll('a[href*="contactinfo"]'))
.map((anchor) => anchor.href)
.filter(Boolean),
}))()
`);
}
async function readItemSeed(page, offerId) {
const itemUrl = buildDetailUrl(offerId);
const state = await gotoAndReadState(page, itemUrl, 2500, 'store seed item');
assertAuthenticatedState(state, 'store seed item');
const seed = await page.evaluate(`
(() => {
const model = window.context?.result?.global?.globalData?.model ?? null;
const toJson = (value) => JSON.parse(JSON.stringify(value ?? null));
return {
href: window.location.href,
bodyText: document.body ? document.body.innerText || '' : '',
seller: toJson(model?.sellerModel),
services: toJson(model?.shippingServices?.fields?.buyerProtectionModel ?? []),
};
})()
`);
const hasSellerContext = !!cleanText(seed?.seller?.memberId) || !!cleanText(seed?.seller?.winportUrl);
if (!hasSellerContext) {
throw new CommandExecutionError('1688 store seed item did not expose seller context', '当前 tab 非商品详情上下文,请切到 detail.1688.com 商品页并重试');
}
return seed;
}
function hasAnyEvidence(storePayload, contactPayload, seed) {
return !!cleanText(storePayload?.bodyText)
|| !!cleanText(contactPayload?.bodyText)
|| !!cleanText(seed?.bodyText);
}
cli({
site: '1688',
name: 'store',
description: '1688 店铺/供应商公开信息(联系方式、主营、入驻年限、公开服务信号)',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 店铺 URL 或 member ID(如 b2b-22154705262941f196',
},
],
columns: ['store_name', 'years_on_platform_text', 'location', 'return_rate_text'],
func: async (page, kwargs) => {
const rawInput = String(kwargs.input ?? '');
const resolvedUrl = resolveStoreUrl(rawInput);
const explicitMemberId = extractMemberId(rawInput);
const storePayload = await readStorePayload(page, resolvedUrl, 'store');
const contactUrl = firstContactUrl(storePayload.contactLinks ?? []) || buildContactUrl(storePayload.href || resolvedUrl);
const contactPayload = contactUrl ? await readStorePayload(page, contactUrl, 'store contact') : null;
const offerId = extractOfferId(rawInput)
|| firstOfferId(storePayload.offerLinks ?? [])
|| firstOfferId(contactPayload?.offerLinks ?? []);
let seed = null;
if (offerId) {
try {
seed = await readItemSeed(page, offerId);
}
catch (error) {
if (!(error instanceof CommandExecutionError))
throw error;
}
}
if (!hasAnyEvidence(storePayload, contactPayload, seed)) {
throw new EmptyResultError('1688 store', 'Store page is reachable but no visible fields were extracted. Open the store page in Chrome and retry.');
}
return [
normalizeStorePayload({
resolvedUrl,
storePayload,
contactPayload,
seed,
explicitMemberId,
}),
];
},
});
export const __test__ = {
normalizeStorePayload,
safeCanonicalStoreUrl,
buildContactUrl,
firstNamedLine,
firstMetric,
extractReturnRate,
firstOfferId,
firstContactUrl,
};
+62
View File
@@ -0,0 +1,62 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './store.js';
describe('1688 store normalization', () => {
it('merges store contact text with seller seed data', () => {
const result = __test__.normalizeStorePayload({
resolvedUrl: 'https://yinuoweierfushi.1688.com/?offerId=887904326744',
explicitMemberId: null,
storePayload: {
href: 'https://yinuoweierfushi.1688.com/page/index.html',
bodyText: `
青岛沁澜衣品服装有限公司
联系方式
地址:山东省青岛市即墨区环秀街道办事处湘江二路97号甲
`,
offerLinks: ['https://detail.1688.com/offer/887904326744.html'],
},
contactPayload: {
href: 'https://yinuoweierfushi.1688.com/page/contactinfo.html',
bodyText: `
青岛沁澜衣品服装有限公司
电话:86 0532 86655366
手机:15963238678
地址:山东省青岛市即墨区环秀街道办事处湘江二路97号甲
`,
},
seed: {
bodyText: `
入驻13年
主营:大码女装
店铺回头率
87%
延期必赔
品质保障
`,
seller: {
companyName: '青岛沁澜衣品服装有限公司',
memberId: 'b2b-1641351767',
winportUrl: 'https://yinuoweierfushi.1688.com/page/index.html?spm=abc',
},
services: [{ serviceName: '延期必赔' }, { serviceName: '品质保障' }],
},
});
expect(result.member_id).toBe('b2b-1641351767');
expect(result.store_url).toBe('https://yinuoweierfushi.1688.com');
expect(result.company_url).toBe('https://yinuoweierfushi.1688.com/page/contactinfo.html');
expect(result.years_on_platform_text).toBe('入驻13年');
expect(result.location).toBe('山东省青岛市即墨区环秀街道办事处湘江二路97号甲');
expect(result.return_rate_text).toContain('87%');
expect(result.top_categories).toEqual(['大码女装']);
expect(result.service_badges).toEqual(['延期必赔', '品质保障']);
});
it('builds contact urls and extracts offer ids', () => {
expect(__test__.safeCanonicalStoreUrl('https://yinuoweierfushi.1688.com/page/index.html?spm=foo')).toBe('https://yinuoweierfushi.1688.com');
expect(__test__.buildContactUrl('https://yinuoweierfushi.1688.com')).toBe('https://yinuoweierfushi.1688.com/page/contactinfo.html');
expect(__test__.firstOfferId([
'https://detail.1688.com/offer/887904326744.html',
])).toBe('887904326744');
expect(__test__.firstContactUrl([
'https://yinuoweierfushi.1688.com/page/contactinfo.html?spm=1',
])).toBe('https://yinuoweierfushi.1688.com/page/contactinfo.html');
});
});
-69
View File
@@ -1,69 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './store.js';
describe('1688 store normalization', () => {
it('merges store contact text with seller seed data', () => {
const result = __test__.normalizeStorePayload({
resolvedUrl: 'https://yinuoweierfushi.1688.com/?offerId=887904326744',
explicitMemberId: null,
storePayload: {
href: 'https://yinuoweierfushi.1688.com/page/index.html',
bodyText: `
青岛沁澜衣品服装有限公司
联系方式
地址:山东省青岛市即墨区环秀街道办事处湘江二路97号甲
`,
offerLinks: ['https://detail.1688.com/offer/887904326744.html'],
},
contactPayload: {
href: 'https://yinuoweierfushi.1688.com/page/contactinfo.html',
bodyText: `
青岛沁澜衣品服装有限公司
电话:86 0532 86655366
手机:15963238678
地址:山东省青岛市即墨区环秀街道办事处湘江二路97号甲
`,
},
seed: {
bodyText: `
入驻13年
主营:大码女装
店铺回头率
87%
延期必赔
品质保障
`,
seller: {
companyName: '青岛沁澜衣品服装有限公司',
memberId: 'b2b-1641351767',
winportUrl: 'https://yinuoweierfushi.1688.com/page/index.html?spm=abc',
},
services: [{ serviceName: '延期必赔' }, { serviceName: '品质保障' }],
},
});
expect(result.member_id).toBe('b2b-1641351767');
expect(result.store_url).toBe('https://yinuoweierfushi.1688.com');
expect(result.company_url).toBe('https://yinuoweierfushi.1688.com/page/contactinfo.html');
expect(result.years_on_platform_text).toBe('入驻13年');
expect(result.location).toBe('山东省青岛市即墨区环秀街道办事处湘江二路97号甲');
expect(result.return_rate_text).toContain('87%');
expect(result.top_categories).toEqual(['大码女装']);
expect(result.service_badges).toEqual(['延期必赔', '品质保障']);
});
it('builds contact urls and extracts offer ids', () => {
expect(__test__.safeCanonicalStoreUrl('https://yinuoweierfushi.1688.com/page/index.html?spm=foo')).toBe(
'https://yinuoweierfushi.1688.com',
);
expect(__test__.buildContactUrl('https://yinuoweierfushi.1688.com')).toBe(
'https://yinuoweierfushi.1688.com/page/contactinfo.html',
);
expect(__test__.firstOfferId([
'https://detail.1688.com/offer/887904326744.html',
])).toBe('887904326744');
expect(__test__.firstContactUrl([
'https://yinuoweierfushi.1688.com/page/contactinfo.html?spm=1',
])).toBe('https://yinuoweierfushi.1688.com/page/contactinfo.html');
});
});
-300
View File
@@ -1,300 +0,0 @@
import { CommandExecutionError, EmptyResultError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
FACTORY_BADGE_PATTERNS,
SERVICE_BADGE_PATTERNS,
assertAuthenticatedState,
buildDetailUrl,
buildProvenance,
canonicalizeSellerUrl,
canonicalizeStoreUrl,
cleanMultilineText,
cleanText,
extractAddress,
extractBadges,
extractMemberId,
extractMetric,
extractOfferId,
extractShopId,
extractYearsOnPlatform,
gotoAndReadState,
guessTopCategories,
resolveStoreUrl,
uniqueNonEmpty,
} from './shared.js';
interface StoreBrowserPayload {
href?: string;
title?: string;
bodyText?: string;
offerLinks?: string[];
contactLinks?: string[];
}
interface StoreItemSeed {
href?: string;
bodyText?: string;
seller?: {
companyName?: string;
memberId?: string;
winportUrl?: string;
sellerWinportUrlMap?: Record<string, string>;
};
services?: Array<{ serviceName?: string }>;
}
function normalizeStorePayload(input: {
resolvedUrl: string;
storePayload: StoreBrowserPayload | null;
contactPayload: StoreBrowserPayload | null;
seed: StoreItemSeed | null;
explicitMemberId: string | null;
}): Record<string, unknown> {
const storePayload = input.storePayload;
const contactPayload = input.contactPayload;
const seed = input.seed;
const contactText = cleanMultilineText(contactPayload?.bodyText);
const storeText = cleanMultilineText(storePayload?.bodyText);
const seedText = cleanMultilineText(seed?.bodyText);
const combinedText = [contactText, storeText, seedText].filter(Boolean).join('\n');
const sellerUrlRaw = cleanText(
seed?.seller?.winportUrl
?? seed?.seller?.sellerWinportUrlMap?.defaultUrl
?? storePayload?.href
?? input.resolvedUrl,
);
const storeUrl = safeCanonicalStoreUrl(sellerUrlRaw || input.resolvedUrl) ?? input.resolvedUrl;
const sellerUrl = canonicalizeSellerUrl(sellerUrlRaw) ?? storeUrl;
const companyUrl = pickCompanyUrl(contactPayload?.href, storeUrl);
const memberId = cleanText(seed?.seller?.memberId)
|| input.explicitMemberId
|| extractMemberId(input.resolvedUrl)
|| extractMemberId(storePayload?.href ?? '')
|| null;
const shopId = extractShopId(sellerUrl) ?? extractShopId(storeUrl);
const companyName = cleanText(seed?.seller?.companyName)
|| firstNamedLine(contactText)
|| firstNamedLine(storeText)
|| null;
const serviceBadges = uniqueNonEmpty([
...extractBadges(combinedText, SERVICE_BADGE_PATTERNS),
...((seed?.services ?? []).map((service) => cleanText(service.serviceName))),
]);
const factoryBadges = extractBadges(combinedText, FACTORY_BADGE_PATTERNS);
return {
member_id: memberId,
shop_id: shopId,
store_name: companyName,
store_url: storeUrl,
company_name: companyName,
company_url: companyUrl,
business_model_text: firstMetric(combinedText, ['经营模式', '生产加工', '主营产品']),
years_on_platform_text: extractYearsOnPlatform(combinedText),
location: extractAddress(contactText) ?? extractAddress(storeText),
staff_size_text: firstMetric(combinedText, ['员工人数', '员工总数']),
factory_badges: factoryBadges,
service_badges: serviceBadges,
response_rate_text: firstMetric(combinedText, ['响应率', '回复率', '响应速度']),
return_rate_text: extractReturnRate(combinedText),
top_categories: guessTopCategories(combinedText),
phone_text: extractMetric(contactText, '电话'),
mobile_text: extractMetric(contactText, '手机'),
...buildProvenance(cleanText(contactPayload?.href) || cleanText(storePayload?.href) || input.resolvedUrl),
};
}
function safeCanonicalStoreUrl(url: string): string | null {
try {
return canonicalizeStoreUrl(url);
} catch {
return null;
}
}
function pickCompanyUrl(contactHref: string | undefined, storeUrl: string): string | null {
const fromPage = cleanText(contactHref);
if (fromPage) {
const normalized = buildContactUrl(fromPage);
if (normalized) return normalized;
}
return buildContactUrl(storeUrl);
}
function buildContactUrl(storeUrl: string): string | null {
try {
const parsed = new URL(storeUrl);
if (!parsed.hostname.endsWith('.1688.com')) return null;
return `${parsed.protocol}//${parsed.hostname}/page/contactinfo.html`;
} catch {
return null;
}
}
function firstNamedLine(text: string): string | null {
return text
.split('\n')
.map((line) => cleanText(line))
.find((line) => line.includes('有限公司') || line.includes('商行') || line.includes('工厂'))
?? null;
}
function firstMetric(text: string, labels: string[]): string | null {
for (const label of labels) {
const value = extractMetric(text, label);
if (value) return value;
}
return null;
}
function extractReturnRate(text: string): string | null {
const inline = text.match(/回头率\s*([0-9.]+%)/);
if (inline) return cleanText(inline[0]);
const multiline = text.match(/回头率\s*\n\s*([0-9.]+%)/);
if (!multiline) return null;
return `回头率${cleanText(multiline[1])}`;
}
function firstOfferId(links: string[]): string | null {
for (const link of links) {
const offerId = extractOfferId(link);
if (offerId) return offerId;
}
return null;
}
function firstContactUrl(links: string[]): string | null {
for (const link of links) {
const url = buildContactUrl(link);
if (url) return url;
}
return null;
}
async function readStorePayload(page: IPage, url: string, action: string): Promise<StoreBrowserPayload> {
const state = await gotoAndReadState(page, url, 2500, action);
assertAuthenticatedState(state, action);
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
bodyText: document.body ? document.body.innerText || '' : '',
offerLinks: Array.from(document.querySelectorAll('a[href*="detail.1688.com/offer/"], a[href*="offerId="]'))
.map((anchor) => anchor.href)
.filter(Boolean),
contactLinks: Array.from(document.querySelectorAll('a[href*="contactinfo"]'))
.map((anchor) => anchor.href)
.filter(Boolean),
}))()
`) as StoreBrowserPayload;
}
async function readItemSeed(page: IPage, offerId: string): Promise<StoreItemSeed> {
const itemUrl = buildDetailUrl(offerId);
const state = await gotoAndReadState(page, itemUrl, 2500, 'store seed item');
assertAuthenticatedState(state, 'store seed item');
const seed = await page.evaluate(`
(() => {
const model = window.context?.result?.global?.globalData?.model ?? null;
const toJson = (value) => JSON.parse(JSON.stringify(value ?? null));
return {
href: window.location.href,
bodyText: document.body ? document.body.innerText || '' : '',
seller: toJson(model?.sellerModel),
services: toJson(model?.shippingServices?.fields?.buyerProtectionModel ?? []),
};
})()
`) as StoreItemSeed;
const hasSellerContext = !!cleanText(seed?.seller?.memberId) || !!cleanText(seed?.seller?.winportUrl);
if (!hasSellerContext) {
throw new CommandExecutionError(
'1688 store seed item did not expose seller context',
'当前 tab 非商品详情上下文,请切到 detail.1688.com 商品页并重试',
);
}
return seed;
}
function hasAnyEvidence(
storePayload: StoreBrowserPayload | null,
contactPayload: StoreBrowserPayload | null,
seed: StoreItemSeed | null,
): boolean {
return !!cleanText(storePayload?.bodyText)
|| !!cleanText(contactPayload?.bodyText)
|| !!cleanText(seed?.bodyText);
}
cli({
site: '1688',
name: 'store',
description: '1688 店铺/供应商公开信息(联系方式、主营、入驻年限、公开服务信号)',
domain: 'www.1688.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: '1688 店铺 URL 或 member ID(如 b2b-22154705262941f196',
},
],
columns: ['store_name', 'years_on_platform_text', 'location', 'return_rate_text'],
func: async (page, kwargs) => {
const rawInput = String(kwargs.input ?? '');
const resolvedUrl = resolveStoreUrl(rawInput);
const explicitMemberId = extractMemberId(rawInput);
const storePayload = await readStorePayload(page, resolvedUrl, 'store');
const contactUrl = firstContactUrl(storePayload.contactLinks ?? []) || buildContactUrl(storePayload.href || resolvedUrl);
const contactPayload = contactUrl ? await readStorePayload(page, contactUrl, 'store contact') : null;
const offerId = extractOfferId(rawInput)
|| firstOfferId(storePayload.offerLinks ?? [])
|| firstOfferId(contactPayload?.offerLinks ?? []);
let seed: StoreItemSeed | null = null;
if (offerId) {
try {
seed = await readItemSeed(page, offerId);
} catch (error) {
if (!(error instanceof CommandExecutionError)) throw error;
}
}
if (!hasAnyEvidence(storePayload, contactPayload, seed)) {
throw new EmptyResultError(
'1688 store',
'Store page is reachable but no visible fields were extracted. Open the store page in Chrome and retry.',
);
}
return [
normalizeStorePayload({
resolvedUrl,
storePayload,
contactPayload,
seed,
explicitMemberId,
}),
];
},
});
export const __test__ = {
normalizeStorePayload,
safeCanonicalStoreUrl,
buildContactUrl,
firstNamedLine,
firstMetric,
extractReturnRate,
firstOfferId,
firstContactUrl,
};
+32 -39
View File
@@ -5,35 +5,30 @@
*/
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
import type { IPage } from '@jackwener/opencli/types';
/** Extract article ID from a full URL or a bare numeric ID string */
function parseArticleId(input: string): string {
const m = input.match(/\/p\/(\d+)/);
return m ? m[1] : input.replace(/\D/g, '');
function parseArticleId(input) {
const m = input.match(/\/p\/(\d+)/);
return m ? m[1] : input.replace(/\D/g, '');
}
cli({
site: '36kr',
name: 'article',
description: '获取36氪文章正文内容',
domain: 'www.36kr.com',
strategy: Strategy.INTERCEPT,
args: [
{ name: 'id', positional: true, required: true, help: 'Article ID or full 36kr article URL' },
],
columns: ['field', 'value'],
func: async (page: IPage, args) => {
const articleId = parseArticleId(String(args.id ?? ''));
if (!articleId) {
throw new CliError('INVALID_ARGUMENT', 'Invalid article ID or URL');
}
await page.installInterceptor('36kr.com/api');
await page.goto(`https://www.36kr.com/p/${articleId}`);
await page.wait(5);
const data: any = await page.evaluate(`
site: '36kr',
name: 'article',
description: '获取36氪文章正文内容',
domain: 'www.36kr.com',
strategy: Strategy.INTERCEPT,
args: [
{ name: 'id', positional: true, required: true, help: 'Article ID or full 36kr article URL' },
],
columns: ['field', 'value'],
func: async (page, args) => {
const articleId = parseArticleId(String(args.id ?? ''));
if (!articleId) {
throw new CliError('INVALID_ARGUMENT', 'Invalid article ID or URL');
}
await page.installInterceptor('36kr.com/api');
await page.goto(`https://www.36kr.com/p/${articleId}`);
await page.wait(5);
const data = await page.evaluate(`
(() => {
// Title: 36kr uses class "article-title" on h1
const title = document.querySelector('.article-title, h1')?.textContent?.trim() || '';
@@ -53,17 +48,15 @@ cli({
return { title, author, date, body };
})()
`);
if (!data?.title) {
throw new CliError('NOT_FOUND', 'Article not found or failed to load', 'Check the article ID');
}
return [
{ field: 'title', value: data.title },
{ field: 'author', value: data.author || '-' },
{ field: 'date', value: data.date || '-' },
{ field: 'url', value: `https://36kr.com/p/${articleId}` },
{ field: 'body', value: data.body || '-' },
];
},
if (!data?.title) {
throw new CliError('NOT_FOUND', 'Article not found or failed to load', 'Check the article ID');
}
return [
{ field: 'title', value: data.title },
{ field: 'author', value: data.author || '-' },
{ field: 'date', value: data.date || '-' },
{ field: 'url', value: `https://36kr.com/p/${articleId}` },
{ field: 'body', value: data.body || '-' },
];
},
});
+86
View File
@@ -0,0 +1,86 @@
/**
* 36kr hot-list — DOM scraping.
*
* Navigates to the 36kr hot-list page and scrapes rendered article links.
* Supports category types: renqi (人气), zonghe (综合), shoucang (收藏), catalog (综合热门).
*/
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
const TYPE_MAP = {
renqi: '人气榜',
zonghe: '综合榜',
shoucang: '收藏榜',
catalog: '热门资讯',
};
function getShanghaiDate(date = new Date()) {
// Shanghai stays on UTC+8 year-round, so a fixed offset is sufficient here
// and avoids the slow Intl timezone path that timed out on Windows CI.
return new Date(date.getTime() + 8 * 60 * 60 * 1000).toISOString().slice(0, 10);
}
function buildHotListUrl(listType, date = new Date()) {
if (listType === 'catalog') {
return 'https://www.36kr.com/hot-list/catalog';
}
return `https://www.36kr.com/hot-list/${listType}/${getShanghaiDate(date)}/1`;
}
cli({
site: '36kr',
name: 'hot',
description: '36氪热榜 — trending articles (renqi/zonghe/shoucang/catalog)',
domain: 'www.36kr.com',
strategy: Strategy.PUBLIC,
browser: true,
args: [
{ name: 'limit', type: 'int', default: 20, help: 'Number of items (max 50)' },
{
name: 'type',
type: 'string',
default: 'catalog',
help: 'List type: renqi (人气), zonghe (综合), shoucang (收藏), catalog (热门资讯)',
},
],
columns: ['rank', 'title', 'url'],
func: async (page, args) => {
const count = Math.min(Number(args.limit) || 20, 50);
const listType = String(args.type ?? 'catalog');
if (!TYPE_MAP[listType]) {
throw new CliError('INVALID_ARGUMENT', `Unknown type "${listType}". Valid types: ${Object.keys(TYPE_MAP).join(', ')}`);
}
const url = buildHotListUrl(listType);
await page.goto(url);
// Poll DOM until article links appear (36kr renders client-side)
const deadline = Date.now() + 5000;
while (Date.now() < deadline) {
if (await page.evaluate('document.querySelectorAll("a[href*=\\"/p/\\"]").length'))
break;
await new Promise(r => setTimeout(r, 300));
}
// Scrape rendered article links from DOM (deduplicated)
const domItems = await page.evaluate(`
(() => {
const seen = new Set();
const results = [];
const links = document.querySelectorAll('a[href*="/p/"]');
for (const el of links) {
const href = el.getAttribute('href') || '';
const title = el.textContent?.trim() || '';
if (!title || title.length < 5 || seen.has(href) || seen.has(title)) continue;
seen.add(href);
seen.add(title);
results.push({ title, url: href.startsWith('http') ? href : 'https://36kr.com' + href });
}
return results;
})()
`);
const items = Array.isArray(domItems) ? domItems : [];
if (items.length === 0) {
throw new CliError('NO_DATA', 'Could not retrieve 36kr hot list', '36kr may have changed its DOM structure');
}
return items.slice(0, count).map((item, i) => ({
rank: i + 1,
title: item.title,
url: item.url,
}));
},
});
export { buildHotListUrl, getShanghaiDate };
+15
View File
@@ -0,0 +1,15 @@
import { describe, expect, it } from 'vitest';
import { buildHotListUrl, getShanghaiDate } from './hot.js';
describe('36kr/hot date routing', () => {
it('formats dates in Asia/Shanghai instead of UTC', () => {
const date = new Date('2026-03-25T18:30:00.000Z');
expect(getShanghaiDate(date)).toBe('2026-03-26');
});
it('builds dated hot-list routes with Shanghai-local date', () => {
const date = new Date('2026-03-25T18:30:00.000Z');
expect(buildHotListUrl('renqi', date)).toBe('https://www.36kr.com/hot-list/renqi/2026-03-26/1');
});
it('keeps catalog on the static route', () => {
expect(buildHotListUrl('catalog')).toBe('https://www.36kr.com/hot-list/catalog');
});
});
-19
View File
@@ -1,19 +0,0 @@
import { describe, expect, it } from 'vitest';
import { buildHotListUrl, getShanghaiDate } from './hot.js';
describe('36kr/hot date routing', () => {
it('formats dates in Asia/Shanghai instead of UTC', () => {
const date = new Date('2026-03-25T18:30:00.000Z');
expect(getShanghaiDate(date)).toBe('2026-03-26');
});
it('builds dated hot-list routes with Shanghai-local date', () => {
const date = new Date('2026-03-25T18:30:00.000Z');
expect(buildHotListUrl('renqi', date)).toBe('https://www.36kr.com/hot-list/renqi/2026-03-26/1');
});
it('keeps catalog on the static route', () => {
expect(buildHotListUrl('catalog')).toBe('https://www.36kr.com/hot-list/catalog');
});
});
-105
View File
@@ -1,105 +0,0 @@
/**
* 36kr hot-list — DOM scraping.
*
* Navigates to the 36kr hot-list page and scrapes rendered article links.
* Supports category types: renqi (人气), zonghe (综合), shoucang (收藏), catalog (综合热门).
*/
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
import type { IPage } from '@jackwener/opencli/types';
const TYPE_MAP: Record<string, string> = {
renqi: '人气榜',
zonghe: '综合榜',
shoucang: '收藏榜',
catalog: '热门资讯',
};
function getShanghaiDate(date = new Date()): string {
// Shanghai stays on UTC+8 year-round, so a fixed offset is sufficient here
// and avoids the slow Intl timezone path that timed out on Windows CI.
return new Date(date.getTime() + 8 * 60 * 60 * 1000).toISOString().slice(0, 10);
}
function buildHotListUrl(listType: string, date = new Date()): string {
if (listType === 'catalog') {
return 'https://www.36kr.com/hot-list/catalog';
}
return `https://www.36kr.com/hot-list/${listType}/${getShanghaiDate(date)}/1`;
}
cli({
site: '36kr',
name: 'hot',
description: '36氪热榜 — trending articles (renqi/zonghe/shoucang/catalog)',
domain: 'www.36kr.com',
strategy: Strategy.PUBLIC,
browser: true,
args: [
{ name: 'limit', type: 'int', default: 20, help: 'Number of items (max 50)' },
{
name: 'type',
type: 'string',
default: 'catalog',
help: 'List type: renqi (人气), zonghe (综合), shoucang (收藏), catalog (热门资讯)',
},
],
columns: ['rank', 'title', 'url'],
func: async (page: IPage, args) => {
const count = Math.min(Number(args.limit) || 20, 50);
const listType = String(args.type ?? 'catalog');
if (!TYPE_MAP[listType]) {
throw new CliError(
'INVALID_ARGUMENT',
`Unknown type "${listType}". Valid types: ${Object.keys(TYPE_MAP).join(', ')}`,
);
}
const url = buildHotListUrl(listType);
await page.goto(url);
// Poll DOM until article links appear (36kr renders client-side)
const deadline = Date.now() + 5000;
while (Date.now() < deadline) {
if (await page.evaluate('document.querySelectorAll("a[href*=\\"/p/\\"]").length')) break;
await new Promise(r => setTimeout(r, 300));
}
// Scrape rendered article links from DOM (deduplicated)
const domItems: any = await page.evaluate(`
(() => {
const seen = new Set();
const results = [];
const links = document.querySelectorAll('a[href*="/p/"]');
for (const el of links) {
const href = el.getAttribute('href') || '';
const title = el.textContent?.trim() || '';
if (!title || title.length < 5 || seen.has(href) || seen.has(title)) continue;
seen.add(href);
seen.add(title);
results.push({ title, url: href.startsWith('http') ? href : 'https://36kr.com' + href });
}
return results;
})()
`);
const items = Array.isArray(domItems) ? (domItems as any[]) : [];
if (items.length === 0) {
throw new CliError(
'NO_DATA',
'Could not retrieve 36kr hot list',
'36kr may have changed its DOM structure',
);
}
return items.slice(0, count).map((item: any, i: number) => ({
rank: i + 1,
title: item.title,
url: item.url,
}));
},
});
export { buildHotListUrl, getShanghaiDate };
+51
View File
@@ -0,0 +1,51 @@
/**
* 36kr latest news — public RSS feed, no browser needed.
*/
import { cli, Strategy } from '@jackwener/opencli/registry';
cli({
site: '36kr',
name: 'news',
description: 'Latest tech/startup news from 36kr (36氪)',
domain: 'www.36kr.com',
strategy: Strategy.PUBLIC,
args: [
{ name: 'limit', type: 'int', default: 20, help: 'Number of articles (max 50)' },
],
columns: ['rank', 'title', 'summary', 'date', 'url'],
func: async (_page, kwargs) => {
const count = Math.min(kwargs.limit || 20, 50);
const resp = await fetch('https://www.36kr.com/feed', {
headers: { 'User-Agent': 'Mozilla/5.0 (compatible; opencli/1.0)' },
});
if (!resp.ok)
return [];
const xml = await resp.text();
const items = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml)) && items.length < count) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
const url = block.match(/<link><!\[CDATA\[(.*?)\]\]>/)?.[1] ??
block.match(/<link>(.*?)<\/link>/)?.[1] ??
'';
const pubDate = block.match(/<pubDate>(.*?)<\/pubDate>/)?.[1]?.trim() ?? '';
const date = pubDate.slice(0, 10);
// Extract plain-text summary from HTML description (first ~120 chars)
const rawDesc = block.match(/<description><!\[CDATA\[([\s\S]*?)\]\]>/)?.[1] ?? '';
const summary = rawDesc
.replace(/<[^>]+>/g, ' ')
.replace(/&nbsp;/g, ' ')
.replace(/&amp;/g, '&')
.replace(/&lt;/g, '<')
.replace(/&gt;/g, '>')
.replace(/\s+/g, ' ')
.trim()
.slice(0, 120);
if (title) {
items.push({ rank: items.length + 1, title, summary, date, url: url.trim() });
}
}
return items;
},
});
+85
View File
@@ -0,0 +1,85 @@
import { describe, it, expect, vi, afterEach } from 'vitest';
const SAMPLE_RSS = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>36氪</title>
<item>
<title>红杉中国领投AI公司「示例」,金额近2亿元</title>
<link><![CDATA[https://36kr.com/p/1111111111111111?f=rss]]></link>
<pubDate>2026-03-26 10:00:00 +0800</pubDate>
</item>
<item>
<title>马斯克旗下xAI估值突破1000亿美元</title>
<link><![CDATA[https://36kr.com/p/2222222222222222?f=rss]]></link>
<pubDate>2026-03-26 09:00:00 +0800</pubDate>
</item>
<item>
<title>OpenAI发布GPT-5,多模态能力大幅提升</title>
<link><![CDATA[https://36kr.com/p/3333333333333333?f=rss]]></link>
<pubDate>2026-03-25 20:00:00 +0800</pubDate>
</item>
</channel></rss>`;
afterEach(() => {
vi.restoreAllMocks();
});
describe('36kr/news RSS parsing', () => {
it('parses RSS feed into ranked news items', async () => {
vi.spyOn(globalThis, 'fetch').mockResolvedValue({
ok: true,
text: async () => SAMPLE_RSS,
});
// Direct RSS parse test using the same regex logic as news.ts
const xml = SAMPLE_RSS;
const items = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml)) && items.length < 10) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
const url = block.match(/<link><!\[CDATA\[(.*?)\]\]>/)?.[1] ??
block.match(/<link>(.*?)<\/link>/)?.[1] ??
'';
const pubDate = block.match(/<pubDate>(.*?)<\/pubDate>/)?.[1]?.trim() ?? '';
const date = pubDate.slice(0, 10);
if (title)
items.push({ rank: items.length + 1, title, date, url: url.trim() });
}
expect(items).toHaveLength(3);
expect(items[0].rank).toBe(1);
expect(items[0].title).toBe('红杉中国领投AI公司「示例」,金额近2亿元');
expect(items[0].date).toBe('2026-03-26');
expect(items[0].url).toBe('https://36kr.com/p/1111111111111111?f=rss');
});
it('respects limit — returns at most N items', async () => {
const xml = SAMPLE_RSS;
const limit = 2;
const items = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml)) && items.length < limit) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
const url = block.match(/<link><!\[CDATA\[(.*?)\]\]>/)?.[1] ?? '';
const pubDate = block.match(/<pubDate>(.*?)<\/pubDate>/)?.[1]?.trim() ?? '';
const date = pubDate.slice(0, 10);
if (title)
items.push({ rank: items.length + 1, title, date, url: url.trim() });
}
expect(items).toHaveLength(2);
});
it('skips items with empty title', async () => {
const xml = `<rss><channel>
<item><title></title><link>https://36kr.com/p/0</link><pubDate>2026-01-01</pubDate></item>
<item><title>有标题的文章</title><link>https://36kr.com/p/1</link><pubDate>2026-01-01</pubDate></item>
</channel></rss>`;
const items = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml))) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
if (title)
items.push({ title });
}
expect(items).toHaveLength(1);
expect(items[0].title).toBe('有标题的文章');
});
});
-90
View File
@@ -1,90 +0,0 @@
import { describe, it, expect, vi, afterEach } from 'vitest';
const SAMPLE_RSS = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>36氪</title>
<item>
<title>红杉中国领投AI公司「示例」,金额近2亿元</title>
<link><![CDATA[https://36kr.com/p/1111111111111111?f=rss]]></link>
<pubDate>2026-03-26 10:00:00 +0800</pubDate>
</item>
<item>
<title>马斯克旗下xAI估值突破1000亿美元</title>
<link><![CDATA[https://36kr.com/p/2222222222222222?f=rss]]></link>
<pubDate>2026-03-26 09:00:00 +0800</pubDate>
</item>
<item>
<title>OpenAI发布GPT-5,多模态能力大幅提升</title>
<link><![CDATA[https://36kr.com/p/3333333333333333?f=rss]]></link>
<pubDate>2026-03-25 20:00:00 +0800</pubDate>
</item>
</channel></rss>`;
afterEach(() => {
vi.restoreAllMocks();
});
describe('36kr/news RSS parsing', () => {
it('parses RSS feed into ranked news items', async () => {
vi.spyOn(globalThis, 'fetch').mockResolvedValue({
ok: true,
text: async () => SAMPLE_RSS,
} as Response);
// Direct RSS parse test using the same regex logic as news.ts
const xml = SAMPLE_RSS;
const items: { rank: number; title: string; date: string; url: string }[] = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml)) && items.length < 10) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
const url =
block.match(/<link><!\[CDATA\[(.*?)\]\]>/)?.[1] ??
block.match(/<link>(.*?)<\/link>/)?.[1] ??
'';
const pubDate = block.match(/<pubDate>(.*?)<\/pubDate>/)?.[1]?.trim() ?? '';
const date = pubDate.slice(0, 10);
if (title) items.push({ rank: items.length + 1, title, date, url: url.trim() });
}
expect(items).toHaveLength(3);
expect(items[0].rank).toBe(1);
expect(items[0].title).toBe('红杉中国领投AI公司「示例」,金额近2亿元');
expect(items[0].date).toBe('2026-03-26');
expect(items[0].url).toBe('https://36kr.com/p/1111111111111111?f=rss');
});
it('respects limit — returns at most N items', async () => {
const xml = SAMPLE_RSS;
const limit = 2;
const items: { rank: number; title: string; date: string; url: string }[] = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml)) && items.length < limit) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
const url = block.match(/<link><!\[CDATA\[(.*?)\]\]>/)?.[1] ?? '';
const pubDate = block.match(/<pubDate>(.*?)<\/pubDate>/)?.[1]?.trim() ?? '';
const date = pubDate.slice(0, 10);
if (title) items.push({ rank: items.length + 1, title, date, url: url.trim() });
}
expect(items).toHaveLength(2);
});
it('skips items with empty title', async () => {
const xml = `<rss><channel>
<item><title></title><link>https://36kr.com/p/0</link><pubDate>2026-01-01</pubDate></item>
<item><title>有标题的文章</title><link>https://36kr.com/p/1</link><pubDate>2026-01-01</pubDate></item>
</channel></rss>`;
const items: any[] = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml))) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
if (title) items.push({ title });
}
expect(items).toHaveLength(1);
expect(items[0].title).toBe('有标题的文章');
});
});
-54
View File
@@ -1,54 +0,0 @@
/**
* 36kr latest news — public RSS feed, no browser needed.
*/
import { cli, Strategy } from '@jackwener/opencli/registry';
cli({
site: '36kr',
name: 'news',
description: 'Latest tech/startup news from 36kr (36氪)',
domain: 'www.36kr.com',
strategy: Strategy.PUBLIC,
args: [
{ name: 'limit', type: 'int', default: 20, help: 'Number of articles (max 50)' },
],
columns: ['rank', 'title', 'summary', 'date', 'url'],
func: async (_page, kwargs) => {
const count = Math.min(kwargs.limit || 20, 50);
const resp = await fetch('https://www.36kr.com/feed', {
headers: { 'User-Agent': 'Mozilla/5.0 (compatible; opencli/1.0)' },
});
if (!resp.ok) return [];
const xml = await resp.text();
const items: { rank: number; title: string; summary: string; date: string; url: string }[] = [];
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
let match;
while ((match = itemRegex.exec(xml)) && items.length < count) {
const block = match[1];
const title = block.match(/<title>([\s\S]*?)<\/title>/)?.[1]?.trim() ?? '';
const url =
block.match(/<link><!\[CDATA\[(.*?)\]\]>/)?.[1] ??
block.match(/<link>(.*?)<\/link>/)?.[1] ??
'';
const pubDate = block.match(/<pubDate>(.*?)<\/pubDate>/)?.[1]?.trim() ?? '';
const date = pubDate.slice(0, 10);
// Extract plain-text summary from HTML description (first ~120 chars)
const rawDesc = block.match(/<description><!\[CDATA\[([\s\S]*?)\]\]>/)?.[1] ?? '';
const summary = rawDesc
.replace(/<[^>]+>/g, ' ')
.replace(/&nbsp;/g, ' ')
.replace(/&amp;/g, '&')
.replace(/&lt;/g, '<')
.replace(/&gt;/g, '>')
.replace(/\s+/g, ' ')
.trim()
.slice(0, 120);
if (title) {
items.push({ rank: items.length + 1, title, summary, date, url: url.trim() });
}
}
return items;
},
});
+34 -39
View File
@@ -5,33 +5,30 @@
*/
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
import type { IPage } from '@jackwener/opencli/types';
cli({
site: '36kr',
name: 'search',
description: '搜索36氪文章',
domain: 'www.36kr.com',
strategy: Strategy.PUBLIC,
browser: true,
args: [
{ name: 'query', positional: true, required: true, help: 'Search keyword (e.g. "AI", "OpenAI")' },
{ name: 'limit', type: 'int', default: 20, help: 'Number of results (max 50)' },
],
columns: ['rank', 'title', 'date', 'url'],
func: async (page: IPage, args) => {
const count = Math.min(Number(args.limit) || 20, 50);
const query = encodeURIComponent(String(args.query ?? ''));
await page.goto(`https://www.36kr.com/search/articles/${query}`);
// Poll DOM until article links appear (36kr renders client-side)
const deadline = Date.now() + 5000;
while (Date.now() < deadline) {
if (await page.evaluate('document.querySelectorAll("a[href*=\\"/p/\\"]").length')) break;
await new Promise(r => setTimeout(r, 300));
}
const domItems: any = await page.evaluate(`
site: '36kr',
name: 'search',
description: '搜索36氪文章',
domain: 'www.36kr.com',
strategy: Strategy.PUBLIC,
browser: true,
args: [
{ name: 'query', positional: true, required: true, help: 'Search keyword (e.g. "AI", "OpenAI")' },
{ name: 'limit', type: 'int', default: 20, help: 'Number of results (max 50)' },
],
columns: ['rank', 'title', 'date', 'url'],
func: async (page, args) => {
const count = Math.min(Number(args.limit) || 20, 50);
const query = encodeURIComponent(String(args.query ?? ''));
await page.goto(`https://www.36kr.com/search/articles/${query}`);
// Poll DOM until article links appear (36kr renders client-side)
const deadline = Date.now() + 5000;
while (Date.now() < deadline) {
if (await page.evaluate('document.querySelectorAll("a[href*=\\"/p/\\"]").length'))
break;
await new Promise(r => setTimeout(r, 300));
}
const domItems = await page.evaluate(`
(() => {
const seen = new Set();
const results = [];
@@ -67,17 +64,15 @@ cli({
return results;
})()
`);
const items = Array.isArray(domItems) ? (domItems as any[]) : [];
if (items.length === 0) {
throw new CliError('NO_DATA', 'No results found', `Try a different query or check your keyword`);
}
return items.slice(0, count).map((item: any, i: number) => ({
rank: i + 1,
title: item.title,
date: item.date,
url: item.url,
}));
},
const items = Array.isArray(domItems) ? domItems : [];
if (items.length === 0) {
throw new CliError('NO_DATA', 'No results found', `Try a different query or check your keyword`);
}
return items.slice(0, count).map((item, i) => ({
rank: i + 1,
title: item.title,
date: item.date,
url: item.url,
}));
},
});
+32
View File
@@ -0,0 +1,32 @@
/**
* Shared utilities for CLI adapters.
*/
import { ArgumentError } from '@jackwener/opencli/errors';
/**
* Clamp a numeric value to [min, max].
* Matches the signature of lodash.clamp and Rust's clamp.
*/
export function clamp(value, min, max) {
return Math.max(min, Math.min(value, max));
}
export function clampInt(raw, fallback, min, max) {
const parsed = Number(raw);
if (!Number.isFinite(parsed)) {
return fallback;
}
return clamp(Math.floor(parsed), min, max);
}
export function normalizeNumericId(value, label, example) {
const normalized = String(value ?? '').trim();
if (!/^\d+$/.test(normalized)) {
throw new ArgumentError(`${label} must be a numeric ID`, `Pass a numeric ${label}, for example: ${example}`);
}
return normalized;
}
export function requireNonEmptyQuery(value, label = 'query') {
const normalized = String(value ?? '').trim();
if (!normalized) {
throw new ArgumentError(`${label} cannot be empty`);
}
return normalized;
}
-37
View File
@@ -1,37 +0,0 @@
/**
* Shared utilities for CLI adapters.
*/
import { ArgumentError } from '@jackwener/opencli/errors';
/**
* Clamp a numeric value to [min, max].
* Matches the signature of lodash.clamp and Rust's clamp.
*/
export function clamp(value: number, min: number, max: number): number {
return Math.max(min, Math.min(value, max));
}
export function clampInt(raw: unknown, fallback: number, min: number, max: number): number {
const parsed = Number(raw);
if (!Number.isFinite(parsed)) {
return fallback;
}
return clamp(Math.floor(parsed), min, max);
}
export function normalizeNumericId(value: unknown, label: string, example: string): string {
const normalized = String(value ?? '').trim();
if (!/^\d+$/.test(normalized)) {
throw new ArgumentError(`${label} must be a numeric ID`, `Pass a numeric ${label}, for example: ${example}`);
}
return normalized;
}
export function requireNonEmptyQuery(value: unknown, label = 'query'): string {
const normalized = String(value ?? '').trim();
if (!normalized) {
throw new ArgumentError(`${label} cannot be empty`);
}
return normalized;
}
+108
View File
@@ -0,0 +1,108 @@
/**
* Shared command factories for Electron/desktop app adapters.
* Eliminates duplicate screenshot/status/new/dump implementations
* across cursor, codex, chatwise, etc.
*/
import * as fs from 'node:fs';
import { cli, Strategy } from '@jackwener/opencli/registry';
/**
* Factory: capture DOM HTML + accessibility snapshot.
*/
export function makeScreenshotCommand(site, displayName, extra = {}) {
const label = displayName ?? site;
return cli({
...extra,
site,
name: 'screenshot',
description: `Capture a snapshot of the current ${label} window (DOM + Accessibility tree)`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'output', required: false, help: `Output file path (default: /tmp/${site}-snapshot.txt)` },
],
columns: ['Status', 'File'],
func: async (page, kwargs) => {
const outputPath = kwargs.output || `/tmp/${site}-snapshot.txt`;
const snap = await page.snapshot({ compact: true });
const html = await page.evaluate('document.documentElement.outerHTML');
const htmlPath = outputPath.replace(/\.\w+$/, '') + '-dom.html';
const snapPath = outputPath.replace(/\.\w+$/, '') + '-a11y.txt';
fs.writeFileSync(htmlPath, html);
fs.writeFileSync(snapPath, typeof snap === 'string' ? snap : JSON.stringify(snap, null, 2));
return [
{ Status: 'Success', File: htmlPath },
{ Status: 'Success', File: snapPath },
];
},
});
}
/**
* Factory: check CDP connection status.
*/
export function makeStatusCommand(site, displayName, extra = {}) {
const label = displayName ?? site;
return cli({
...extra,
site,
name: 'status',
description: `Check active CDP connection to ${label}`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
columns: ['Status', 'Url', 'Title'],
func: async (page) => {
const url = await page.evaluate('window.location.href');
const title = await page.evaluate('document.title');
return [{ Status: 'Connected', Url: url, Title: title }];
},
});
}
/**
* Factory: start a new session via Cmd/Ctrl+N.
*/
export function makeNewCommand(site, displayName, extra = {}) {
const label = displayName ?? site;
return cli({
...extra,
site,
name: 'new',
description: `Start a new ${label} session`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
columns: ['Status'],
func: async (page) => {
const isMac = process.platform === 'darwin';
await page.pressKey(isMac ? 'Meta+N' : 'Control+N');
await page.wait(1);
return [{ Status: 'Success' }];
},
});
}
/**
* Factory: dump DOM + snapshot for reverse-engineering.
*/
export function makeDumpCommand(site) {
return cli({
site,
name: 'dump',
description: `Dump the DOM and Accessibility tree of ${site} for reverse-engineering`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
columns: ['action', 'files'],
func: async (page) => {
const dom = await page.evaluate('document.body.innerHTML');
fs.writeFileSync(`/tmp/${site}-dom.html`, dom);
const snap = await page.snapshot({ interactive: false });
fs.writeFileSync(`/tmp/${site}-snapshot.json`, JSON.stringify(snap, null, 2));
return [
{
action: 'Dom extraction finished',
files: `/tmp/${site}-dom.html, /tmp/${site}-snapshot.json`,
},
];
},
});
}
-121
View File
@@ -1,121 +0,0 @@
/**
* Shared command factories for Electron/desktop app adapters.
* Eliminates duplicate screenshot/status/new/dump implementations
* across cursor, codex, chatwise, etc.
*/
import * as fs from 'node:fs';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import type { CliOptions } from '@jackwener/opencli/registry';
/**
* Factory: capture DOM HTML + accessibility snapshot.
*/
export function makeScreenshotCommand(site: string, displayName?: string, extra: Partial<CliOptions> = {}) {
const label = displayName ?? site;
return cli({
...extra,
site,
name: 'screenshot',
description: `Capture a snapshot of the current ${label} window (DOM + Accessibility tree)`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'output', required: false, help: `Output file path (default: /tmp/${site}-snapshot.txt)` },
],
columns: ['Status', 'File'],
func: async (page: IPage, kwargs: any) => {
const outputPath = (kwargs.output as string) || `/tmp/${site}-snapshot.txt`;
const snap = await page.snapshot({ compact: true });
const html = await page.evaluate('document.documentElement.outerHTML');
const htmlPath = outputPath.replace(/\.\w+$/, '') + '-dom.html';
const snapPath = outputPath.replace(/\.\w+$/, '') + '-a11y.txt';
fs.writeFileSync(htmlPath, html);
fs.writeFileSync(snapPath, typeof snap === 'string' ? snap : JSON.stringify(snap, null, 2));
return [
{ Status: 'Success', File: htmlPath },
{ Status: 'Success', File: snapPath },
];
},
});
}
/**
* Factory: check CDP connection status.
*/
export function makeStatusCommand(site: string, displayName?: string, extra: Partial<CliOptions> = {}) {
const label = displayName ?? site;
return cli({
...extra,
site,
name: 'status',
description: `Check active CDP connection to ${label}`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
columns: ['Status', 'Url', 'Title'],
func: async (page: IPage) => {
const url = await page.evaluate('window.location.href');
const title = await page.evaluate('document.title');
return [{ Status: 'Connected', Url: url, Title: title }];
},
});
}
/**
* Factory: start a new session via Cmd/Ctrl+N.
*/
export function makeNewCommand(site: string, displayName?: string, extra: Partial<CliOptions> = {}) {
const label = displayName ?? site;
return cli({
...extra,
site,
name: 'new',
description: `Start a new ${label} session`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
columns: ['Status'],
func: async (page: IPage) => {
const isMac = process.platform === 'darwin';
await page.pressKey(isMac ? 'Meta+N' : 'Control+N');
await page.wait(1);
return [{ Status: 'Success' }];
},
});
}
/**
* Factory: dump DOM + snapshot for reverse-engineering.
*/
export function makeDumpCommand(site: string) {
return cli({
site,
name: 'dump',
description: `Dump the DOM and Accessibility tree of ${site} for reverse-engineering`,
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
columns: ['action', 'files'],
func: async (page: IPage) => {
const dom = await page.evaluate('document.body.innerHTML');
fs.writeFileSync(`/tmp/${site}-dom.html`, dom);
const snap = await page.snapshot({ interactive: false });
fs.writeFileSync(`/tmp/${site}-snapshot.json`, JSON.stringify(snap, null, 2));
return [
{
action: 'Dom extraction finished',
files: `/tmp/${site}-dom.html, /tmp/${site}-snapshot.json`,
},
];
},
});
}
@@ -1,8 +1,7 @@
import { cli } from '@jackwener/opencli/registry';
import { createRankingCliOptions } from './rankings.js';
cli(createRankingCliOptions({
commandName: 'new-releases',
listType: 'new_releases',
description: 'Amazon New Releases pages for early momentum discovery',
commandName: 'bestsellers',
listType: 'bestsellers',
description: 'Amazon Best Sellers pages for category candidate discovery',
}));
+29
View File
@@ -0,0 +1,29 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './rankings.js';
describe('amazon bestsellers normalization', () => {
it('normalizes bestseller cards and infers review counts from card text', () => {
const result = __test__.normalizeRankingCandidate({
asin: 'B0DR31GC3D',
title: '',
href: 'https://www.amazon.com/NUTIKAS-Shelves-Desktop-Orgnizer-Shlef/dp/B0DR31GC3D/ref=zg_bs',
price_text: '$25.92',
rating_text: '4.3 out of 5 stars',
review_count_text: '',
card_text: 'Desk Shelves Desktop Organizer Shlef\n4.3 out of 5 stars\n435\n$25.92',
}, {
listType: 'bestsellers',
rankFallback: 2,
listTitle: 'Amazon Best Sellers: Best Desktop & Off-Surface Shelves',
sourceUrl: 'https://www.amazon.com/example',
categoryTitle: null,
categoryUrl: 'https://www.amazon.com/example',
categoryPath: [],
visibleCategoryLinks: [],
});
expect(result.rank).toBe(2);
expect(result.asin).toBe('B0DR31GC3D');
expect(result.title).toBe('Desk Shelves Desktop Organizer Shlef');
expect(result.review_count).toBe(435);
expect(result.list_title).toBe('Amazon Best Sellers: Best Desktop & Off-Surface Shelves');
});
});
-31
View File
@@ -1,31 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './rankings.js';
describe('amazon bestsellers normalization', () => {
it('normalizes bestseller cards and infers review counts from card text', () => {
const result = __test__.normalizeRankingCandidate({
asin: 'B0DR31GC3D',
title: '',
href: 'https://www.amazon.com/NUTIKAS-Shelves-Desktop-Orgnizer-Shlef/dp/B0DR31GC3D/ref=zg_bs',
price_text: '$25.92',
rating_text: '4.3 out of 5 stars',
review_count_text: '',
card_text: 'Desk Shelves Desktop Organizer Shlef\n4.3 out of 5 stars\n435\n$25.92',
}, {
listType: 'bestsellers',
rankFallback: 2,
listTitle: 'Amazon Best Sellers: Best Desktop & Off-Surface Shelves',
sourceUrl: 'https://www.amazon.com/example',
categoryTitle: null,
categoryUrl: 'https://www.amazon.com/example',
categoryPath: [],
visibleCategoryLinks: [],
});
expect(result.rank).toBe(2);
expect(result.asin).toBe('B0DR31GC3D');
expect(result.title).toBe('Desk Shelves Desktop Organizer Shlef');
expect(result.review_count).toBe(435);
expect(result.list_title).toBe('Amazon Best Sellers: Best Desktop & Off-Surface Shelves');
});
});
+91
View File
@@ -0,0 +1,91 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { buildDiscussionUrl, buildProvenance, cleanText, extractAsin, normalizeProductUrl, parseRatingValue, parseReviewCount, trimRatingPrefix, uniqueNonEmpty, assertUsableState, gotoAndReadState, } from './shared.js';
function normalizeDiscussionPayload(payload) {
const sourceUrl = cleanText(payload.href) || buildDiscussionUrl(payload.href ?? '');
const asin = extractAsin(payload.href ?? '') ?? null;
const averageRatingText = cleanText(payload.average_rating_text) || null;
const totalReviewCountText = cleanText(payload.total_review_count_text) || null;
const provenance = buildProvenance(sourceUrl);
return {
asin,
product_url: asin ? normalizeProductUrl(asin) : null,
discussion_url: sourceUrl,
...provenance,
average_rating_text: averageRatingText,
average_rating_value: parseRatingValue(averageRatingText),
total_review_count_text: totalReviewCountText,
total_review_count: parseReviewCount(totalReviewCountText),
qa_urls: uniqueNonEmpty(payload.qa_links ?? []),
review_samples: (payload.review_samples ?? []).map((sample) => ({
title: trimRatingPrefix(sample.title) || null,
rating_text: cleanText(sample.rating_text) || null,
rating_value: parseRatingValue(sample.rating_text),
author: cleanText(sample.author) || null,
date_text: cleanText(sample.date_text) || null,
body: cleanText(sample.body) || null,
verified_purchase: sample.verified === true,
})),
};
}
async function readDiscussionPayload(page, input, limit) {
const url = buildDiscussionUrl(input);
const state = await gotoAndReadState(page, url, 2500, 'discussion');
assertUsableState(state, 'discussion');
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
average_rating_text: document.querySelector('[data-hook="rating-out-of-text"]')?.textContent || '',
total_review_count_text: document.querySelector('[data-hook="total-review-count"]')?.textContent || '',
qa_links: Array.from(document.querySelectorAll('a[href*="ask/questions"]')).map((anchor) => anchor.href || ''),
review_samples: Array.from(document.querySelectorAll('[data-hook="review"]')).slice(0, ${limit}).map((card) => ({
title: card.querySelector('[data-hook="review-title"]')?.textContent || '',
rating_text:
card.querySelector('[data-hook="review-star-rating"]')?.textContent
|| card.querySelector('[data-hook="cmps-review-star-rating"]')?.textContent
|| '',
author: card.querySelector('.a-profile-name')?.textContent || '',
date_text: card.querySelector('[data-hook="review-date"]')?.textContent || '',
body: card.querySelector('[data-hook="review-body"]')?.textContent || '',
verified: !!card.querySelector('[data-hook="avp-badge"]'),
})),
}))()
`);
}
cli({
site: 'amazon',
name: 'discussion',
description: 'Amazon review summary and sample customer discussion from product review pages',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: 'ASIN or product URL, for example B0FJS72893',
},
{
name: 'limit',
type: 'int',
default: 10,
help: 'Maximum number of review samples to return (default 10)',
},
],
columns: ['asin', 'average_rating_value', 'total_review_count'],
func: async (page, kwargs) => {
const input = String(kwargs.input ?? '');
const limit = Math.max(1, Number(kwargs.limit) || 10);
const payload = await readDiscussionPayload(page, input, limit);
const normalized = normalizeDiscussionPayload(payload);
if (!normalized.average_rating_text && !normalized.total_review_count_text) {
throw new CommandExecutionError('amazon discussion page did not expose review summary', 'The review page may have changed or hit a robot check. Open the review page in Chrome and retry.');
}
return [normalized];
},
});
export const __test__ = {
normalizeDiscussionPayload,
};
+36
View File
@@ -0,0 +1,36 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './discussion.js';
describe('amazon discussion normalization', () => {
it('normalizes review summary and sample reviews', () => {
const result = __test__.normalizeDiscussionPayload({
href: 'https://www.amazon.com/product-reviews/B0FJS72893',
average_rating_text: '3.9 out of 5',
total_review_count_text: '27 global ratings',
qa_links: [],
review_samples: [
{
title: '5.0 out of 5 stars Great value and quality',
rating_text: '5.0 out of 5 stars',
author: 'GTreader2',
date_text: 'Reviewed in the United States on February 21, 2026',
body: 'Small but mighty.',
verified: true,
},
],
});
expect(result.asin).toBe('B0FJS72893');
expect(result.average_rating_value).toBe(3.9);
expect(result.total_review_count).toBe(27);
expect(result.review_samples).toEqual([
{
title: 'Great value and quality',
rating_text: '5.0 out of 5 stars',
rating_value: 5,
author: 'GTreader2',
date_text: 'Reviewed in the United States on February 21, 2026',
body: 'Small but mighty.',
verified_purchase: true,
},
]);
});
});
-38
View File
@@ -1,38 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './discussion.js';
describe('amazon discussion normalization', () => {
it('normalizes review summary and sample reviews', () => {
const result = __test__.normalizeDiscussionPayload({
href: 'https://www.amazon.com/product-reviews/B0FJS72893',
average_rating_text: '3.9 out of 5',
total_review_count_text: '27 global ratings',
qa_links: [],
review_samples: [
{
title: '5.0 out of 5 stars Great value and quality',
rating_text: '5.0 out of 5 stars',
author: 'GTreader2',
date_text: 'Reviewed in the United States on February 21, 2026',
body: 'Small but mighty.',
verified: true,
},
],
});
expect(result.asin).toBe('B0FJS72893');
expect(result.average_rating_value).toBe(3.9);
expect(result.total_review_count).toBe(27);
expect(result.review_samples).toEqual([
{
title: 'Great value and quality',
rating_text: '5.0 out of 5 stars',
rating_value: 5,
author: 'GTreader2',
date_text: 'Reviewed in the United States on February 21, 2026',
body: 'Small but mighty.',
verified_purchase: true,
},
]);
});
});
-131
View File
@@ -1,131 +0,0 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
buildDiscussionUrl,
buildProvenance,
cleanText,
extractAsin,
normalizeProductUrl,
parseRatingValue,
parseReviewCount,
trimRatingPrefix,
uniqueNonEmpty,
assertUsableState,
gotoAndReadState,
} from './shared.js';
interface DiscussionPayload {
href?: string;
title?: string;
average_rating_text?: string | null;
total_review_count_text?: string | null;
qa_links?: string[];
review_samples?: Array<{
title?: string | null;
rating_text?: string | null;
author?: string | null;
date_text?: string | null;
body?: string | null;
verified?: boolean;
}>;
}
function normalizeDiscussionPayload(payload: DiscussionPayload): Record<string, unknown> {
const sourceUrl = cleanText(payload.href) || buildDiscussionUrl(payload.href ?? '');
const asin = extractAsin(payload.href ?? '') ?? null;
const averageRatingText = cleanText(payload.average_rating_text) || null;
const totalReviewCountText = cleanText(payload.total_review_count_text) || null;
const provenance = buildProvenance(sourceUrl);
return {
asin,
product_url: asin ? normalizeProductUrl(asin) : null,
discussion_url: sourceUrl,
...provenance,
average_rating_text: averageRatingText,
average_rating_value: parseRatingValue(averageRatingText),
total_review_count_text: totalReviewCountText,
total_review_count: parseReviewCount(totalReviewCountText),
qa_urls: uniqueNonEmpty(payload.qa_links ?? []),
review_samples: (payload.review_samples ?? []).map((sample) => ({
title: trimRatingPrefix(sample.title) || null,
rating_text: cleanText(sample.rating_text) || null,
rating_value: parseRatingValue(sample.rating_text),
author: cleanText(sample.author) || null,
date_text: cleanText(sample.date_text) || null,
body: cleanText(sample.body) || null,
verified_purchase: sample.verified === true,
})),
};
}
async function readDiscussionPayload(page: IPage, input: string, limit: number): Promise<DiscussionPayload> {
const url = buildDiscussionUrl(input);
const state = await gotoAndReadState(page, url, 2500, 'discussion');
assertUsableState(state, 'discussion');
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
average_rating_text: document.querySelector('[data-hook="rating-out-of-text"]')?.textContent || '',
total_review_count_text: document.querySelector('[data-hook="total-review-count"]')?.textContent || '',
qa_links: Array.from(document.querySelectorAll('a[href*="ask/questions"]')).map((anchor) => anchor.href || ''),
review_samples: Array.from(document.querySelectorAll('[data-hook="review"]')).slice(0, ${limit}).map((card) => ({
title: card.querySelector('[data-hook="review-title"]')?.textContent || '',
rating_text:
card.querySelector('[data-hook="review-star-rating"]')?.textContent
|| card.querySelector('[data-hook="cmps-review-star-rating"]')?.textContent
|| '',
author: card.querySelector('.a-profile-name')?.textContent || '',
date_text: card.querySelector('[data-hook="review-date"]')?.textContent || '',
body: card.querySelector('[data-hook="review-body"]')?.textContent || '',
verified: !!card.querySelector('[data-hook="avp-badge"]'),
})),
}))()
`) as DiscussionPayload;
}
cli({
site: 'amazon',
name: 'discussion',
description: 'Amazon review summary and sample customer discussion from product review pages',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: 'ASIN or product URL, for example B0FJS72893',
},
{
name: 'limit',
type: 'int',
default: 10,
help: 'Maximum number of review samples to return (default 10)',
},
],
columns: ['asin', 'average_rating_value', 'total_review_count'],
func: async (page, kwargs) => {
const input = String(kwargs.input ?? '');
const limit = Math.max(1, Number(kwargs.limit) || 10);
const payload = await readDiscussionPayload(page, input, limit);
const normalized = normalizeDiscussionPayload(payload);
if (!normalized.average_rating_text && !normalized.total_review_count_text) {
throw new CommandExecutionError(
'amazon discussion page did not expose review summary',
'The review page may have changed or hit a robot check. Open the review page in Chrome and retry.',
);
}
return [normalized];
},
});
export const __test__ = {
normalizeDiscussionPayload,
};
+7
View File
@@ -0,0 +1,7 @@
import { cli } from '@jackwener/opencli/registry';
import { createRankingCliOptions } from './rankings.js';
cli(createRankingCliOptions({
commandName: 'movers-shakers',
listType: 'movers_shakers',
description: 'Amazon Movers & Shakers pages for short-term growth signals',
}));
-8
View File
@@ -1,8 +0,0 @@
import { cli } from '@jackwener/opencli/registry';
import { createRankingCliOptions } from './rankings.js';
cli(createRankingCliOptions({
commandName: 'movers-shakers',
listType: 'movers_shakers',
description: 'Amazon Movers & Shakers pages for short-term growth signals',
}));
@@ -1,8 +1,7 @@
import { cli } from '@jackwener/opencli/registry';
import { createRankingCliOptions } from './rankings.js';
cli(createRankingCliOptions({
commandName: 'bestsellers',
listType: 'bestsellers',
description: 'Amazon Best Sellers pages for category candidate discovery',
commandName: 'new-releases',
listType: 'new_releases',
description: 'Amazon New Releases pages for early momentum discovery',
}));
+140
View File
@@ -0,0 +1,140 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { buildProductUrl, buildProvenance, cleanText, extractAsin, isAmazonEntity, normalizeProductUrl, PRIMARY_PRICE_SELECTORS, parsePriceText, assertUsableState, gotoAndReadState, } from './shared.js';
const OFFER_FACT_SELECTOR = [
'#sellerProfileTriggerId',
'#shipsFromSoldByInsideBuyBox_feature_div',
'#fulfillerInfoFeature_feature_div',
'#merchantInfoFeature_feature_div',
'#tabular-buybox-container',
'#merchant-info',
].join(', ');
function collapseAdjacentWords(text) {
const parts = cleanText(text).split(' ').filter(Boolean);
const deduped = [];
for (const part of parts) {
if (deduped[deduped.length - 1] === part)
continue;
deduped.push(part);
}
return deduped.join(' ');
}
function extractShipsFrom(text) {
const normalized = cleanText(text);
const match = normalized.match(/Ships from\s+(.+?)(?=Sold by|and Fulfilled by|$)/i);
return match ? collapseAdjacentWords(match[1].replace(/Ships from/ig, '')) : null;
}
function extractSoldBy(text) {
const normalized = cleanText(text);
const match = normalized.match(/Sold by\s+(.+?)(?=and Fulfilled by|Ships from|$)/i);
return match ? collapseAdjacentWords(match[1]) : null;
}
function isDeliveryLocationBlocked(text) {
const normalized = cleanText(text).toLowerCase();
return normalized.includes('cannot be shipped to your selected delivery location')
|| normalized.includes('similar items shipping to')
|| normalized.includes('deliver to hong kong');
}
function normalizeOfferPayload(payload) {
const asin = extractAsin(payload.href ?? '') ?? null;
const sourceUrl = cleanText(payload.href) || buildProductUrl(payload.href ?? '');
const price = parsePriceText(payload.price_text);
const merchantInfo = cleanText(payload.merchant_info) || null;
const soldBy = cleanText(payload.sold_by)
|| extractSoldBy(payload.ships_from_text ?? '')
|| extractSoldBy(merchantInfo ?? '')
|| null;
const shipsFrom = extractShipsFrom(payload.ships_from_text ?? '')
|| extractShipsFrom(merchantInfo ?? '')
|| cleanText(payload.ships_from_text)
|| null;
const provenance = buildProvenance(sourceUrl);
return {
asin,
product_url: normalizeProductUrl(payload.href),
...provenance,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
merchant_info_text: merchantInfo,
sold_by: soldBy,
ships_from: shipsFrom,
offer_listing_url: cleanText(payload.offer_link) || null,
review_url: cleanText(payload.review_url) || null,
qa_url: cleanText(payload.qa_url) || null,
is_amazon_sold: isAmazonEntity(soldBy),
is_amazon_fulfilled: isAmazonEntity(shipsFrom) || /fulfilled by amazon/i.test(merchantInfo ?? ''),
};
}
async function readOfferPayload(page, input) {
const url = buildProductUrl(input);
const state = await gotoAndReadState(page, url, 2500, 'offer');
assertUsableState(state, 'offer');
// Reconnecting to an existing Amazon target can surface the product page
// before the buy-box / merchant blocks are reattached to the DOM.
await page.wait({ selector: OFFER_FACT_SELECTOR, timeout: 6 }).catch(() => { });
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
price_text: (() => {
const selectors = ${JSON.stringify(PRIMARY_PRICE_SELECTORS)};
for (const selector of selectors) {
const text = document.querySelector(selector)?.textContent || '';
if (text.trim()) return text;
}
return '';
})(),
merchant_info: document.querySelector('#merchant-info')?.textContent || '',
sold_by: document.querySelector('#sellerProfileTriggerId')?.textContent || '',
ships_from_text:
document.querySelector('#shipsFromSoldByInsideBuyBox_feature_div')?.textContent
|| document.querySelector('#fulfillerInfoFeature_feature_div')?.textContent
|| document.querySelector('#merchantInfoFeature_feature_div')?.textContent
|| document.querySelector('#tabular-buybox-container')?.textContent
|| '',
offer_link: document.querySelector('a[href*="/gp/offer-listing/"]')?.href || '',
review_url: document.querySelector('a[href*="#customerReviews"]')?.href || '',
qa_url: document.querySelector('a[href*="ask/questions"]')?.href || '',
buybox_text:
document.querySelector('#desktop_qualifiedBuyBox')?.textContent
|| document.querySelector('#buybox')?.textContent
|| '',
}))()
`);
}
cli({
site: 'amazon',
name: 'offer',
description: 'Amazon seller, buy box, and fulfillment facts from the product page',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: 'ASIN or product URL, for example B0FJS72893',
},
],
columns: ['asin', 'price_text', 'sold_by', 'ships_from', 'is_amazon_sold', 'is_amazon_fulfilled'],
func: async (page, kwargs) => {
const input = String(kwargs.input ?? '');
const payload = await readOfferPayload(page, input);
const normalized = normalizeOfferPayload(payload);
if (!normalized.sold_by && !normalized.ships_from && !normalized.merchant_info_text) {
if (isDeliveryLocationBlocked(payload.buybox_text)) {
throw new CommandExecutionError('amazon offer buy box is blocked by the current delivery location', 'The shared Chrome profile is not set to the target US delivery address. Switch Amazon delivery location to the requested US destination, reopen the product page, and retry.');
}
throw new CommandExecutionError('amazon offer surface did not expose seller or fulfillment facts', 'The product page may have changed. Open the product page in Chrome, make sure the buy box is visible, and retry.');
}
return [normalized];
},
});
export const __test__ = {
extractShipsFrom,
extractSoldBy,
isDeliveryLocationBlocked,
normalizeOfferPayload,
};
+29
View File
@@ -0,0 +1,29 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './offer.js';
describe('amazon offer normalization', () => {
it('extracts sold-by and fulfillment facts from product offer text', () => {
const result = __test__.normalizeOfferPayload({
href: 'https://www.amazon.com/dp/B0FJS72893',
price_text: '$15.99',
merchant_info: '',
sold_by: 'KUATUDIRECT',
ships_from_text: 'Ships from Amazon',
offer_link: null,
review_url: 'https://www.amazon.com/dp/B0FJS72893#customerReviews',
qa_url: null,
});
expect(result.asin).toBe('B0FJS72893');
expect(result.sold_by).toBe('KUATUDIRECT');
expect(result.ships_from).toBe('Amazon');
expect(result.is_amazon_sold).toBe(false);
expect(result.is_amazon_fulfilled).toBe(true);
});
it('parses merchant info fallback text', () => {
expect(__test__.extractSoldBy('Sold by Example Seller and Fulfilled by Amazon.')).toBe('Example Seller');
expect(__test__.extractShipsFrom('Ships from Amazon')).toBe('Amazon');
});
it('detects delivery-location blocking in the buy box text', () => {
expect(__test__.isDeliveryLocationBlocked('This item cannot be shipped to your selected delivery location. Similar items shipping to Hong Kong')).toBe(true);
expect(__test__.isDeliveryLocationBlocked('Ships from Amazon')).toBe(false);
});
});
-35
View File
@@ -1,35 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './offer.js';
describe('amazon offer normalization', () => {
it('extracts sold-by and fulfillment facts from product offer text', () => {
const result = __test__.normalizeOfferPayload({
href: 'https://www.amazon.com/dp/B0FJS72893',
price_text: '$15.99',
merchant_info: '',
sold_by: 'KUATUDIRECT',
ships_from_text: 'Ships from Amazon',
offer_link: null,
review_url: 'https://www.amazon.com/dp/B0FJS72893#customerReviews',
qa_url: null,
});
expect(result.asin).toBe('B0FJS72893');
expect(result.sold_by).toBe('KUATUDIRECT');
expect(result.ships_from).toBe('Amazon');
expect(result.is_amazon_sold).toBe(false);
expect(result.is_amazon_fulfilled).toBe(true);
});
it('parses merchant info fallback text', () => {
expect(__test__.extractSoldBy('Sold by Example Seller and Fulfilled by Amazon.')).toBe('Example Seller');
expect(__test__.extractShipsFrom('Ships from Amazon')).toBe('Amazon');
});
it('detects delivery-location blocking in the buy box text', () => {
expect(__test__.isDeliveryLocationBlocked(
'This item cannot be shipped to your selected delivery location. Similar items shipping to Hong Kong',
)).toBe(true);
expect(__test__.isDeliveryLocationBlocked('Ships from Amazon')).toBe(false);
});
});
-185
View File
@@ -1,185 +0,0 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
buildProductUrl,
buildProvenance,
cleanText,
extractAsin,
isAmazonEntity,
normalizeProductUrl,
PRIMARY_PRICE_SELECTORS,
parsePriceText,
assertUsableState,
gotoAndReadState,
} from './shared.js';
interface OfferPayload {
href?: string;
title?: string;
price_text?: string | null;
merchant_info?: string | null;
sold_by?: string | null;
ships_from_text?: string | null;
offer_link?: string | null;
review_url?: string | null;
qa_url?: string | null;
buybox_text?: string | null;
}
const OFFER_FACT_SELECTOR = [
'#sellerProfileTriggerId',
'#shipsFromSoldByInsideBuyBox_feature_div',
'#fulfillerInfoFeature_feature_div',
'#merchantInfoFeature_feature_div',
'#tabular-buybox-container',
'#merchant-info',
].join(', ');
function collapseAdjacentWords(text: string): string {
const parts = cleanText(text).split(' ').filter(Boolean);
const deduped: string[] = [];
for (const part of parts) {
if (deduped[deduped.length - 1] === part) continue;
deduped.push(part);
}
return deduped.join(' ');
}
function extractShipsFrom(text: string): string | null {
const normalized = cleanText(text);
const match = normalized.match(/Ships from\s+(.+?)(?=Sold by|and Fulfilled by|$)/i);
return match ? collapseAdjacentWords(match[1].replace(/Ships from/ig, '')) : null;
}
function extractSoldBy(text: string): string | null {
const normalized = cleanText(text);
const match = normalized.match(/Sold by\s+(.+?)(?=and Fulfilled by|Ships from|$)/i);
return match ? collapseAdjacentWords(match[1]) : null;
}
function isDeliveryLocationBlocked(text: string | null | undefined): boolean {
const normalized = cleanText(text).toLowerCase();
return normalized.includes('cannot be shipped to your selected delivery location')
|| normalized.includes('similar items shipping to')
|| normalized.includes('deliver to hong kong');
}
function normalizeOfferPayload(payload: OfferPayload): Record<string, unknown> {
const asin = extractAsin(payload.href ?? '') ?? null;
const sourceUrl = cleanText(payload.href) || buildProductUrl(payload.href ?? '');
const price = parsePriceText(payload.price_text);
const merchantInfo = cleanText(payload.merchant_info) || null;
const soldBy = cleanText(payload.sold_by)
|| extractSoldBy(payload.ships_from_text ?? '')
|| extractSoldBy(merchantInfo ?? '')
|| null;
const shipsFrom = extractShipsFrom(payload.ships_from_text ?? '')
|| extractShipsFrom(merchantInfo ?? '')
|| cleanText(payload.ships_from_text)
|| null;
const provenance = buildProvenance(sourceUrl);
return {
asin,
product_url: normalizeProductUrl(payload.href),
...provenance,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
merchant_info_text: merchantInfo,
sold_by: soldBy,
ships_from: shipsFrom,
offer_listing_url: cleanText(payload.offer_link) || null,
review_url: cleanText(payload.review_url) || null,
qa_url: cleanText(payload.qa_url) || null,
is_amazon_sold: isAmazonEntity(soldBy),
is_amazon_fulfilled: isAmazonEntity(shipsFrom) || /fulfilled by amazon/i.test(merchantInfo ?? ''),
};
}
async function readOfferPayload(page: IPage, input: string): Promise<OfferPayload> {
const url = buildProductUrl(input);
const state = await gotoAndReadState(page, url, 2500, 'offer');
assertUsableState(state, 'offer');
// Reconnecting to an existing Amazon target can surface the product page
// before the buy-box / merchant blocks are reattached to the DOM.
await page.wait({ selector: OFFER_FACT_SELECTOR, timeout: 6 }).catch(() => {});
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
price_text: (() => {
const selectors = ${JSON.stringify(PRIMARY_PRICE_SELECTORS)};
for (const selector of selectors) {
const text = document.querySelector(selector)?.textContent || '';
if (text.trim()) return text;
}
return '';
})(),
merchant_info: document.querySelector('#merchant-info')?.textContent || '',
sold_by: document.querySelector('#sellerProfileTriggerId')?.textContent || '',
ships_from_text:
document.querySelector('#shipsFromSoldByInsideBuyBox_feature_div')?.textContent
|| document.querySelector('#fulfillerInfoFeature_feature_div')?.textContent
|| document.querySelector('#merchantInfoFeature_feature_div')?.textContent
|| document.querySelector('#tabular-buybox-container')?.textContent
|| '',
offer_link: document.querySelector('a[href*="/gp/offer-listing/"]')?.href || '',
review_url: document.querySelector('a[href*="#customerReviews"]')?.href || '',
qa_url: document.querySelector('a[href*="ask/questions"]')?.href || '',
buybox_text:
document.querySelector('#desktop_qualifiedBuyBox')?.textContent
|| document.querySelector('#buybox')?.textContent
|| '',
}))()
`) as OfferPayload;
}
cli({
site: 'amazon',
name: 'offer',
description: 'Amazon seller, buy box, and fulfillment facts from the product page',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: 'ASIN or product URL, for example B0FJS72893',
},
],
columns: ['asin', 'price_text', 'sold_by', 'ships_from', 'is_amazon_sold', 'is_amazon_fulfilled'],
func: async (page, kwargs) => {
const input = String(kwargs.input ?? '');
const payload = await readOfferPayload(page, input);
const normalized = normalizeOfferPayload(payload);
if (!normalized.sold_by && !normalized.ships_from && !normalized.merchant_info_text) {
if (isDeliveryLocationBlocked(payload.buybox_text)) {
throw new CommandExecutionError(
'amazon offer buy box is blocked by the current delivery location',
'The shared Chrome profile is not set to the target US delivery address. Switch Amazon delivery location to the requested US destination, reopen the product page, and retry.',
);
}
throw new CommandExecutionError(
'amazon offer surface did not expose seller or fulfillment facts',
'The product page may have changed. Open the product page in Chrome, make sure the buy box is visible, and retry.',
);
}
return [normalized];
},
});
export const __test__ = {
extractShipsFrom,
extractSoldBy,
isDeliveryLocationBlocked,
normalizeOfferPayload,
};
+92
View File
@@ -0,0 +1,92 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { buildProductUrl, buildProvenance, cleanText, extractAsin, PRIMARY_PRICE_SELECTORS, parsePriceText, parseRatingValue, parseReviewCount, normalizeProductUrl, uniqueNonEmpty, assertUsableState, gotoAndReadState, } from './shared.js';
const PRODUCT_TITLE_SELECTOR = '#productTitle, #title span, [data-feature-name="title"] h1 span';
const BYLINE_SELECTOR = '#bylineInfo, [data-feature-name="bylineInfo"] #bylineInfo';
function normalizeProductPayload(payload) {
const sourceUrl = cleanText(payload.href) || buildProductUrl(cleanText(payload.product_title) || cleanText(payload.href));
const asin = extractAsin(payload.href ?? '') ?? null;
const price = parsePriceText(payload.price_text);
const ratingText = cleanText(payload.rating_text) || null;
const reviewCountText = cleanText(payload.review_count_text) || null;
const provenance = buildProvenance(sourceUrl);
return {
asin,
title: cleanText(payload.product_title) || cleanText(payload.title) || null,
product_url: normalizeProductUrl(payload.href),
...provenance,
brand_text: cleanText(payload.byline) || null,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
rating_text: ratingText,
rating_value: parseRatingValue(ratingText),
review_count_text: reviewCountText,
review_count: parseReviewCount(reviewCountText),
review_url: cleanText(payload.review_url) || null,
qa_url: cleanText(payload.qa_url) || null,
breadcrumbs: uniqueNonEmpty(payload.breadcrumbs ?? []),
bullet_points: uniqueNonEmpty(payload.bullets ?? []),
};
}
async function readProductPayload(page, input) {
const url = buildProductUrl(input);
const state = await gotoAndReadState(page, url, 2500, 'product');
assertUsableState(state, 'product');
// Amazon can report a "stable" DOM before the product title block hydrates,
// especially when reconnecting to an existing shared CDP target.
await page.wait({ selector: PRODUCT_TITLE_SELECTOR, timeout: 6 }).catch(() => { });
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
product_title: document.querySelector(${JSON.stringify(PRODUCT_TITLE_SELECTOR)})?.textContent || '',
byline: document.querySelector(${JSON.stringify(BYLINE_SELECTOR)})?.textContent || '',
price_text: (() => {
const selectors = ${JSON.stringify(PRIMARY_PRICE_SELECTORS)};
for (const selector of selectors) {
const text = document.querySelector(selector)?.textContent || '';
if (text.trim()) return text;
}
return '';
})(),
rating_text:
document.querySelector('#acrPopover')?.getAttribute('title')
|| document.querySelector('#acrPopover')?.textContent
|| '',
review_count_text: document.querySelector('#acrCustomerReviewText')?.textContent || '',
review_url: document.querySelector('a[href*="#customerReviews"]')?.href || '',
qa_url: document.querySelector('a[href*="ask/questions"]')?.href || '',
bullets: Array.from(document.querySelectorAll('#feature-bullets li .a-list-item')).map((node) => node.textContent || ''),
breadcrumbs: Array.from(document.querySelectorAll('#wayfinding-breadcrumbs_feature_div a')).map((node) => node.textContent || ''),
}))()
`);
}
cli({
site: 'amazon',
name: 'product',
description: 'Amazon product page facts for candidate validation',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: 'ASIN or product URL, for example B0FJS72893',
},
],
columns: ['asin', 'title', 'price_text', 'rating_value', 'review_count'],
func: async (page, kwargs) => {
const input = String(kwargs.input ?? '');
const payload = await readProductPayload(page, input);
if (!cleanText(payload.product_title)) {
throw new CommandExecutionError('amazon product page did not expose product content', 'The product page may have changed or hit a robot check. Open the product page in Chrome and retry.');
}
return [normalizeProductPayload(payload)];
},
});
export const __test__ = {
normalizeProductPayload,
};
+24
View File
@@ -0,0 +1,24 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './product.js';
describe('amazon product normalization', () => {
it('normalizes product facts from the product page', () => {
const result = __test__.normalizeProductPayload({
href: 'https://www.amazon.com/dp/B0FJS72893',
title: 'Amazon.com: KVTUKIAIT Desktop Shelf Organizer',
product_title: 'White Desktop Shelf Organizer for Top of Desk',
byline: 'Visit the KVTUKIAIT Store',
price_text: '$15.99',
rating_text: '3.9 out of 5 stars',
review_count_text: '27 ratings',
review_url: 'https://www.amazon.com/dp/B0FJS72893#customerReviews',
qa_url: null,
bullets: ['SPACE-SAVING DESK SHELF ORGANIZER', 'SMALL AND STYLISH AESTHETIC DECOR'],
breadcrumbs: ['Office Products', 'Desktop & Off-Surface Shelves'],
});
expect(result.asin).toBe('B0FJS72893');
expect(result.price_value).toBe(15.99);
expect(result.rating_value).toBe(3.9);
expect(result.review_count).toBe(27);
expect(result.breadcrumbs).toEqual(['Office Products', 'Desktop & Off-Surface Shelves']);
});
});
-26
View File
@@ -1,26 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './product.js';
describe('amazon product normalization', () => {
it('normalizes product facts from the product page', () => {
const result = __test__.normalizeProductPayload({
href: 'https://www.amazon.com/dp/B0FJS72893',
title: 'Amazon.com: KVTUKIAIT Desktop Shelf Organizer',
product_title: 'White Desktop Shelf Organizer for Top of Desk',
byline: 'Visit the KVTUKIAIT Store',
price_text: '$15.99',
rating_text: '3.9 out of 5 stars',
review_count_text: '27 ratings',
review_url: 'https://www.amazon.com/dp/B0FJS72893#customerReviews',
qa_url: null,
bullets: ['SPACE-SAVING DESK SHELF ORGANIZER', 'SMALL AND STYLISH AESTHETIC DECOR'],
breadcrumbs: ['Office Products', 'Desktop & Off-Surface Shelves'],
});
expect(result.asin).toBe('B0FJS72893');
expect(result.price_value).toBe(15.99);
expect(result.rating_value).toBe(3.9);
expect(result.review_count).toBe(27);
expect(result.breadcrumbs).toEqual(['Office Products', 'Desktop & Off-Surface Shelves']);
});
});
-131
View File
@@ -1,131 +0,0 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
buildProductUrl,
buildProvenance,
cleanText,
extractAsin,
PRIMARY_PRICE_SELECTORS,
parsePriceText,
parseRatingValue,
parseReviewCount,
normalizeProductUrl,
uniqueNonEmpty,
assertUsableState,
gotoAndReadState,
} from './shared.js';
interface ProductPayload {
href?: string;
title?: string;
product_title?: string | null;
byline?: string | null;
price_text?: string | null;
rating_text?: string | null;
review_count_text?: string | null;
review_url?: string | null;
qa_url?: string | null;
bullets?: string[];
breadcrumbs?: string[];
}
const PRODUCT_TITLE_SELECTOR = '#productTitle, #title span, [data-feature-name="title"] h1 span';
const BYLINE_SELECTOR = '#bylineInfo, [data-feature-name="bylineInfo"] #bylineInfo';
function normalizeProductPayload(payload: ProductPayload): Record<string, unknown> {
const sourceUrl = cleanText(payload.href) || buildProductUrl(cleanText(payload.product_title) || cleanText(payload.href));
const asin = extractAsin(payload.href ?? '') ?? null;
const price = parsePriceText(payload.price_text);
const ratingText = cleanText(payload.rating_text) || null;
const reviewCountText = cleanText(payload.review_count_text) || null;
const provenance = buildProvenance(sourceUrl);
return {
asin,
title: cleanText(payload.product_title) || cleanText(payload.title) || null,
product_url: normalizeProductUrl(payload.href),
...provenance,
brand_text: cleanText(payload.byline) || null,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
rating_text: ratingText,
rating_value: parseRatingValue(ratingText),
review_count_text: reviewCountText,
review_count: parseReviewCount(reviewCountText),
review_url: cleanText(payload.review_url) || null,
qa_url: cleanText(payload.qa_url) || null,
breadcrumbs: uniqueNonEmpty(payload.breadcrumbs ?? []),
bullet_points: uniqueNonEmpty(payload.bullets ?? []),
};
}
async function readProductPayload(page: IPage, input: string): Promise<ProductPayload> {
const url = buildProductUrl(input);
const state = await gotoAndReadState(page, url, 2500, 'product');
assertUsableState(state, 'product');
// Amazon can report a "stable" DOM before the product title block hydrates,
// especially when reconnecting to an existing shared CDP target.
await page.wait({ selector: PRODUCT_TITLE_SELECTOR, timeout: 6 }).catch(() => {});
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
product_title: document.querySelector(${JSON.stringify(PRODUCT_TITLE_SELECTOR)})?.textContent || '',
byline: document.querySelector(${JSON.stringify(BYLINE_SELECTOR)})?.textContent || '',
price_text: (() => {
const selectors = ${JSON.stringify(PRIMARY_PRICE_SELECTORS)};
for (const selector of selectors) {
const text = document.querySelector(selector)?.textContent || '';
if (text.trim()) return text;
}
return '';
})(),
rating_text:
document.querySelector('#acrPopover')?.getAttribute('title')
|| document.querySelector('#acrPopover')?.textContent
|| '',
review_count_text: document.querySelector('#acrCustomerReviewText')?.textContent || '',
review_url: document.querySelector('a[href*="#customerReviews"]')?.href || '',
qa_url: document.querySelector('a[href*="ask/questions"]')?.href || '',
bullets: Array.from(document.querySelectorAll('#feature-bullets li .a-list-item')).map((node) => node.textContent || ''),
breadcrumbs: Array.from(document.querySelectorAll('#wayfinding-breadcrumbs_feature_div a')).map((node) => node.textContent || ''),
}))()
`) as ProductPayload;
}
cli({
site: 'amazon',
name: 'product',
description: 'Amazon product page facts for candidate validation',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
required: true,
positional: true,
help: 'ASIN or product URL, for example B0FJS72893',
},
],
columns: ['asin', 'title', 'price_text', 'rating_value', 'review_count'],
func: async (page, kwargs) => {
const input = String(kwargs.input ?? '');
const payload = await readProductPayload(page, input);
if (!cleanText(payload.product_title)) {
throw new CommandExecutionError(
'amazon product page did not expose product content',
'The product page may have changed or hit a robot check. Open the product page in Chrome and retry.',
);
}
return [normalizeProductPayload(payload)];
},
});
export const __test__ = {
normalizeProductPayload,
};
+226
View File
@@ -0,0 +1,226 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { Strategy } from '@jackwener/opencli/registry';
import { assertUsableState, buildProvenance, cleanText, extractAsin, extractCategoryNodeId, extractReviewCountFromCardText, firstMeaningfulLine, gotoAndReadState, isRankingPaginationUrl, normalizeProductUrl, parsePriceText, parseRatingValue, parseReviewCount, resolveRankingUrl, toAbsoluteAmazonUrl, uniqueNonEmpty, } from './shared.js';
function parseRank(rawRank, fallback) {
const normalized = cleanText(rawRank);
const match = normalized.match(/(\d{1,4})/);
if (!match)
return fallback;
const parsed = Number.parseInt(match[1], 10);
return Number.isFinite(parsed) && parsed > 0 ? parsed : fallback;
}
function normalizeVisibleCategoryLinks(links) {
const normalized = (links ?? [])
.map((entry) => ({
title: cleanText(entry?.title),
url: toAbsoluteAmazonUrl(entry?.url) ?? '',
node_id: cleanText(entry?.node_id) || extractCategoryNodeId(entry?.url) || null,
}))
.filter((entry) => Boolean(entry.title) && Boolean(entry.url));
const seen = new Set();
const deduped = [];
for (const entry of normalized) {
if (seen.has(entry.url))
continue;
seen.add(entry.url);
deduped.push(entry);
}
return deduped;
}
export function normalizeRankingCandidate(candidate, context) {
const productUrl = normalizeProductUrl(candidate.href);
const asin = extractAsin(candidate.asin ?? '') ?? extractAsin(productUrl ?? '') ?? null;
const title = cleanText(candidate.title) || firstMeaningfulLine(candidate.card_text);
const price = parsePriceText(cleanText(candidate.price_text) || candidate.card_text);
const ratingText = cleanText(candidate.rating_text) || null;
const reviewCountText = cleanText(candidate.review_count_text)
|| extractReviewCountFromCardText(candidate.card_text)
|| null;
const provenance = buildProvenance(context.sourceUrl);
const categoryUrl = context.categoryUrl || context.sourceUrl;
return {
list_type: context.listType,
rank: parseRank(candidate.rank_text, context.rankFallback),
asin,
title: title || null,
product_url: productUrl,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
rating_text: ratingText,
rating_value: parseRatingValue(ratingText),
review_count_text: reviewCountText,
review_count: parseReviewCount(reviewCountText),
list_title: context.listTitle,
category_title: context.categoryTitle,
category_url: categoryUrl,
category_node_id: extractCategoryNodeId(categoryUrl),
category_path: context.categoryPath,
visible_category_links: context.visibleCategoryLinks,
...provenance,
};
}
async function readRankingPage(page, listType, url) {
const state = await gotoAndReadState(page, url, 2500, listType);
assertUsableState(state, listType);
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
list_title:
document.querySelector('#zg_banner_text')?.textContent
|| document.querySelector('h1')?.textContent
|| '',
category_title:
document.querySelector('#zg_browseRoot .zg_selected')?.textContent
|| document.querySelector('#wayfinding-breadcrumbs_feature_div ul li:last-child')?.textContent
|| document.querySelector('#wayfinding-breadcrumbs_container ul li:last-child')?.textContent
|| '',
category_path: Array.from(document.querySelectorAll(
'#zg_browseRoot ul li a, #zg_browseRoot ul li span, ' +
'#wayfinding-breadcrumbs_feature_div ul li a, #wayfinding-breadcrumbs_feature_div ul li span.a-list-item, ' +
'#wayfinding-breadcrumbs_container ul li a, #wayfinding-breadcrumbs_container ul li span.a-list-item'
))
.map((entry) => (entry.textContent || '').trim())
.filter(Boolean),
cards: Array.from(document.querySelectorAll(
'.p13n-sc-uncoverable-faceout, .zg-grid-general-faceout, [data-asin][class*="p13n"]'
)).map((card) => ({
rank_text:
card.querySelector('.zg-bdg-text')?.textContent
|| card.querySelector('[class*="rank"]')?.textContent
|| '',
asin:
card.getAttribute('data-asin')
|| card.getAttribute('id')
|| '',
title:
card.querySelector('[class*="line-clamp"]')?.textContent
|| card.querySelector('img')?.getAttribute('alt')
|| card.querySelector('a[href*="/dp/"]')?.textContent
|| '',
href:
card.querySelector('a[href*="/dp/"], a[href*="/gp/product/"]')?.href
|| '',
price_text:
card.querySelector('.a-price .a-offscreen')?.textContent
|| card.querySelector('.a-color-price')?.textContent
|| '',
rating_text:
card.querySelector('[aria-label*="out of 5 stars"]')?.getAttribute('aria-label')
|| '',
review_count_text:
card.querySelector('a[href*="#customerReviews"]')?.textContent
|| card.querySelector('.a-size-small')?.textContent
|| '',
card_text: card.innerText || '',
})),
page_links: Array.from(document.querySelectorAll('.a-pagination a[href], li.a-normal a[href], li.a-selected a[href]'))
.map((anchor) => anchor.href || '')
.filter(Boolean),
visible_category_links: Array.from(document.querySelectorAll(
'#zg_browseRoot a[href], #zg-left-col a[href], [class*="zg-browse"] a[href]'
)).map((anchor) => ({
title: (anchor.textContent || '').trim(),
url: anchor.href || '',
node_id:
anchor.getAttribute('data-node-id')
|| anchor.dataset?.nodeid
|| '',
}))
.filter((entry) => entry.title && entry.url),
}))()
`);
}
function createEmptyResultHint(commandName) {
return [
`Open the same Amazon ${commandName} page in shared Chrome and verify ranked items are visible.`,
'If the page shows a robot check, clear it manually and retry.',
].join(' ');
}
export function createRankingCliOptions(definition) {
return {
site: 'amazon',
name: definition.commandName,
description: definition.description,
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
positional: true,
help: 'Ranking URL or supported Amazon path. Omit to use the list root.',
},
{
name: 'limit',
type: 'int',
default: 100,
help: 'Maximum number of ranked items to return (default 100)',
},
],
columns: ['list_type', 'rank', 'asin', 'title', 'price_text', 'rating_value', 'review_count'],
func: async (page, kwargs) => {
const limit = Math.max(1, Number(kwargs.limit) || 100);
const initialUrl = resolveRankingUrl(definition.listType, typeof kwargs.input === 'string' ? kwargs.input : undefined);
const queue = [initialUrl];
const visited = new Set();
const seenEntityKeys = new Set();
const results = [];
let listTitle = null;
while (queue.length > 0 && results.length < limit) {
const nextUrl = queue.shift();
if (visited.has(nextUrl))
continue;
visited.add(nextUrl);
const payload = await readRankingPage(page, definition.listType, nextUrl);
const sourceUrl = cleanText(payload.href) || nextUrl;
listTitle = cleanText(payload.list_title) || cleanText(payload.title) || listTitle;
const categoryPath = uniqueNonEmpty(payload.category_path ?? []);
const categoryTitle = cleanText(payload.category_title)
|| (categoryPath.length > 0 ? categoryPath[categoryPath.length - 1] : '');
const visibleCategoryLinks = normalizeVisibleCategoryLinks(payload.visible_category_links);
const cards = payload.cards ?? [];
for (const card of cards) {
const normalized = normalizeRankingCandidate(card, {
listType: definition.listType,
rankFallback: results.length + 1,
listTitle,
sourceUrl,
categoryTitle: categoryTitle || null,
categoryUrl: sourceUrl,
categoryPath,
visibleCategoryLinks,
});
const dedupeKey = cleanText(String(normalized.asin ?? ''))
|| cleanText(String(normalized.product_url ?? ''));
if (dedupeKey && seenEntityKeys.has(dedupeKey))
continue;
if (dedupeKey)
seenEntityKeys.add(dedupeKey);
results.push(normalized);
if (results.length >= limit)
break;
}
const pageLinks = uniqueNonEmpty(payload.page_links ?? []);
for (const href of pageLinks) {
const absolute = toAbsoluteAmazonUrl(href);
if (!absolute || !isRankingPaginationUrl(definition.listType, absolute))
continue;
if (!visited.has(absolute) && !queue.includes(absolute)) {
queue.push(absolute);
}
}
}
if (results.length === 0) {
throw new CommandExecutionError(`amazon ${definition.commandName} did not expose any ranked items`, createEmptyResultHint(definition.commandName));
}
return results.slice(0, limit);
},
};
}
export const __test__ = {
parseRank,
normalizeVisibleCategoryLinks,
normalizeRankingCandidate,
};
+41
View File
@@ -0,0 +1,41 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './rankings.js';
describe('amazon rankings helpers', () => {
it('normalizes ranking candidates with unified schema', () => {
const result = __test__.normalizeRankingCandidate({
rank_text: '#3',
asin: 'B0DR31GC3D',
title: 'Desk Shelves Desktop Organizer',
href: 'https://www.amazon.com/dp/B0DR31GC3D/ref=zg_bs',
price_text: '$25.92',
rating_text: '4.3 out of 5 stars',
review_count_text: '435',
}, {
listType: 'new_releases',
rankFallback: 3,
listTitle: 'Amazon New Releases',
sourceUrl: 'https://www.amazon.com/gp/new-releases',
categoryTitle: 'Home & Kitchen',
categoryUrl: 'https://www.amazon.com/gp/new-releases/home-garden',
categoryPath: ['Home & Kitchen'],
visibleCategoryLinks: [{ title: 'Storage', url: 'https://www.amazon.com/gp/new-releases/storage', node_id: null }],
});
expect(result.list_type).toBe('new_releases');
expect(result.rank).toBe(3);
expect(result.asin).toBe('B0DR31GC3D');
expect(result.product_url).toBe('https://www.amazon.com/dp/B0DR31GC3D');
expect(result.category_title).toBe('Home & Kitchen');
expect(result.visible_category_links).toEqual([
{ title: 'Storage', url: 'https://www.amazon.com/gp/new-releases/storage', node_id: null },
]);
});
it('deduplicates category links and parses rank fallback', () => {
const links = __test__.normalizeVisibleCategoryLinks([
{ title: 'Kitchen', url: '/gp/new-releases/home-garden' },
{ title: 'Kitchen', url: 'https://www.amazon.com/gp/new-releases/home-garden' },
{ title: 'Storage', url: '/gp/new-releases/storage', node_id: '1064954' },
]);
expect(links.length).toBe(2);
expect(__test__.parseRank('N/A', 8)).toBe(8);
});
});
-47
View File
@@ -1,47 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './rankings.js';
describe('amazon rankings helpers', () => {
it('normalizes ranking candidates with unified schema', () => {
const result = __test__.normalizeRankingCandidate(
{
rank_text: '#3',
asin: 'B0DR31GC3D',
title: 'Desk Shelves Desktop Organizer',
href: 'https://www.amazon.com/dp/B0DR31GC3D/ref=zg_bs',
price_text: '$25.92',
rating_text: '4.3 out of 5 stars',
review_count_text: '435',
},
{
listType: 'new_releases',
rankFallback: 3,
listTitle: 'Amazon New Releases',
sourceUrl: 'https://www.amazon.com/gp/new-releases',
categoryTitle: 'Home & Kitchen',
categoryUrl: 'https://www.amazon.com/gp/new-releases/home-garden',
categoryPath: ['Home & Kitchen'],
visibleCategoryLinks: [{ title: 'Storage', url: 'https://www.amazon.com/gp/new-releases/storage', node_id: null }],
},
);
expect(result.list_type).toBe('new_releases');
expect(result.rank).toBe(3);
expect(result.asin).toBe('B0DR31GC3D');
expect(result.product_url).toBe('https://www.amazon.com/dp/B0DR31GC3D');
expect(result.category_title).toBe('Home & Kitchen');
expect(result.visible_category_links).toEqual([
{ title: 'Storage', url: 'https://www.amazon.com/gp/new-releases/storage', node_id: null },
]);
});
it('deduplicates category links and parses rank fallback', () => {
const links = __test__.normalizeVisibleCategoryLinks([
{ title: 'Kitchen', url: '/gp/new-releases/home-garden' },
{ title: 'Kitchen', url: 'https://www.amazon.com/gp/new-releases/home-garden' },
{ title: 'Storage', url: '/gp/new-releases/storage', node_id: '1064954' },
]);
expect(links.length).toBe(2);
expect(__test__.parseRank('N/A', 8)).toBe(8);
});
});
-312
View File
@@ -1,312 +0,0 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { Strategy, type CliOptions } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
assertUsableState,
buildProvenance,
cleanText,
extractAsin,
extractCategoryNodeId,
extractReviewCountFromCardText,
firstMeaningfulLine,
gotoAndReadState,
isRankingPaginationUrl,
normalizeProductUrl,
parsePriceText,
parseRatingValue,
parseReviewCount,
resolveRankingUrl,
toAbsoluteAmazonUrl,
uniqueNonEmpty,
type AmazonRankingListType,
} from './shared.js';
export interface RankingCardPayload {
rank_text?: string | null;
asin?: string | null;
title?: string | null;
href?: string | null;
price_text?: string | null;
rating_text?: string | null;
review_count_text?: string | null;
card_text?: string | null;
}
interface RankingPagePayload {
href?: string;
title?: string;
list_title?: string;
category_title?: string;
category_path?: string[];
cards?: RankingCardPayload[];
page_links?: string[];
visible_category_links?: Array<{
title?: string | null;
url?: string | null;
node_id?: string | null;
}>;
}
interface RankingCommandDefinition {
commandName: string;
listType: AmazonRankingListType;
description: string;
}
interface RankingNormalizeContext {
listType: AmazonRankingListType;
rankFallback: number;
listTitle: string | null;
sourceUrl: string;
categoryTitle: string | null;
categoryUrl: string | null;
categoryPath: string[];
visibleCategoryLinks: Array<{ title: string; url: string; node_id: string | null }>;
}
function parseRank(rawRank: string | null | undefined, fallback: number): number {
const normalized = cleanText(rawRank);
const match = normalized.match(/(\d{1,4})/);
if (!match) return fallback;
const parsed = Number.parseInt(match[1], 10);
return Number.isFinite(parsed) && parsed > 0 ? parsed : fallback;
}
function normalizeVisibleCategoryLinks(
links: RankingPagePayload['visible_category_links'],
): Array<{ title: string; url: string; node_id: string | null }> {
const normalized = (links ?? [])
.map((entry) => ({
title: cleanText(entry?.title),
url: toAbsoluteAmazonUrl(entry?.url) ?? '',
node_id: cleanText(entry?.node_id) || extractCategoryNodeId(entry?.url) || null,
}))
.filter((entry) => Boolean(entry.title) && Boolean(entry.url));
const seen = new Set<string>();
const deduped: Array<{ title: string; url: string; node_id: string | null }> = [];
for (const entry of normalized) {
if (seen.has(entry.url)) continue;
seen.add(entry.url);
deduped.push(entry);
}
return deduped;
}
export function normalizeRankingCandidate(
candidate: RankingCardPayload,
context: RankingNormalizeContext,
): Record<string, unknown> {
const productUrl = normalizeProductUrl(candidate.href);
const asin = extractAsin(candidate.asin ?? '') ?? extractAsin(productUrl ?? '') ?? null;
const title = cleanText(candidate.title) || firstMeaningfulLine(candidate.card_text);
const price = parsePriceText(cleanText(candidate.price_text) || candidate.card_text);
const ratingText = cleanText(candidate.rating_text) || null;
const reviewCountText = cleanText(candidate.review_count_text)
|| extractReviewCountFromCardText(candidate.card_text)
|| null;
const provenance = buildProvenance(context.sourceUrl);
const categoryUrl = context.categoryUrl || context.sourceUrl;
return {
list_type: context.listType,
rank: parseRank(candidate.rank_text, context.rankFallback),
asin,
title: title || null,
product_url: productUrl,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
rating_text: ratingText,
rating_value: parseRatingValue(ratingText),
review_count_text: reviewCountText,
review_count: parseReviewCount(reviewCountText),
list_title: context.listTitle,
category_title: context.categoryTitle,
category_url: categoryUrl,
category_node_id: extractCategoryNodeId(categoryUrl),
category_path: context.categoryPath,
visible_category_links: context.visibleCategoryLinks,
...provenance,
};
}
async function readRankingPage(
page: IPage,
listType: AmazonRankingListType,
url: string,
): Promise<RankingPagePayload> {
const state = await gotoAndReadState(page, url, 2500, listType);
assertUsableState(state, listType);
return await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
list_title:
document.querySelector('#zg_banner_text')?.textContent
|| document.querySelector('h1')?.textContent
|| '',
category_title:
document.querySelector('#zg_browseRoot .zg_selected')?.textContent
|| document.querySelector('#wayfinding-breadcrumbs_feature_div ul li:last-child')?.textContent
|| document.querySelector('#wayfinding-breadcrumbs_container ul li:last-child')?.textContent
|| '',
category_path: Array.from(document.querySelectorAll(
'#zg_browseRoot ul li a, #zg_browseRoot ul li span, ' +
'#wayfinding-breadcrumbs_feature_div ul li a, #wayfinding-breadcrumbs_feature_div ul li span.a-list-item, ' +
'#wayfinding-breadcrumbs_container ul li a, #wayfinding-breadcrumbs_container ul li span.a-list-item'
))
.map((entry) => (entry.textContent || '').trim())
.filter(Boolean),
cards: Array.from(document.querySelectorAll(
'.p13n-sc-uncoverable-faceout, .zg-grid-general-faceout, [data-asin][class*="p13n"]'
)).map((card) => ({
rank_text:
card.querySelector('.zg-bdg-text')?.textContent
|| card.querySelector('[class*="rank"]')?.textContent
|| '',
asin:
card.getAttribute('data-asin')
|| card.getAttribute('id')
|| '',
title:
card.querySelector('[class*="line-clamp"]')?.textContent
|| card.querySelector('img')?.getAttribute('alt')
|| card.querySelector('a[href*="/dp/"]')?.textContent
|| '',
href:
card.querySelector('a[href*="/dp/"], a[href*="/gp/product/"]')?.href
|| '',
price_text:
card.querySelector('.a-price .a-offscreen')?.textContent
|| card.querySelector('.a-color-price')?.textContent
|| '',
rating_text:
card.querySelector('[aria-label*="out of 5 stars"]')?.getAttribute('aria-label')
|| '',
review_count_text:
card.querySelector('a[href*="#customerReviews"]')?.textContent
|| card.querySelector('.a-size-small')?.textContent
|| '',
card_text: card.innerText || '',
})),
page_links: Array.from(document.querySelectorAll('.a-pagination a[href], li.a-normal a[href], li.a-selected a[href]'))
.map((anchor) => anchor.href || '')
.filter(Boolean),
visible_category_links: Array.from(document.querySelectorAll(
'#zg_browseRoot a[href], #zg-left-col a[href], [class*="zg-browse"] a[href]'
)).map((anchor) => ({
title: (anchor.textContent || '').trim(),
url: anchor.href || '',
node_id:
anchor.getAttribute('data-node-id')
|| anchor.dataset?.nodeid
|| '',
}))
.filter((entry) => entry.title && entry.url),
}))()
`) as RankingPagePayload;
}
function createEmptyResultHint(commandName: string): string {
return [
`Open the same Amazon ${commandName} page in shared Chrome and verify ranked items are visible.`,
'If the page shows a robot check, clear it manually and retry.',
].join(' ');
}
export function createRankingCliOptions(definition: RankingCommandDefinition): CliOptions {
return {
site: 'amazon',
name: definition.commandName,
description: definition.description,
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'input',
positional: true,
help: 'Ranking URL or supported Amazon path. Omit to use the list root.',
},
{
name: 'limit',
type: 'int',
default: 100,
help: 'Maximum number of ranked items to return (default 100)',
},
],
columns: ['list_type', 'rank', 'asin', 'title', 'price_text', 'rating_value', 'review_count'],
func: async (page, kwargs) => {
const limit = Math.max(1, Number(kwargs.limit) || 100);
const initialUrl = resolveRankingUrl(definition.listType, typeof kwargs.input === 'string' ? kwargs.input : undefined);
const queue = [initialUrl];
const visited = new Set<string>();
const seenEntityKeys = new Set<string>();
const results: Record<string, unknown>[] = [];
let listTitle: string | null = null;
while (queue.length > 0 && results.length < limit) {
const nextUrl = queue.shift()!;
if (visited.has(nextUrl)) continue;
visited.add(nextUrl);
const payload = await readRankingPage(page, definition.listType, nextUrl);
const sourceUrl = cleanText(payload.href) || nextUrl;
listTitle = cleanText(payload.list_title) || cleanText(payload.title) || listTitle;
const categoryPath = uniqueNonEmpty(payload.category_path ?? []);
const categoryTitle = cleanText(payload.category_title)
|| (categoryPath.length > 0 ? categoryPath[categoryPath.length - 1] : '');
const visibleCategoryLinks = normalizeVisibleCategoryLinks(payload.visible_category_links);
const cards = payload.cards ?? [];
for (const card of cards) {
const normalized = normalizeRankingCandidate(card, {
listType: definition.listType,
rankFallback: results.length + 1,
listTitle,
sourceUrl,
categoryTitle: categoryTitle || null,
categoryUrl: sourceUrl,
categoryPath,
visibleCategoryLinks,
});
const dedupeKey = cleanText(String(normalized.asin ?? ''))
|| cleanText(String(normalized.product_url ?? ''));
if (dedupeKey && seenEntityKeys.has(dedupeKey)) continue;
if (dedupeKey) seenEntityKeys.add(dedupeKey);
results.push(normalized);
if (results.length >= limit) break;
}
const pageLinks = uniqueNonEmpty(payload.page_links ?? []);
for (const href of pageLinks) {
const absolute = toAbsoluteAmazonUrl(href);
if (!absolute || !isRankingPaginationUrl(definition.listType, absolute)) continue;
if (!visited.has(absolute) && !queue.includes(absolute)) {
queue.push(absolute);
}
}
}
if (results.length === 0) {
throw new CommandExecutionError(
`amazon ${definition.commandName} did not expose any ranked items`,
createEmptyResultHint(definition.commandName),
);
}
return results.slice(0, limit);
},
};
}
export const __test__ = {
parseRank,
normalizeVisibleCategoryLinks,
normalizeRankingCandidate,
};
+87
View File
@@ -0,0 +1,87 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import { buildProvenance, buildSearchUrl, cleanText, extractAsin, normalizeProductUrl, parsePriceText, parseRatingValue, parseReviewCount, assertUsableState, gotoAndReadState, } from './shared.js';
function normalizeSearchCandidate(candidate, rank, sourceUrl) {
const productUrl = normalizeProductUrl(candidate.href);
const asin = extractAsin(candidate.asin ?? '') ?? extractAsin(productUrl ?? '') ?? null;
const price = parsePriceText(candidate.price_text);
const ratingText = cleanText(candidate.rating_text) || null;
const reviewCountText = cleanText(candidate.review_count_text) || null;
const provenance = buildProvenance(sourceUrl);
return {
rank,
asin,
title: cleanText(candidate.title) || null,
product_url: productUrl,
...provenance,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
rating_text: ratingText,
rating_value: parseRatingValue(ratingText),
review_count_text: reviewCountText,
review_count: parseReviewCount(reviewCountText),
is_sponsored: candidate.sponsored === true,
badges: (candidate.badge_texts ?? []).map((value) => cleanText(value)).filter(Boolean),
};
}
async function readSearchPayload(page, query) {
const url = buildSearchUrl(query);
const state = await gotoAndReadState(page, url, 2500, 'search');
assertUsableState(state, 'search');
return await page.evaluate(`
(() => ({
href: window.location.href,
cards: Array.from(document.querySelectorAll('[data-component-type="s-search-result"]'))
.map((card) => ({
asin: card.getAttribute('data-asin') || '',
title: card.querySelector('h2')?.textContent || '',
href: card.querySelector('a.a-link-normal[href*="/dp/"]')?.href || '',
price_text: card.querySelector('.a-price .a-offscreen')?.textContent || '',
rating_text: card.querySelector('[aria-label*="out of 5 stars"]')?.getAttribute('aria-label') || '',
review_count_text: card.querySelector('a[href*="#customerReviews"]')?.textContent || '',
sponsored: /sponsored/i.test(card.innerText || ''),
badge_texts: Array.from(card.querySelectorAll('.a-badge-text')).map((node) => node.textContent || ''),
})),
}))()
`);
}
cli({
site: 'amazon',
name: 'search',
description: 'Amazon search results for product discovery and coarse filtering',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'query',
required: true,
positional: true,
help: 'Search query, for example "desk shelf organizer"',
},
{
name: 'limit',
type: 'int',
default: 20,
help: 'Maximum number of results to return (default 20)',
},
],
columns: ['rank', 'asin', 'title', 'price_text', 'rating_value', 'review_count'],
func: async (page, kwargs) => {
const query = String(kwargs.query ?? '');
const limit = Math.max(1, Number(kwargs.limit) || 20);
const payload = await readSearchPayload(page, query);
const sourceUrl = cleanText(payload.href) || buildSearchUrl(query);
const cards = (payload.cards ?? [])
.filter((card) => cleanText(card.asin) && cleanText(card.title))
.slice(0, limit);
if (cards.length === 0) {
throw new CommandExecutionError('amazon search did not expose any product cards', 'The search page may have changed or hit a robot check. Open the same query in Chrome, verify the page is visible, and retry.');
}
return cards.map((card, index) => normalizeSearchCandidate(card, index + 1, sourceUrl));
},
});
export const __test__ = {
normalizeSearchCandidate,
};
+22
View File
@@ -0,0 +1,22 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './search.js';
describe('amazon search normalization', () => {
it('normalizes search cards into research-friendly fields', () => {
const result = __test__.normalizeSearchCandidate({
asin: 'B0FJS72893',
title: 'White Desktop Shelf Organizer for Top of Desk',
href: 'https://www.amazon.com/KVTUKIAIT-White-Desktop-Shelf-Organizer/dp/B0FJS72893/ref=sr_1_1',
price_text: '$15.99',
rating_text: '3.9 out of 5 stars, rating details',
review_count_text: '(27)',
sponsored: false,
badge_texts: ['Limited time deal'],
}, 1, 'https://www.amazon.com/s?k=desk+shelf+organizer');
expect(result.asin).toBe('B0FJS72893');
expect(result.product_url).toBe('https://www.amazon.com/dp/B0FJS72893');
expect(result.price_value).toBe(15.99);
expect(result.rating_value).toBe(3.9);
expect(result.review_count).toBe(27);
expect(result.badges).toEqual(['Limited time deal']);
});
});
-24
View File
@@ -1,24 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './search.js';
describe('amazon search normalization', () => {
it('normalizes search cards into research-friendly fields', () => {
const result = __test__.normalizeSearchCandidate({
asin: 'B0FJS72893',
title: 'White Desktop Shelf Organizer for Top of Desk',
href: 'https://www.amazon.com/KVTUKIAIT-White-Desktop-Shelf-Organizer/dp/B0FJS72893/ref=sr_1_1',
price_text: '$15.99',
rating_text: '3.9 out of 5 stars, rating details',
review_count_text: '(27)',
sponsored: false,
badge_texts: ['Limited time deal'],
}, 1, 'https://www.amazon.com/s?k=desk+shelf+organizer');
expect(result.asin).toBe('B0FJS72893');
expect(result.product_url).toBe('https://www.amazon.com/dp/B0FJS72893');
expect(result.price_value).toBe(15.99);
expect(result.rating_value).toBe(3.9);
expect(result.review_count).toBe(27);
expect(result.badges).toEqual(['Limited time deal']);
});
});
-128
View File
@@ -1,128 +0,0 @@
import { CommandExecutionError } from '@jackwener/opencli/errors';
import { cli, Strategy } from '@jackwener/opencli/registry';
import type { IPage } from '@jackwener/opencli/types';
import {
buildProvenance,
buildSearchUrl,
cleanText,
extractAsin,
normalizeProductUrl,
parsePriceText,
parseRatingValue,
parseReviewCount,
assertUsableState,
gotoAndReadState,
} from './shared.js';
interface SearchPayload {
href?: string;
cards?: Array<{
asin?: string;
title?: string;
href?: string;
price_text?: string | null;
rating_text?: string | null;
review_count_text?: string | null;
sponsored?: boolean;
badge_texts?: string[];
}>;
}
function normalizeSearchCandidate(
candidate: NonNullable<SearchPayload['cards']>[number],
rank: number,
sourceUrl: string,
): Record<string, unknown> {
const productUrl = normalizeProductUrl(candidate.href);
const asin = extractAsin(candidate.asin ?? '') ?? extractAsin(productUrl ?? '') ?? null;
const price = parsePriceText(candidate.price_text);
const ratingText = cleanText(candidate.rating_text) || null;
const reviewCountText = cleanText(candidate.review_count_text) || null;
const provenance = buildProvenance(sourceUrl);
return {
rank,
asin,
title: cleanText(candidate.title) || null,
product_url: productUrl,
...provenance,
price_text: price.price_text,
price_value: price.price_value,
currency: price.currency,
rating_text: ratingText,
rating_value: parseRatingValue(ratingText),
review_count_text: reviewCountText,
review_count: parseReviewCount(reviewCountText),
is_sponsored: candidate.sponsored === true,
badges: (candidate.badge_texts ?? []).map((value) => cleanText(value)).filter(Boolean),
};
}
async function readSearchPayload(page: IPage, query: string): Promise<SearchPayload> {
const url = buildSearchUrl(query);
const state = await gotoAndReadState(page, url, 2500, 'search');
assertUsableState(state, 'search');
return await page.evaluate(`
(() => ({
href: window.location.href,
cards: Array.from(document.querySelectorAll('[data-component-type="s-search-result"]'))
.map((card) => ({
asin: card.getAttribute('data-asin') || '',
title: card.querySelector('h2')?.textContent || '',
href: card.querySelector('a.a-link-normal[href*="/dp/"]')?.href || '',
price_text: card.querySelector('.a-price .a-offscreen')?.textContent || '',
rating_text: card.querySelector('[aria-label*="out of 5 stars"]')?.getAttribute('aria-label') || '',
review_count_text: card.querySelector('a[href*="#customerReviews"]')?.textContent || '',
sponsored: /sponsored/i.test(card.innerText || ''),
badge_texts: Array.from(card.querySelectorAll('.a-badge-text')).map((node) => node.textContent || ''),
})),
}))()
`) as SearchPayload;
}
cli({
site: 'amazon',
name: 'search',
description: 'Amazon search results for product discovery and coarse filtering',
domain: 'amazon.com',
strategy: Strategy.COOKIE,
navigateBefore: false,
args: [
{
name: 'query',
required: true,
positional: true,
help: 'Search query, for example "desk shelf organizer"',
},
{
name: 'limit',
type: 'int',
default: 20,
help: 'Maximum number of results to return (default 20)',
},
],
columns: ['rank', 'asin', 'title', 'price_text', 'rating_value', 'review_count'],
func: async (page, kwargs) => {
const query = String(kwargs.query ?? '');
const limit = Math.max(1, Number(kwargs.limit) || 20);
const payload = await readSearchPayload(page, query);
const sourceUrl = cleanText(payload.href) || buildSearchUrl(query);
const cards = (payload.cards ?? [])
.filter((card) => cleanText(card.asin) && cleanText(card.title))
.slice(0, limit);
if (cards.length === 0) {
throw new CommandExecutionError(
'amazon search did not expose any product cards',
'The search page may have changed or hit a robot check. Open the same query in Chrome, verify the page is visible, and retry.',
);
}
return cards.map((card, index) => normalizeSearchCandidate(card, index + 1, sourceUrl));
},
});
export const __test__ = {
normalizeSearchCandidate,
};
+365
View File
@@ -0,0 +1,365 @@
import { ArgumentError, CommandExecutionError } from '@jackwener/opencli/errors';
export const SITE = 'amazon';
export const DOMAIN = 'amazon.com';
export const HOME_URL = 'https://www.amazon.com/';
export const BESTSELLERS_URL = 'https://www.amazon.com/Best-Sellers/zgbs';
export const NEW_RELEASES_URL = 'https://www.amazon.com/gp/new-releases';
export const MOVERS_SHAKERS_URL = 'https://www.amazon.com/gp/movers-and-shakers';
export const SEARCH_URL_PREFIX = 'https://www.amazon.com/s?k=';
export const PRODUCT_URL_PREFIX = 'https://www.amazon.com/dp/';
export const DISCUSSION_URL_PREFIX = 'https://www.amazon.com/product-reviews/';
export const STRATEGY = 'cookie';
export const PRIMARY_PRICE_SELECTORS = [
'#corePrice_feature_div .a-offscreen',
'#corePriceDisplay_desktop_feature_div .a-offscreen',
'#corePrice_desktop .a-offscreen',
'#apex_desktop .a-offscreen',
'#newAccordionRow_0 .a-offscreen',
'#price_inside_buybox',
'#priceblock_ourprice',
'#priceblock_dealprice',
'#tp_price_block_total_price_ww',
];
const ROBOT_TEXT_PATTERNS = [
'Sorry, we just need to make sure you\'re not a robot',
'Enter the characters you see below',
'Type the characters you see in this image',
'To discuss automated access to Amazon data please contact',
];
const AMAZON_RANKING_SPECS = {
bestsellers: {
commandName: 'bestsellers',
rootUrl: BESTSELLERS_URL,
pathPattern: /(?:^|\/)zgbs(?:\/|$)/i,
invalidInputMessage: 'amazon bestsellers expects a best sellers URL or /zgbs path',
invalidInputHint: 'Example: opencli amazon bestsellers https://www.amazon.com/Best-Sellers/zgbs',
},
new_releases: {
commandName: 'new-releases',
rootUrl: NEW_RELEASES_URL,
pathPattern: /\/gp\/new-releases(?:\/|$)/i,
invalidInputMessage: 'amazon new-releases expects a new releases URL or /gp/new-releases path',
invalidInputHint: 'Example: opencli amazon new-releases https://www.amazon.com/gp/new-releases',
},
movers_shakers: {
commandName: 'movers-shakers',
rootUrl: MOVERS_SHAKERS_URL,
pathPattern: /\/gp\/movers-and-shakers(?:\/|$)/i,
invalidInputMessage: 'amazon movers-shakers expects a movers-and-shakers URL or /gp/movers-and-shakers path',
invalidInputHint: 'Example: opencli amazon movers-shakers https://www.amazon.com/gp/movers-and-shakers',
},
};
export function cleanText(value) {
return typeof value === 'string'
? value.replace(/\u00a0/g, ' ').replace(/\s+/g, ' ').trim()
: '';
}
export function cleanMultilineText(value) {
return typeof value === 'string'
? value
.replace(/\u00a0/g, ' ')
.split('\n')
.map((line) => line.replace(/\s+/g, ' ').trim())
.filter(Boolean)
.join('\n')
: '';
}
export function uniqueNonEmpty(values) {
return [...new Set(values.map((value) => cleanText(value)).filter(Boolean))];
}
export function buildProvenance(sourceUrl) {
return {
source_url: sourceUrl,
fetched_at: new Date().toISOString(),
strategy: STRATEGY,
};
}
export function buildSearchUrl(query) {
const normalized = cleanText(query);
if (!normalized) {
throw new ArgumentError('amazon search query cannot be empty');
}
return `${SEARCH_URL_PREFIX}${encodeURIComponent(normalized)}`;
}
export function extractAsin(input) {
const normalized = cleanText(input);
if (!normalized)
return null;
if (/^[A-Z0-9]{10}$/i.test(normalized)) {
return normalized.toUpperCase();
}
const match = normalized.match(/\/(?:dp|gp\/product|product-reviews)\/([A-Z0-9]{10})/i);
return match ? match[1].toUpperCase() : null;
}
export function buildProductUrl(input) {
const asin = extractAsin(input);
if (!asin) {
throw new ArgumentError('amazon product expects an ASIN or product URL', 'Example: opencli amazon product B0FJS72893');
}
return `${PRODUCT_URL_PREFIX}${asin}`;
}
export function buildDiscussionUrl(input) {
const asin = extractAsin(input);
if (!asin) {
throw new ArgumentError('amazon discussion expects an ASIN or product URL', 'Example: opencli amazon discussion B0FJS72893');
}
return `${DISCUSSION_URL_PREFIX}${asin}`;
}
function getRankingSpec(listType) {
return AMAZON_RANKING_SPECS[listType];
}
export function isSupportedRankingPath(listType, inputUrl) {
try {
const url = new URL(inputUrl);
return getRankingSpec(listType).pathPattern.test(url.pathname);
}
catch {
return false;
}
}
export function resolveRankingUrl(listType, input) {
const spec = getRankingSpec(listType);
const normalized = cleanText(input);
if (!normalized || normalized === 'root')
return spec.rootUrl;
let candidateUrl;
if (normalized.startsWith('/')) {
candidateUrl = new URL(normalized, HOME_URL).toString();
}
else if (/^https?:\/\//i.test(normalized)) {
candidateUrl = canonicalizeAmazonUrl(normalized);
}
else if (normalized.includes('amazon.') && normalized.includes('/')) {
candidateUrl = canonicalizeAmazonUrl(`https://${normalized.replace(/^\/+/, '')}`);
}
else {
throw new ArgumentError(spec.invalidInputMessage, spec.invalidInputHint);
}
if (!isSupportedRankingPath(listType, candidateUrl)) {
throw new ArgumentError(spec.invalidInputMessage, spec.invalidInputHint);
}
return normalizeRankingInputUrl(candidateUrl);
}
function normalizeRankingInputUrl(inputUrl) {
try {
const url = new URL(inputUrl);
const normalizedPathSegments = url.pathname
.split('/')
.filter(Boolean)
.filter((segment) => !/^ref=/i.test(segment));
url.pathname = `/${normalizedPathSegments.join('/')}`;
url.hash = '';
// Ranking pages are frequently shared with tracking refs that can land on unstable variants.
// Dropping ref keeps the canonical ranking path while preserving useful params (for example pg=2).
url.searchParams.delete('ref');
return url.toString();
}
catch {
return inputUrl;
}
}
export function isRankingPaginationUrl(listType, inputUrl) {
const absolute = toAbsoluteAmazonUrl(inputUrl);
if (!absolute || !isSupportedRankingPath(listType, absolute))
return false;
try {
const url = new URL(absolute);
const ref = cleanText(url.searchParams.get('ref')).toLowerCase();
// pg= query param is the most reliable pagination indicator across all ranking lists
return url.searchParams.has('pg')
|| /(?:^|_)pg(?:_|$)/.test(ref)
// Amazon ranking pagination refs: zg_bs_pg_ (bestsellers), zg_bsnr_pg_ (new releases), zg_bsms_pg_ (movers & shakers)
|| /zg_bs(?:nr|ms)?_pg_/.test(ref);
}
catch {
return false;
}
}
export function extractCategoryNodeId(inputUrl) {
const absolute = toAbsoluteAmazonUrl(inputUrl);
if (!absolute)
return null;
try {
const url = new URL(absolute);
for (const key of ['node', 'nodeid', 'nodeId', 'browseNode']) {
const value = cleanText(url.searchParams.get(key));
if (/^\d{4,}$/.test(value))
return value;
}
const rhValue = cleanText(url.searchParams.get('rh'));
const rhMatch = decodeURIComponent(rhValue).match(/(?:^|,)\s*n:(\d{4,})(?:,|$)/i);
if (rhMatch)
return rhMatch[1];
const pathMatches = [...url.pathname.matchAll(/\/(\d{4,})(?=\/|$)/g)];
if (pathMatches.length > 0) {
return pathMatches[pathMatches.length - 1][1];
}
}
catch {
return null;
}
return null;
}
export function resolveBestsellersUrl(input) {
return resolveRankingUrl('bestsellers', input);
}
export function canonicalizeAmazonUrl(input) {
try {
const url = new URL(input);
if (!url.hostname.endsWith(DOMAIN)) {
throw new Error('not-amazon');
}
return url.toString();
}
catch {
throw new ArgumentError('Invalid Amazon URL');
}
}
export function toAbsoluteAmazonUrl(value) {
const normalized = cleanText(value);
if (!normalized)
return null;
try {
return new URL(normalized, HOME_URL).toString();
}
catch {
return null;
}
}
export function normalizeProductUrl(value) {
const normalized = cleanText(value);
const asin = extractAsin(normalized);
if (asin)
return buildProductUrl(asin);
return toAbsoluteAmazonUrl(normalized);
}
export function parsePriceText(text) {
const normalized = cleanText(text);
const match = normalized.match(/([$€£])\s*(\d+(?:,\d{3})*(?:\.\d+)?)/);
if (!match) {
return {
price_text: normalized || null,
price_value: null,
currency: null,
};
}
const currencyMap = {
'$': 'USD',
'€': 'EUR',
'£': 'GBP',
};
return {
price_text: `${match[1]}${match[2]}`,
price_value: Number.parseFloat(match[2].replace(/,/g, '')),
currency: currencyMap[match[1]] ?? null,
};
}
export function parseRatingValue(text) {
const normalized = cleanText(text);
const match = normalized.match(/(\d+(?:\.\d+)?)\s*out of 5/i);
return match ? Number.parseFloat(match[1]) : null;
}
export function parseReviewCount(text) {
const normalized = cleanText(text);
const compactMatch = normalized.match(/(\d+(?:\.\d+)?)\s*([kKmM])/);
if (compactMatch) {
const value = Number.parseFloat(compactMatch[1]);
const multiplier = /m/i.test(compactMatch[2]) ? 1_000_000 : 1_000;
return Number.isFinite(value) ? Math.round(value * multiplier) : null;
}
const match = normalized.match(/([\d,]+)/);
return match ? Number.parseInt(match[1].replace(/,/g, ''), 10) : null;
}
export function extractReviewCountFromCardText(text) {
const normalized = cleanMultilineText(text);
const match = normalized.match(/out of 5 stars(?:, rating details)?\s*([\d,]+)/i);
if (match)
return match[1];
const numericLine = normalized
.split('\n')
.map((line) => cleanText(line))
.find((line) => /^[\d,]+$/.test(line));
return numericLine ?? null;
}
export function isAmazonEntity(text) {
const normalized = cleanText(text).toLowerCase();
return normalized.includes('amazon');
}
export function firstMeaningfulLine(text) {
return cleanMultilineText(text)
.split('\n')
.map((line) => cleanText(line))
.find(Boolean)
?? '';
}
export function trimRatingPrefix(text) {
const normalized = cleanText(text);
if (!normalized)
return null;
return normalized.replace(/^\d+(?:\.\d+)?\s*out of 5 stars\s*/i, '').trim() || normalized;
}
export function isRobotState(state) {
const title = cleanText(state.title);
const bodyText = cleanMultilineText(state.body_text);
return ROBOT_TEXT_PATTERNS.some((pattern) => title.includes(pattern) || bodyText.includes(pattern));
}
export function buildChallengeHint(action) {
return [
`Open a clean Amazon ${action} page in the shared Chrome profile and clear any robot check first.`,
'If you are using CDP, set OPENCLI_CDP_TARGET=amazon.com and avoid parallel Amazon commands against the same browser target.',
].join(' ');
}
export async function readPageState(page) {
const result = await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
body_text: document.body ? document.body.innerText || '' : '',
}))()
`);
return {
href: cleanText(result.href),
title: cleanText(result.title),
body_text: cleanMultilineText(result.body_text),
};
}
export async function gotoAndReadState(page, url, settleMs = 2500, action = 'page') {
try {
await page.goto(url, { settleMs });
await page.wait(1.5);
return await readPageState(page);
}
catch (error) {
const message = error instanceof Error ? error.message : String(error);
if (message.includes('Inspected target navigated or closed')
|| message.includes('Cannot find context with specified id')
|| message.includes('Target closed')) {
throw new CommandExecutionError(`amazon ${action} navigation lost the current browser target`, `${buildChallengeHint(action)} If CDP is attached to a stale tab, open a fresh Amazon tab and retry.`);
}
throw error;
}
}
export function assertUsableState(state, action) {
if (!isRobotState(state))
return;
throw new CommandExecutionError(`amazon ${action} hit a robot check`, buildChallengeHint(action));
}
export const __test__ = {
buildSearchUrl,
extractAsin,
buildProductUrl,
buildDiscussionUrl,
resolveBestsellersUrl,
resolveRankingUrl,
isSupportedRankingPath,
isRankingPaginationUrl,
extractCategoryNodeId,
parsePriceText,
parseRatingValue,
parseReviewCount,
extractReviewCountFromCardText,
isAmazonEntity,
trimRatingPrefix,
isRobotState,
PRIMARY_PRICE_SELECTORS,
};
+44
View File
@@ -0,0 +1,44 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './shared.js';
describe('amazon shared helpers', () => {
it('builds canonical product and discussion URLs from ASINs and product URLs', () => {
expect(__test__.buildProductUrl('B0FJS72893')).toBe('https://www.amazon.com/dp/B0FJS72893');
expect(__test__.buildProductUrl('https://www.amazon.com/dp/B0FJS72893/ref=something')).toBe('https://www.amazon.com/dp/B0FJS72893');
expect(__test__.buildDiscussionUrl('https://www.amazon.com/dp/B0FJS72893')).toBe('https://www.amazon.com/product-reviews/B0FJS72893');
});
it('parses price, rating, and review-count text', () => {
expect(__test__.parsePriceText('1 offer from $34.11')).toEqual({
price_text: '$34.11',
price_value: 34.11,
currency: 'USD',
});
expect(__test__.parseRatingValue('3.9 out of 5 stars, rating details')).toBe(3.9);
expect(__test__.parseReviewCount('27 global ratings')).toBe(27);
expect(__test__.parseReviewCount('(2.9K)')).toBe(2900);
expect(__test__.parseReviewCount('1.2M global ratings')).toBe(1200000);
expect(__test__.extractReviewCountFromCardText('Desk Shelf\n4.3 out of 5 stars\n435\n$25.92')).toBe('435');
});
it('recognizes robot checks and Amazon-owned merchants', () => {
expect(__test__.isAmazonEntity('Ships from Amazon')).toBe(true);
expect(__test__.trimRatingPrefix('5.0 out of 5 stars Great value and quality')).toBe('Great value and quality');
expect(__test__.isRobotState({
title: 'Robot Check',
body_text: 'Sorry, we just need to make sure you\'re not a robot',
})).toBe(true);
});
it('requires a real best-sellers URL or path', () => {
expect(__test__.resolveBestsellersUrl('/Best-Sellers/zgbs')).toBe('https://www.amazon.com/Best-Sellers/zgbs');
expect(() => __test__.resolveBestsellersUrl('desk shelf organizer')).toThrow('amazon bestsellers expects a best sellers URL or /zgbs path');
});
it('resolves and validates all ranking list URLs', () => {
expect(__test__.resolveRankingUrl('new_releases')).toBe('https://www.amazon.com/gp/new-releases');
expect(__test__.resolveRankingUrl('movers_shakers')).toBe('https://www.amazon.com/gp/movers-and-shakers');
expect(__test__.resolveRankingUrl('new_releases', '/gp/new-releases/kitchen')).toBe('https://www.amazon.com/gp/new-releases/kitchen');
expect(__test__.resolveRankingUrl('bestsellers', 'https://www.amazon.com/Best-Sellers/zgbs/ref=zg_bsnr_tab_bs')).toBe('https://www.amazon.com/Best-Sellers/zgbs');
expect(() => __test__.resolveRankingUrl('movers_shakers', 'https://example.com/gp/movers-and-shakers')).toThrow('Invalid Amazon URL');
});
it('extracts category node id from URL best effort', () => {
expect(__test__.extractCategoryNodeId('https://www.amazon.com/Best-Sellers-Home-Kitchen/zgbs/home-garden/3744371')).toBe('3744371');
expect(__test__.extractCategoryNodeId('https://www.amazon.com/s?k=desk+organizer&rh=n%3A1064954')).toBe('1064954');
});
});
-53
View File
@@ -1,53 +0,0 @@
import { describe, expect, it } from 'vitest';
import { __test__ } from './shared.js';
describe('amazon shared helpers', () => {
it('builds canonical product and discussion URLs from ASINs and product URLs', () => {
expect(__test__.buildProductUrl('B0FJS72893')).toBe('https://www.amazon.com/dp/B0FJS72893');
expect(__test__.buildProductUrl('https://www.amazon.com/dp/B0FJS72893/ref=something')).toBe('https://www.amazon.com/dp/B0FJS72893');
expect(__test__.buildDiscussionUrl('https://www.amazon.com/dp/B0FJS72893')).toBe('https://www.amazon.com/product-reviews/B0FJS72893');
});
it('parses price, rating, and review-count text', () => {
expect(__test__.parsePriceText('1 offer from $34.11')).toEqual({
price_text: '$34.11',
price_value: 34.11,
currency: 'USD',
});
expect(__test__.parseRatingValue('3.9 out of 5 stars, rating details')).toBe(3.9);
expect(__test__.parseReviewCount('27 global ratings')).toBe(27);
expect(__test__.parseReviewCount('(2.9K)')).toBe(2900);
expect(__test__.parseReviewCount('1.2M global ratings')).toBe(1200000);
expect(__test__.extractReviewCountFromCardText('Desk Shelf\n4.3 out of 5 stars\n435\n$25.92')).toBe('435');
});
it('recognizes robot checks and Amazon-owned merchants', () => {
expect(__test__.isAmazonEntity('Ships from Amazon')).toBe(true);
expect(__test__.trimRatingPrefix('5.0 out of 5 stars Great value and quality')).toBe('Great value and quality');
expect(__test__.isRobotState({
title: 'Robot Check',
body_text: 'Sorry, we just need to make sure you\'re not a robot',
})).toBe(true);
});
it('requires a real best-sellers URL or path', () => {
expect(__test__.resolveBestsellersUrl('/Best-Sellers/zgbs')).toBe('https://www.amazon.com/Best-Sellers/zgbs');
expect(() => __test__.resolveBestsellersUrl('desk shelf organizer')).toThrow('amazon bestsellers expects a best sellers URL or /zgbs path');
});
it('resolves and validates all ranking list URLs', () => {
expect(__test__.resolveRankingUrl('new_releases')).toBe('https://www.amazon.com/gp/new-releases');
expect(__test__.resolveRankingUrl('movers_shakers')).toBe('https://www.amazon.com/gp/movers-and-shakers');
expect(__test__.resolveRankingUrl('new_releases', '/gp/new-releases/kitchen')).toBe('https://www.amazon.com/gp/new-releases/kitchen');
expect(__test__.resolveRankingUrl(
'bestsellers',
'https://www.amazon.com/Best-Sellers/zgbs/ref=zg_bsnr_tab_bs',
)).toBe('https://www.amazon.com/Best-Sellers/zgbs');
expect(() => __test__.resolveRankingUrl('movers_shakers', 'https://example.com/gp/movers-and-shakers')).toThrow('Invalid Amazon URL');
});
it('extracts category node id from URL best effort', () => {
expect(__test__.extractCategoryNodeId('https://www.amazon.com/Best-Sellers-Home-Kitchen/zgbs/home-garden/3744371')).toBe('3744371');
expect(__test__.extractCategoryNodeId('https://www.amazon.com/s?k=desk+organizer&rh=n%3A1064954')).toBe('1064954');
});
});
-438
View File
@@ -1,438 +0,0 @@
import { ArgumentError, CommandExecutionError } from '@jackwener/opencli/errors';
import type { IPage } from '@jackwener/opencli/types';
export const SITE = 'amazon';
export const DOMAIN = 'amazon.com';
export const HOME_URL = 'https://www.amazon.com/';
export const BESTSELLERS_URL = 'https://www.amazon.com/Best-Sellers/zgbs';
export const NEW_RELEASES_URL = 'https://www.amazon.com/gp/new-releases';
export const MOVERS_SHAKERS_URL = 'https://www.amazon.com/gp/movers-and-shakers';
export const SEARCH_URL_PREFIX = 'https://www.amazon.com/s?k=';
export const PRODUCT_URL_PREFIX = 'https://www.amazon.com/dp/';
export const DISCUSSION_URL_PREFIX = 'https://www.amazon.com/product-reviews/';
export const STRATEGY = 'cookie';
export const PRIMARY_PRICE_SELECTORS = [
'#corePrice_feature_div .a-offscreen',
'#corePriceDisplay_desktop_feature_div .a-offscreen',
'#corePrice_desktop .a-offscreen',
'#apex_desktop .a-offscreen',
'#newAccordionRow_0 .a-offscreen',
'#price_inside_buybox',
'#priceblock_ourprice',
'#priceblock_dealprice',
'#tp_price_block_total_price_ww',
];
const ROBOT_TEXT_PATTERNS = [
'Sorry, we just need to make sure you\'re not a robot',
'Enter the characters you see below',
'Type the characters you see in this image',
'To discuss automated access to Amazon data please contact',
];
export type AmazonRankingListType = 'bestsellers' | 'new_releases' | 'movers_shakers';
interface AmazonRankingSpec {
commandName: string;
rootUrl: string;
pathPattern: RegExp;
invalidInputMessage: string;
invalidInputHint: string;
}
const AMAZON_RANKING_SPECS: Record<AmazonRankingListType, AmazonRankingSpec> = {
bestsellers: {
commandName: 'bestsellers',
rootUrl: BESTSELLERS_URL,
pathPattern: /(?:^|\/)zgbs(?:\/|$)/i,
invalidInputMessage: 'amazon bestsellers expects a best sellers URL or /zgbs path',
invalidInputHint: 'Example: opencli amazon bestsellers https://www.amazon.com/Best-Sellers/zgbs',
},
new_releases: {
commandName: 'new-releases',
rootUrl: NEW_RELEASES_URL,
pathPattern: /\/gp\/new-releases(?:\/|$)/i,
invalidInputMessage: 'amazon new-releases expects a new releases URL or /gp/new-releases path',
invalidInputHint: 'Example: opencli amazon new-releases https://www.amazon.com/gp/new-releases',
},
movers_shakers: {
commandName: 'movers-shakers',
rootUrl: MOVERS_SHAKERS_URL,
pathPattern: /\/gp\/movers-and-shakers(?:\/|$)/i,
invalidInputMessage: 'amazon movers-shakers expects a movers-and-shakers URL or /gp/movers-and-shakers path',
invalidInputHint: 'Example: opencli amazon movers-shakers https://www.amazon.com/gp/movers-and-shakers',
},
};
export interface ProvenanceFields {
source_url: string;
fetched_at: string;
strategy: string;
}
export interface PageState {
href: string;
title: string;
body_text: string;
}
export interface PriceValue {
price_text: string | null;
price_value: number | null;
currency: string | null;
}
export function cleanText(value: unknown): string {
return typeof value === 'string'
? value.replace(/\u00a0/g, ' ').replace(/\s+/g, ' ').trim()
: '';
}
export function cleanMultilineText(value: unknown): string {
return typeof value === 'string'
? value
.replace(/\u00a0/g, ' ')
.split('\n')
.map((line) => line.replace(/\s+/g, ' ').trim())
.filter(Boolean)
.join('\n')
: '';
}
export function uniqueNonEmpty(values: Array<string | null | undefined>): string[] {
return [...new Set(values.map((value) => cleanText(value)).filter(Boolean))];
}
export function buildProvenance(sourceUrl: string): ProvenanceFields {
return {
source_url: sourceUrl,
fetched_at: new Date().toISOString(),
strategy: STRATEGY,
};
}
export function buildSearchUrl(query: string): string {
const normalized = cleanText(query);
if (!normalized) {
throw new ArgumentError('amazon search query cannot be empty');
}
return `${SEARCH_URL_PREFIX}${encodeURIComponent(normalized)}`;
}
export function extractAsin(input: string): string | null {
const normalized = cleanText(input);
if (!normalized) return null;
if (/^[A-Z0-9]{10}$/i.test(normalized)) {
return normalized.toUpperCase();
}
const match = normalized.match(/\/(?:dp|gp\/product|product-reviews)\/([A-Z0-9]{10})/i);
return match ? match[1].toUpperCase() : null;
}
export function buildProductUrl(input: string): string {
const asin = extractAsin(input);
if (!asin) {
throw new ArgumentError(
'amazon product expects an ASIN or product URL',
'Example: opencli amazon product B0FJS72893',
);
}
return `${PRODUCT_URL_PREFIX}${asin}`;
}
export function buildDiscussionUrl(input: string): string {
const asin = extractAsin(input);
if (!asin) {
throw new ArgumentError(
'amazon discussion expects an ASIN or product URL',
'Example: opencli amazon discussion B0FJS72893',
);
}
return `${DISCUSSION_URL_PREFIX}${asin}`;
}
function getRankingSpec(listType: AmazonRankingListType): AmazonRankingSpec {
return AMAZON_RANKING_SPECS[listType];
}
export function isSupportedRankingPath(listType: AmazonRankingListType, inputUrl: string): boolean {
try {
const url = new URL(inputUrl);
return getRankingSpec(listType).pathPattern.test(url.pathname);
} catch {
return false;
}
}
export function resolveRankingUrl(listType: AmazonRankingListType, input?: string): string {
const spec = getRankingSpec(listType);
const normalized = cleanText(input);
if (!normalized || normalized === 'root') return spec.rootUrl;
let candidateUrl: string;
if (normalized.startsWith('/')) {
candidateUrl = new URL(normalized, HOME_URL).toString();
} else if (/^https?:\/\//i.test(normalized)) {
candidateUrl = canonicalizeAmazonUrl(normalized);
} else if (normalized.includes('amazon.') && normalized.includes('/')) {
candidateUrl = canonicalizeAmazonUrl(`https://${normalized.replace(/^\/+/, '')}`);
} else {
throw new ArgumentError(spec.invalidInputMessage, spec.invalidInputHint);
}
if (!isSupportedRankingPath(listType, candidateUrl)) {
throw new ArgumentError(spec.invalidInputMessage, spec.invalidInputHint);
}
return normalizeRankingInputUrl(candidateUrl);
}
function normalizeRankingInputUrl(inputUrl: string): string {
try {
const url = new URL(inputUrl);
const normalizedPathSegments = url.pathname
.split('/')
.filter(Boolean)
.filter((segment) => !/^ref=/i.test(segment));
url.pathname = `/${normalizedPathSegments.join('/')}`;
url.hash = '';
// Ranking pages are frequently shared with tracking refs that can land on unstable variants.
// Dropping ref keeps the canonical ranking path while preserving useful params (for example pg=2).
url.searchParams.delete('ref');
return url.toString();
} catch {
return inputUrl;
}
}
export function isRankingPaginationUrl(listType: AmazonRankingListType, inputUrl: string): boolean {
const absolute = toAbsoluteAmazonUrl(inputUrl);
if (!absolute || !isSupportedRankingPath(listType, absolute)) return false;
try {
const url = new URL(absolute);
const ref = cleanText(url.searchParams.get('ref')).toLowerCase();
// pg= query param is the most reliable pagination indicator across all ranking lists
return url.searchParams.has('pg')
|| /(?:^|_)pg(?:_|$)/.test(ref)
// Amazon ranking pagination refs: zg_bs_pg_ (bestsellers), zg_bsnr_pg_ (new releases), zg_bsms_pg_ (movers & shakers)
|| /zg_bs(?:nr|ms)?_pg_/.test(ref);
} catch {
return false;
}
}
export function extractCategoryNodeId(inputUrl: string | null | undefined): string | null {
const absolute = toAbsoluteAmazonUrl(inputUrl);
if (!absolute) return null;
try {
const url = new URL(absolute);
for (const key of ['node', 'nodeid', 'nodeId', 'browseNode']) {
const value = cleanText(url.searchParams.get(key));
if (/^\d{4,}$/.test(value)) return value;
}
const rhValue = cleanText(url.searchParams.get('rh'));
const rhMatch = decodeURIComponent(rhValue).match(/(?:^|,)\s*n:(\d{4,})(?:,|$)/i);
if (rhMatch) return rhMatch[1];
const pathMatches = [...url.pathname.matchAll(/\/(\d{4,})(?=\/|$)/g)];
if (pathMatches.length > 0) {
return pathMatches[pathMatches.length - 1][1];
}
} catch {
return null;
}
return null;
}
export function resolveBestsellersUrl(input?: string): string {
return resolveRankingUrl('bestsellers', input);
}
export function canonicalizeAmazonUrl(input: string): string {
try {
const url = new URL(input);
if (!url.hostname.endsWith(DOMAIN)) {
throw new Error('not-amazon');
}
return url.toString();
} catch {
throw new ArgumentError('Invalid Amazon URL');
}
}
export function toAbsoluteAmazonUrl(value: string | null | undefined): string | null {
const normalized = cleanText(value);
if (!normalized) return null;
try {
return new URL(normalized, HOME_URL).toString();
} catch {
return null;
}
}
export function normalizeProductUrl(value: string | null | undefined): string | null {
const normalized = cleanText(value);
const asin = extractAsin(normalized);
if (asin) return buildProductUrl(asin);
return toAbsoluteAmazonUrl(normalized);
}
export function parsePriceText(text: string | null | undefined): PriceValue {
const normalized = cleanText(text);
const match = normalized.match(/([$€£])\s*(\d+(?:,\d{3})*(?:\.\d+)?)/);
if (!match) {
return {
price_text: normalized || null,
price_value: null,
currency: null,
};
}
const currencyMap: Record<string, string> = {
'$': 'USD',
'€': 'EUR',
'£': 'GBP',
};
return {
price_text: `${match[1]}${match[2]}`,
price_value: Number.parseFloat(match[2].replace(/,/g, '')),
currency: currencyMap[match[1]] ?? null,
};
}
export function parseRatingValue(text: string | null | undefined): number | null {
const normalized = cleanText(text);
const match = normalized.match(/(\d+(?:\.\d+)?)\s*out of 5/i);
return match ? Number.parseFloat(match[1]) : null;
}
export function parseReviewCount(text: string | null | undefined): number | null {
const normalized = cleanText(text);
const compactMatch = normalized.match(/(\d+(?:\.\d+)?)\s*([kKmM])/);
if (compactMatch) {
const value = Number.parseFloat(compactMatch[1]);
const multiplier = /m/i.test(compactMatch[2]) ? 1_000_000 : 1_000;
return Number.isFinite(value) ? Math.round(value * multiplier) : null;
}
const match = normalized.match(/([\d,]+)/);
return match ? Number.parseInt(match[1].replace(/,/g, ''), 10) : null;
}
export function extractReviewCountFromCardText(text: string | null | undefined): string | null {
const normalized = cleanMultilineText(text);
const match = normalized.match(/out of 5 stars(?:, rating details)?\s*([\d,]+)/i);
if (match) return match[1];
const numericLine = normalized
.split('\n')
.map((line) => cleanText(line))
.find((line) => /^[\d,]+$/.test(line));
return numericLine ?? null;
}
export function isAmazonEntity(text: string | null | undefined): boolean {
const normalized = cleanText(text).toLowerCase();
return normalized.includes('amazon');
}
export function firstMeaningfulLine(text: string | null | undefined): string {
return cleanMultilineText(text)
.split('\n')
.map((line) => cleanText(line))
.find(Boolean)
?? '';
}
export function trimRatingPrefix(text: string | null | undefined): string | null {
const normalized = cleanText(text);
if (!normalized) return null;
return normalized.replace(/^\d+(?:\.\d+)?\s*out of 5 stars\s*/i, '').trim() || normalized;
}
export function isRobotState(state: Partial<PageState>): boolean {
const title = cleanText(state.title);
const bodyText = cleanMultilineText(state.body_text);
return ROBOT_TEXT_PATTERNS.some((pattern) => title.includes(pattern) || bodyText.includes(pattern));
}
export function buildChallengeHint(action: string): string {
return [
`Open a clean Amazon ${action} page in the shared Chrome profile and clear any robot check first.`,
'If you are using CDP, set OPENCLI_CDP_TARGET=amazon.com and avoid parallel Amazon commands against the same browser target.',
].join(' ');
}
export async function readPageState(page: IPage): Promise<PageState> {
const result = await page.evaluate(`
(() => ({
href: window.location.href,
title: document.title || '',
body_text: document.body ? document.body.innerText || '' : '',
}))()
`) as Partial<PageState>;
return {
href: cleanText(result.href),
title: cleanText(result.title),
body_text: cleanMultilineText(result.body_text),
};
}
export async function gotoAndReadState(
page: IPage,
url: string,
settleMs: number = 2500,
action: string = 'page',
): Promise<PageState> {
try {
await page.goto(url, { settleMs });
await page.wait(1.5);
return await readPageState(page);
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
if (
message.includes('Inspected target navigated or closed')
|| message.includes('Cannot find context with specified id')
|| message.includes('Target closed')
) {
throw new CommandExecutionError(
`amazon ${action} navigation lost the current browser target`,
`${buildChallengeHint(action)} If CDP is attached to a stale tab, open a fresh Amazon tab and retry.`,
);
}
throw error;
}
}
export function assertUsableState(state: PageState, action: string): void {
if (!isRobotState(state)) return;
throw new CommandExecutionError(
`amazon ${action} hit a robot check`,
buildChallengeHint(action),
);
}
export const __test__ = {
buildSearchUrl,
extractAsin,
buildProductUrl,
buildDiscussionUrl,
resolveBestsellersUrl,
resolveRankingUrl,
isSupportedRankingPath,
isRankingPaginationUrl,
extractCategoryNodeId,
parsePriceText,
parseRatingValue,
parseReviewCount,
extractReviewCountFromCardText,
isAmazonEntity,
trimRatingPrefix,
isRobotState,
PRIMARY_PRICE_SELECTORS,
};
+28
View File
@@ -0,0 +1,28 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
import * as fs from 'node:fs';
export const dumpCommand = cli({
site: 'antigravity',
name: 'dump',
description: 'Dump the DOM to help AI understand the UI',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['htmlFile', 'snapFile'],
func: async (page) => {
// Extract HTML
const html = await page.evaluate('document.body.innerHTML');
fs.writeFileSync('/tmp/antigravity-dom.html', html);
// Extract Snapshot
let snapFile = '';
try {
const snap = await page.snapshot({ raw: true });
snapFile = '/tmp/antigravity-snapshot.json';
fs.writeFileSync(snapFile, JSON.stringify(snap, null, 2));
}
catch (e) {
snapFile = 'Failed';
}
return [{ htmlFile: '/tmp/antigravity-dom.html', snapFile }];
},
});
-30
View File
@@ -1,30 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
import * as fs from 'node:fs';
export const dumpCommand = cli({
site: 'antigravity',
name: 'dump',
description: 'Dump the DOM to help AI understand the UI',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['htmlFile', 'snapFile'],
func: async (page) => {
// Extract HTML
const html = await page.evaluate('document.body.innerHTML');
fs.writeFileSync('/tmp/antigravity-dom.html', html);
// Extract Snapshot
let snapFile = '';
try {
const snap = await page.snapshot({ raw: true });
snapFile = '/tmp/antigravity-snapshot.json';
fs.writeFileSync(snapFile, JSON.stringify(snap, null, 2));
} catch (e) {
snapFile = 'Failed';
}
return [{ htmlFile: '/tmp/antigravity-dom.html', snapFile }];
},
});
@@ -1,16 +1,15 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const extractCodeCommand = cli({
site: 'antigravity',
name: 'extract-code',
description: 'Extract multi-line code blocks from the current Antigravity conversation',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['code'],
func: async (page) => {
const blocks = await page.evaluate(`
site: 'antigravity',
name: 'extract-code',
description: 'Extract multi-line code blocks from the current Antigravity conversation',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['code'],
func: async (page) => {
const blocks = await page.evaluate(`
async () => {
// Find standard pre/code blocks
let elements = Array.from(document.querySelectorAll('pre code'));
@@ -28,7 +27,6 @@ export const extractCodeCommand = cli({
return elements.map(el => el.innerText).filter(text => text.trim().length > 0);
}
`);
return blocks.map((code: string) => ({ code }));
},
return blocks.map((code) => ({ code }));
},
});
@@ -1,20 +1,18 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const modelCommand = cli({
site: 'antigravity',
name: 'model',
description: 'Switch the active LLM model in Antigravity',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'name', help: 'Target model name (e.g. claude, gemini, o1)', required: true, positional: true }
],
columns: ['Status'],
func: async (page, kwargs) => {
const targetName = kwargs.name.toLowerCase();
await page.evaluate(`
site: 'antigravity',
name: 'model',
description: 'Switch the active LLM model in Antigravity',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'name', help: 'Target model name (e.g. claude, gemini, o1)', required: true, positional: true }
],
columns: ['Status'],
func: async (page, kwargs) => {
const targetName = kwargs.name.toLowerCase();
await page.evaluate(`
async () => {
const targetModelName = ${JSON.stringify(targetName)};
@@ -40,8 +38,7 @@ export const modelCommand = cli({
optionNode.click();
}
`);
await page.wait(0.5);
return [{ Status: `Model switched to: ${kwargs.name}` }];
},
await page.wait(0.5);
return [{ Status: `Model switched to: ${kwargs.name}` }];
},
});
+25
View File
@@ -0,0 +1,25 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const newCommand = cli({
site: 'antigravity',
name: 'new',
description: 'Start a new conversation / clear context in Antigravity',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['status'],
func: async (page) => {
await page.evaluate(`
async () => {
const btn = document.querySelector('[data-tooltip-id="new-conversation-tooltip"]');
if (!btn) throw new Error('Could not find New Conversation button');
// In case it's disabled, we must check, but we'll try to click it anyway
btn.click();
}
`);
// Give it a moment to reset the UI
await page.wait(0.5);
return [{ status: 'Successfully started a new conversation' }];
},
});
-28
View File
@@ -1,28 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const newCommand = cli({
site: 'antigravity',
name: 'new',
description: 'Start a new conversation / clear context in Antigravity',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['status'],
func: async (page) => {
await page.evaluate(`
async () => {
const btn = document.querySelector('[data-tooltip-id="new-conversation-tooltip"]');
if (!btn) throw new Error('Could not find New Conversation button');
// In case it's disabled, we must check, but we'll try to click it anyway
btn.click();
}
`);
// Give it a moment to reset the UI
await page.wait(0.5);
return [{ status: 'Successfully started a new conversation' }];
},
});
+34
View File
@@ -0,0 +1,34 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const readCommand = cli({
site: 'antigravity',
name: 'read',
description: 'Read the latest chat messages from Antigravity AI',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'last', help: 'Number of recent messages to read (not fully implemented due to generic structure, currently returns full history text or latest chunk)' }
],
columns: ['role', 'content'],
func: async (page, kwargs) => {
// We execute a script inside Antigravity's Chromium environment to extract the text
// of the entire conversation pane.
const rawText = await page.evaluate(`
async () => {
const container = document.getElementById('conversation');
if (!container) throw new Error('Could not find conversation container');
// Extract the full visible text of the conversation
// In Electron/Chromium, innerText preserves basic visual line breaks nicely
return container.innerText;
}
`);
// We can do simple heuristic parsing based on typical visual markers if needed.
// For now, we return the entire text blob, or just the last 2000 characters if it's too long.
const cleanText = String(rawText).trim();
return [{
role: 'history',
content: cleanText
}];
},
});
-36
View File
@@ -1,36 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const readCommand = cli({
site: 'antigravity',
name: 'read',
description: 'Read the latest chat messages from Antigravity AI',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'last', help: 'Number of recent messages to read (not fully implemented due to generic structure, currently returns full history text or latest chunk)' }
],
columns: ['role', 'content'],
func: async (page, kwargs) => {
// We execute a script inside Antigravity's Chromium environment to extract the text
// of the entire conversation pane.
const rawText = await page.evaluate(`
async () => {
const container = document.getElementById('conversation');
if (!container) throw new Error('Could not find conversation container');
// Extract the full visible text of the conversation
// In Electron/Chromium, innerText preserves basic visual line breaks nicely
return container.innerText;
}
`);
// We can do simple heuristic parsing based on typical visual markers if needed.
// For now, we return the entire text blob, or just the last 2000 characters if it's too long.
const cleanText = String(rawText).trim();
return [{
role: 'history',
content: cleanText
}];
},
});
+35
View File
@@ -0,0 +1,35 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const sendCommand = cli({
site: 'antigravity',
name: 'send',
description: 'Send a message to Antigravity AI via the internal Lexical editor',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'message', help: 'The message text to send', required: true, positional: true }
],
columns: ['Status', 'Message'],
func: async (page, kwargs) => {
const text = kwargs.message;
// We use evaluate to focus and insert text because Lexical editors maintain
// absolute control over their DOM and don't respond to raw node.textContent.
// document.execCommand simulates a native paste/typing action perfectly.
await page.evaluate(`
async () => {
const container = document.getElementById('antigravity.agentSidePanelInputBox');
if (!container) throw new Error('Could not find antigravity.agentSidePanelInputBox');
const editor = container.querySelector('[data-lexical-editor="true"]');
if (!editor) throw new Error('Could not find Antigravity input box');
editor.focus();
document.execCommand('insertText', false, ${JSON.stringify(text)});
}
`);
// Wait for the React/Lexical state to flush the new input
await page.wait(0.5);
// Press Enter to submit the message
await page.pressKey('Enter');
return [{ Status: 'Sent successfully', Message: text }];
},
});
-40
View File
@@ -1,40 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const sendCommand = cli({
site: 'antigravity',
name: 'send',
description: 'Send a message to Antigravity AI via the internal Lexical editor',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [
{ name: 'message', help: 'The message text to send', required: true, positional: true }
],
columns: ['Status', 'Message'],
func: async (page, kwargs) => {
const text = kwargs.message;
// We use evaluate to focus and insert text because Lexical editors maintain
// absolute control over their DOM and don't respond to raw node.textContent.
// document.execCommand simulates a native paste/typing action perfectly.
await page.evaluate(`
async () => {
const container = document.getElementById('antigravity.agentSidePanelInputBox');
if (!container) throw new Error('Could not find antigravity.agentSidePanelInputBox');
const editor = container.querySelector('[data-lexical-editor="true"]');
if (!editor) throw new Error('Could not find Antigravity input box');
editor.focus();
document.execCommand('insertText', false, ${JSON.stringify(text)});
}
`);
// Wait for the React/Lexical state to flush the new input
await page.wait(0.5);
// Press Enter to submit the message
await page.pressKey('Enter');
return [{ Status: 'Sent successfully', Message: text }];
},
});
+512
View File
@@ -0,0 +1,512 @@
/**
* antigravity serve — Anthropic-compatible `/v1/messages` proxy server.
*
* Starts an HTTP server that accepts Anthropic Messages API requests,
* forwards them to a running Antigravity app via CDP, polls for the response,
* and returns it in Anthropic format.
*
* Usage:
* opencli antigravity serve --port 8082
* ANTHROPIC_BASE_URL=http://localhost:8082 claude
*/
import { createServer } from 'node:http';
import { CDPBridge } from '@jackwener/opencli/browser/cdp';
import { resolveElectronEndpoint } from '@jackwener/opencli/launcher';
import { EXIT_CODES, getErrorMessage } from '@jackwener/opencli/errors';
// ─── Helpers ─────────────────────────────────────────────────────────
function generateMsgId() {
const chars = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789';
let id = 'msg_';
for (let i = 0; i < 24; i++)
id += chars[Math.floor(Math.random() * chars.length)];
return id;
}
function estimateTokens(text) {
// Rough approximation: ~4 chars per token for English, ~2 for CJK
return Math.max(1, Math.ceil(text.length / 3));
}
function extractTextContent(content) {
if (typeof content === 'string')
return content;
return content
.filter(b => b.type === 'text' && b.text)
.map(b => b.text)
.join('\n');
}
function readBody(req) {
return new Promise((resolve, reject) => {
const chunks = [];
req.on('data', (c) => chunks.push(c));
req.on('end', () => resolve(Buffer.concat(chunks).toString('utf-8')));
req.on('error', reject);
});
}
function jsonResponse(res, status, data) {
const body = JSON.stringify(data);
res.writeHead(status, {
'Content-Type': 'application/json',
'Access-Control-Allow-Origin': '*',
'Access-Control-Allow-Methods': 'GET, POST, OPTIONS',
'Access-Control-Allow-Headers': 'Content-Type, x-api-key, anthropic-version, Authorization',
});
res.end(body);
}
function sleep(ms) {
return new Promise(resolve => setTimeout(resolve, ms));
}
// ─── DOM helpers ─────────────────────────────────────────────────────
/**
* Click the 'New Conversation' button to reset context.
*/
async function startNewConversation(page) {
await page.evaluate(`
(() => {
const btn = document.querySelector('[data-tooltip-id="new-conversation-tooltip"]');
if (btn) btn.click();
})()
`);
await sleep(1000); // Give UI time to clear
}
/**
* Switch the active model in Antigravity UI.
*/
async function switchModel(page, anthropicModelId) {
// Map standard model IDs to Antigravity UI names based on actual UI
let targetName = 'claude sonnet 4.6'; // Default fallback
const id = anthropicModelId.toLowerCase();
if (id.includes('sonnet')) {
targetName = 'claude sonnet 4.6';
}
else if (id.includes('opus')) {
targetName = 'claude opus 4.6';
}
else if (id.includes('gemini') && id.includes('pro')) {
targetName = 'gemini 3.1 pro (high)';
}
else if (id.includes('gemini') && id.includes('flash')) {
targetName = 'gemini 3 flash';
}
else if (id.includes('gpt')) {
targetName = 'gpt-oss 120b';
}
try {
await page.evaluate(`
async () => {
const targetModelName = ${JSON.stringify(targetName)};
const trigger = document.querySelector('div[aria-haspopup="dialog"] > div[tabindex="0"]');
if (!trigger) return; // Silent fail if UI changed
// Open dropdown only if not already selected
if (trigger.innerText.toLowerCase().includes(targetModelName)) return;
trigger.click();
await new Promise(r => setTimeout(r, 200));
const spans = Array.from(document.querySelectorAll('[role="dialog"] span'));
const target = spans.find(s => s.innerText.toLowerCase().includes(targetModelName));
if (target) {
const optionNode = target.closest('.cursor-pointer') || target;
optionNode.click();
} else {
// Close if not found
trigger.click();
}
}
`);
await sleep(500); // Wait for switch
}
catch (err) {
console.error(`[serve] Warning: Could not switch to model ${targetName}:`, err);
}
}
/**
* Check if the Antigravity UI is currently generating a response
* by looking for Stop/Cancel buttons or loading indicators.
*/
async function isGenerating(page) {
const result = await page.evaluate(`
(() => {
// Look for a cancel/stop button in the UI
const cancelBtn = document.querySelector('button[aria-label*="cancel" i], button[aria-label*="stop" i], button[title*="cancel" i], button[title*="stop" i]');
return !!cancelBtn;
})()
`);
return Boolean(result);
}
/**
* Walk from the scroll container and find the deepest element that
* has multiple non-empty children (our message container).
*/
function findMessageContainer(root, depth = 0) {
if (!root || depth > 12)
return null;
const nonEmpty = Array.from(root.children).filter(c => c.innerText?.trim().length > 5);
if (nonEmpty.length >= 2)
return root;
if (nonEmpty.length === 1)
return findMessageContainer(nonEmpty[0], depth + 1);
return root;
}
// ─── Antigravity CDP Operations ──────────────────────────────────────
/**
* Get the full chat text for change-detection polling.
*/
async function getConversationText(page) {
const text = await page.evaluate(`
(() => {
const container = document.getElementById('conversation');
if (!container) return '';
// Read only the first child div (actual chat content),
// skipping UI chrome like file change panels, model selectors, etc.
const chatContent = container.children[0];
return chatContent ? chatContent.innerText : container.innerText;
})()
`);
return String(text ?? '');
}
/**
* Get the text of the last assistant reply by navigating to the message container
* and extracting the last non-empty message block.
*/
async function getLastAssistantReply(page, userText) {
const text = await page.evaluate(`
(() => {
const conv = document.getElementById('conversation')?.children[0];
const scroll = conv?.querySelector('.overflow-y-auto');
// Walk down until we find a container with multiple message siblings
function findMsgContainer(el, depth) {
if (!el || depth > 12) return null;
const nonEmpty = Array.from(el.children).filter(c => c.innerText && c.innerText.trim().length > 5);
if (nonEmpty.length >= 2) return el;
if (nonEmpty.length === 1) return findMsgContainer(nonEmpty[0], depth + 1);
return null;
}
const container = findMsgContainer(scroll || conv, 0);
if (!container) return '';
// Get all non-empty children (skip trailing empty UI divs)
const msgs = Array.from(container.children).filter(
c => c.innerText && c.innerText.trim().length > 5
);
if (msgs.length === 0) return '';
// The last element is the last assistant reply
const last = msgs[msgs.length - 1];
return last.innerText || '';
})()
`);
let reply = String(text ?? '').trim();
// Strip echoed user message from the top (Antigravity sometimes includes it)
if (userText && reply.startsWith(userText)) {
reply = reply.slice(userText.length).trim();
}
// Strip thinking block: "Thought for Xs\n..." at the start
reply = reply.replace(/^Thought for[^\n]*\n+/i, '').trim();
// Strip "Copy" button text at the end
reply = reply.replace(/\s*\bCopy\b\s*$/m, '').trim();
// De-duplicate trailing repeated content (e.g., "OK\n\nOK" → "OK")
const half = Math.floor(reply.length / 2);
const firstHalf = reply.slice(0, half).trim();
const secondHalf = reply.slice(half).trim();
if (firstHalf && firstHalf === secondHalf) {
reply = firstHalf;
}
return reply;
}
async function sendMessage(page, message, bridge) {
if (!bridge) {
// Fallback: use JS-based approach
await page.evaluate(`
(() => {
const container = document.getElementById('antigravity.agentSidePanelInputBox');
const editor = container?.querySelector('[data-lexical-editor="true"]');
if (!editor) throw new Error('Could not find input box');
editor.focus();
document.execCommand('insertText', false, ${JSON.stringify(message)});
})()
`);
await sleep(500);
await page.pressKey('Enter');
return;
}
// Get the bounding box of the Lexical editor for a physical mouse click
const rect = await page.evaluate(`
(() => {
const container = document.getElementById('antigravity.agentSidePanelInputBox');
if (!container) throw new Error('Could not find antigravity.agentSidePanelInputBox');
const editor = container.querySelector('[data-lexical-editor="true"]');
if (!editor) throw new Error('Could not find Antigravity input box');
const r = editor.getBoundingClientRect();
return JSON.stringify({ x: r.left + r.width / 2, y: r.top + r.height / 2 });
})()
`);
const { x, y } = JSON.parse(String(rect));
// Physical mouse click to give the element real browser focus
await bridge.send('Input.dispatchMouseEvent', { type: 'mousePressed', x, y, button: 'left', clickCount: 1 });
await sleep(50);
await bridge.send('Input.dispatchMouseEvent', { type: 'mouseReleased', x, y, button: 'left', clickCount: 1 });
await sleep(200);
// Inject text at the CDP level (no deprecated execCommand)
await bridge.send('Input.insertText', { text: message });
await sleep(300);
// Send Enter via native CDP key event
await bridge.send('Input.dispatchKeyEvent', { type: 'keyDown', key: 'Enter', code: 'Enter', windowsVirtualKeyCode: 13, nativeVirtualKeyCode: 13 });
await sleep(50);
await bridge.send('Input.dispatchKeyEvent', { type: 'keyUp', key: 'Enter', code: 'Enter', windowsVirtualKeyCode: 13, nativeVirtualKeyCode: 13 });
}
async function waitForReply(page, beforeText, opts = {}) {
const timeout = opts.timeout ?? 120_000; // 2 minutes max
const pollInterval = opts.pollInterval ?? 500; // 500ms polling
const deadline = Date.now() + timeout;
// Wait a bit to ensure the UI transitions to "generating" state after we hit Enter
await sleep(1000);
let hasStartedGenerating = false;
let lastText = beforeText;
let stableCount = 0;
const stableThreshold = 4; // 4 * 500ms = 2s of stability fallback
while (Date.now() < deadline) {
const generating = await isGenerating(page);
const currentText = await getConversationText(page);
const textChanged = currentText !== beforeText && currentText.length > 0;
if (generating) {
hasStartedGenerating = true;
stableCount = 0; // Reset stability while generating
}
else {
if (hasStartedGenerating) {
// It actively generated and now it stopped -> DONE
// Provide a small buffer to let React render the final message fully
await sleep(500);
return;
}
// Fallback: If it never showed "Generating/Cancel", but text changed and is stable
if (textChanged) {
if (currentText === lastText) {
stableCount++;
if (stableCount >= stableThreshold) {
return; // Text has been stable for 2 seconds -> DONE
}
}
else {
stableCount = 0;
lastText = currentText;
}
}
}
await sleep(pollInterval);
}
throw new Error('Timeout waiting for Antigravity reply');
}
// ─── Request Handlers ────────────────────────────────────────────────
async function handleMessages(body, page, bridge) {
// Extract the last user message
const userMessages = body.messages.filter(m => m.role === 'user');
if (userMessages.length === 0) {
throw new Error('No user message found in request');
}
const lastUserMsg = userMessages[userMessages.length - 1];
const userText = extractTextContent(lastUserMsg.content);
if (!userText.trim()) {
throw new Error('Empty user message');
}
// Optimization 1: New conversation if this is the first message in the session
if (body.messages.length === 1) {
console.error(`[serve] New session detected (1 message). Starting new conversation in UI.`);
await startNewConversation(page);
}
// Optimization 3: Switch model if requested
if (body.model) {
await switchModel(page, body.model);
}
// Get conversation state before sending
const beforeText = await getConversationText(page);
// Send the message
console.error(`[serve] Sending: "${userText.slice(0, 80)}${userText.length > 80 ? '...' : ''}"`);
await sendMessage(page, userText, bridge);
// Poll for reply (change detection)
console.error('[serve] Waiting for reply...');
await waitForReply(page, beforeText);
// Extract the actual reply text precisely from the DOM
const replyText = await getLastAssistantReply(page, userText);
console.error(`[serve] Got reply: "${replyText.slice(0, 80)}${replyText.length > 80 ? '...' : ''}"`);
return {
id: generateMsgId(),
type: 'message',
role: 'assistant',
content: [{ type: 'text', text: replyText }],
model: body.model ?? 'antigravity',
stop_reason: 'end_turn',
stop_sequence: null,
usage: {
input_tokens: estimateTokens(userText),
output_tokens: estimateTokens(replyText),
},
};
}
// ─── Server ──────────────────────────────────────────────────────────
export async function startServe(opts = {}) {
const port = opts.port ?? 8082;
// Lazy CDP connection — connect when first request comes in
let cdp = null;
let page = null;
let requestInFlight = false;
async function ensureConnected() {
if (page) {
try {
await page.evaluate('1+1');
return page;
}
catch {
console.error('[serve] CDP connection lost, reconnecting...');
cdp?.close().catch(() => { });
cdp = null;
page = null;
}
}
const endpoint = await resolveElectronEndpoint('antigravity');
// Note: Antigravity chat panel lives inside editor windows, not in Launchpad.
// If multiple editor windows are open, set OPENCLI_CDP_TARGET to the window title.
if (process.env.OPENCLI_CDP_TARGET) {
console.error(`[serve] Using OPENCLI_CDP_TARGET=${process.env.OPENCLI_CDP_TARGET}`);
}
// List available targets for debugging
try {
const res = await fetch(`${endpoint.replace(/\/$/, '')}/json`);
const targets = await res.json();
const pages = targets.filter(t => t.type === 'page');
console.error(`[serve] Available targets: ${pages.map(t => `"${t.title}"`).join(', ')}`);
}
catch { /* ignore */ }
console.error(`[serve] Connecting via CDP (target pattern: "${process.env.OPENCLI_CDP_TARGET}")...`);
cdp = new CDPBridge();
try {
page = await cdp.connect({ timeout: 15_000, cdpEndpoint: endpoint });
}
catch (err) {
cdp = null;
const errMsg = getErrorMessage(err);
const cause = err instanceof Error ? err.cause : undefined;
const isRefused = cause?.code === 'ECONNREFUSED' || errMsg.includes('ECONNREFUSED');
throw new Error(isRefused
? `Cannot connect to Antigravity at ${endpoint}.\n` +
' 1. Make sure Antigravity is running\n' +
' 2. Launch with: --remote-debugging-port=9234'
: `CDP connection failed: ${errMsg}`);
}
console.error('[serve] ✅ CDP connected.');
// Quick verification
const hasUI = await page.evaluate(`
(() => !!document.getElementById('conversation') || !!document.getElementById('antigravity.agentSidePanelInputBox'))()
`);
if (!hasUI) {
console.error('[serve] ⚠️ Warning: chat UI elements not found in this target. Try setting OPENCLI_CDP_TARGET to the correct window title.');
}
return page;
}
const server = createServer(async (req, res) => {
// CORS preflight
if (req.method === 'OPTIONS') {
res.writeHead(204, {
'Access-Control-Allow-Origin': '*',
'Access-Control-Allow-Methods': 'GET, POST, OPTIONS',
'Access-Control-Allow-Headers': 'Content-Type, x-api-key, anthropic-version, Authorization',
});
res.end();
return;
}
const url = req.url ?? '/';
const pathname = url.split('?')[0];
try {
// GET /v1/models — return available models
if (req.method === 'GET' && pathname === '/v1/models') {
jsonResponse(res, 200, {
data: [
{
id: 'antigravity',
object: 'model',
created: Math.floor(Date.now() / 1000),
owned_by: 'antigravity',
},
],
});
return;
}
// POST /v1/messages — main endpoint
if (req.method === 'POST' && pathname === '/v1/messages') {
if (requestInFlight) {
jsonResponse(res, 429, {
type: 'error',
error: {
type: 'rate_limit_error',
message: 'Another request is currently being processed. Antigravity can only handle one request at a time.',
},
});
return;
}
requestInFlight = true;
try {
const rawBody = await readBody(req);
const body = JSON.parse(rawBody);
if (body.stream) {
jsonResponse(res, 400, {
type: 'error',
error: {
type: 'invalid_request_error',
message: 'Streaming is not supported. Set "stream": false.',
},
});
return;
}
// Lazy connect on first request
const activePage = await ensureConnected();
const response = await handleMessages(body, activePage, cdp ?? undefined);
jsonResponse(res, 200, response);
}
finally {
requestInFlight = false;
}
return;
}
// Health check
if (req.method === 'GET' && (pathname === '/' || pathname === '/health')) {
jsonResponse(res, 200, { ok: true, cdpConnected: page !== null });
return;
}
jsonResponse(res, 404, {
type: 'error',
error: { type: 'not_found_error', message: `Not found: ${pathname}` },
});
}
catch (err) {
console.error('[serve] Error:', err instanceof Error ? err.message : err);
jsonResponse(res, 500, {
type: 'error',
error: {
type: 'api_error',
message: err instanceof Error ? err.message : 'Internal server error',
},
});
}
});
server.listen(port, '127.0.0.1', () => {
console.error(`\n[serve] ✅ Antigravity API proxy running at http://127.0.0.1:${port}`);
console.error(`[serve] Compatible with Anthropic /v1/messages API`);
console.error(`[serve] CDP connection will be established on first request.`);
console.error(`\n[serve] Usage with Claude Code:`);
console.error(` ANTHROPIC_BASE_URL=http://localhost:${port} claude\n`);
});
// Graceful shutdown
const shutdown = () => {
console.error('\n[serve] Shutting down...');
cdp?.close().catch(() => { });
server.close();
process.exit(EXIT_CODES.SUCCESS);
};
process.on('SIGTERM', shutdown);
process.on('SIGINT', shutdown);
// Keep alive
await new Promise(() => { });
}
-600
View File
@@ -1,600 +0,0 @@
/**
* antigravity serve — Anthropic-compatible `/v1/messages` proxy server.
*
* Starts an HTTP server that accepts Anthropic Messages API requests,
* forwards them to a running Antigravity app via CDP, polls for the response,
* and returns it in Anthropic format.
*
* Usage:
* opencli antigravity serve --port 8082
* ANTHROPIC_BASE_URL=http://localhost:8082 claude
*/
import { createServer, type IncomingMessage, type ServerResponse } from 'node:http';
import { CDPBridge } from '@jackwener/opencli/browser/cdp';
import type { IPage } from '@jackwener/opencli/types';
import { resolveElectronEndpoint } from '@jackwener/opencli/launcher';
import { EXIT_CODES, getErrorMessage } from '@jackwener/opencli/errors';
// ─── Types ───────────────────────────────────────────────────────────
interface AnthropicRequest {
model?: string;
max_tokens?: number;
system?: string | Array<{ type: string; text: string }>;
messages: Array<{ role: string; content: string | Array<{ type: string; text?: string }> }>;
stream?: boolean;
}
interface AnthropicResponse {
id: string;
type: 'message';
role: 'assistant';
content: Array<{ type: 'text'; text: string }>;
model: string;
stop_reason: 'end_turn' | 'max_tokens' | 'stop_sequence';
stop_sequence: null;
usage: { input_tokens: number; output_tokens: number };
}
// ─── Helpers ─────────────────────────────────────────────────────────
function generateMsgId(): string {
const chars = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789';
let id = 'msg_';
for (let i = 0; i < 24; i++) id += chars[Math.floor(Math.random() * chars.length)];
return id;
}
function estimateTokens(text: string): number {
// Rough approximation: ~4 chars per token for English, ~2 for CJK
return Math.max(1, Math.ceil(text.length / 3));
}
function extractTextContent(content: string | Array<{ type: string; text?: string }>): string {
if (typeof content === 'string') return content;
return content
.filter(b => b.type === 'text' && b.text)
.map(b => b.text!)
.join('\n');
}
function readBody(req: IncomingMessage): Promise<string> {
return new Promise((resolve, reject) => {
const chunks: Buffer[] = [];
req.on('data', (c: Buffer) => chunks.push(c));
req.on('end', () => resolve(Buffer.concat(chunks).toString('utf-8')));
req.on('error', reject);
});
}
function jsonResponse(res: ServerResponse, status: number, data: unknown): void {
const body = JSON.stringify(data);
res.writeHead(status, {
'Content-Type': 'application/json',
'Access-Control-Allow-Origin': '*',
'Access-Control-Allow-Methods': 'GET, POST, OPTIONS',
'Access-Control-Allow-Headers': 'Content-Type, x-api-key, anthropic-version, Authorization',
});
res.end(body);
}
function sleep(ms: number): Promise<void> {
return new Promise(resolve => setTimeout(resolve, ms));
}
// ─── DOM helpers ─────────────────────────────────────────────────────
/**
* Click the 'New Conversation' button to reset context.
*/
async function startNewConversation(page: IPage): Promise<void> {
await page.evaluate(`
(() => {
const btn = document.querySelector('[data-tooltip-id="new-conversation-tooltip"]');
if (btn) btn.click();
})()
`);
await sleep(1000); // Give UI time to clear
}
/**
* Switch the active model in Antigravity UI.
*/
async function switchModel(page: IPage, anthropicModelId: string): Promise<void> {
// Map standard model IDs to Antigravity UI names based on actual UI
let targetName = 'claude sonnet 4.6'; // Default fallback
const id = anthropicModelId.toLowerCase();
if (id.includes('sonnet')) {
targetName = 'claude sonnet 4.6';
} else if (id.includes('opus')) {
targetName = 'claude opus 4.6';
} else if (id.includes('gemini') && id.includes('pro')) {
targetName = 'gemini 3.1 pro (high)';
} else if (id.includes('gemini') && id.includes('flash')) {
targetName = 'gemini 3 flash';
} else if (id.includes('gpt')) {
targetName = 'gpt-oss 120b';
}
try {
await page.evaluate(`
async () => {
const targetModelName = ${JSON.stringify(targetName)};
const trigger = document.querySelector('div[aria-haspopup="dialog"] > div[tabindex="0"]');
if (!trigger) return; // Silent fail if UI changed
// Open dropdown only if not already selected
if (trigger.innerText.toLowerCase().includes(targetModelName)) return;
trigger.click();
await new Promise(r => setTimeout(r, 200));
const spans = Array.from(document.querySelectorAll('[role="dialog"] span'));
const target = spans.find(s => s.innerText.toLowerCase().includes(targetModelName));
if (target) {
const optionNode = target.closest('.cursor-pointer') || target;
optionNode.click();
} else {
// Close if not found
trigger.click();
}
}
`);
await sleep(500); // Wait for switch
} catch (err) {
console.error(`[serve] Warning: Could not switch to model ${targetName}:`, err);
}
}
/**
* Check if the Antigravity UI is currently generating a response
* by looking for Stop/Cancel buttons or loading indicators.
*/
async function isGenerating(page: IPage): Promise<boolean> {
const result = await page.evaluate(`
(() => {
// Look for a cancel/stop button in the UI
const cancelBtn = document.querySelector('button[aria-label*="cancel" i], button[aria-label*="stop" i], button[title*="cancel" i], button[title*="stop" i]');
return !!cancelBtn;
})()
`);
return Boolean(result);
}
/**
* Walk from the scroll container and find the deepest element that
* has multiple non-empty children (our message container).
*/
function findMessageContainer(root: Element | null, depth = 0): Element | null {
if (!root || depth > 12) return null;
const nonEmpty = Array.from(root.children).filter(
c => (c as HTMLElement).innerText?.trim().length > 5
);
if (nonEmpty.length >= 2) return root;
if (nonEmpty.length === 1) return findMessageContainer(nonEmpty[0], depth + 1);
return root;
}
// ─── Antigravity CDP Operations ──────────────────────────────────────
/**
* Get the full chat text for change-detection polling.
*/
async function getConversationText(page: IPage): Promise<string> {
const text = await page.evaluate(`
(() => {
const container = document.getElementById('conversation');
if (!container) return '';
// Read only the first child div (actual chat content),
// skipping UI chrome like file change panels, model selectors, etc.
const chatContent = container.children[0];
return chatContent ? chatContent.innerText : container.innerText;
})()
`);
return String(text ?? '');
}
/**
* Get the text of the last assistant reply by navigating to the message container
* and extracting the last non-empty message block.
*/
async function getLastAssistantReply(page: IPage, userText?: string): Promise<string> {
const text = await page.evaluate(`
(() => {
const conv = document.getElementById('conversation')?.children[0];
const scroll = conv?.querySelector('.overflow-y-auto');
// Walk down until we find a container with multiple message siblings
function findMsgContainer(el, depth) {
if (!el || depth > 12) return null;
const nonEmpty = Array.from(el.children).filter(c => c.innerText && c.innerText.trim().length > 5);
if (nonEmpty.length >= 2) return el;
if (nonEmpty.length === 1) return findMsgContainer(nonEmpty[0], depth + 1);
return null;
}
const container = findMsgContainer(scroll || conv, 0);
if (!container) return '';
// Get all non-empty children (skip trailing empty UI divs)
const msgs = Array.from(container.children).filter(
c => c.innerText && c.innerText.trim().length > 5
);
if (msgs.length === 0) return '';
// The last element is the last assistant reply
const last = msgs[msgs.length - 1];
return last.innerText || '';
})()
`);
let reply = String(text ?? '').trim();
// Strip echoed user message from the top (Antigravity sometimes includes it)
if (userText && reply.startsWith(userText)) {
reply = reply.slice(userText.length).trim();
}
// Strip thinking block: "Thought for Xs\n..." at the start
reply = reply.replace(/^Thought for[^\n]*\n+/i, '').trim();
// Strip "Copy" button text at the end
reply = reply.replace(/\s*\bCopy\b\s*$/m, '').trim();
// De-duplicate trailing repeated content (e.g., "OK\n\nOK" → "OK")
const half = Math.floor(reply.length / 2);
const firstHalf = reply.slice(0, half).trim();
const secondHalf = reply.slice(half).trim();
if (firstHalf && firstHalf === secondHalf) {
reply = firstHalf;
}
return reply;
}
async function sendMessage(page: IPage, message: string, bridge?: CDPBridge): Promise<void> {
if (!bridge) {
// Fallback: use JS-based approach
await page.evaluate(`
(() => {
const container = document.getElementById('antigravity.agentSidePanelInputBox');
const editor = container?.querySelector('[data-lexical-editor="true"]');
if (!editor) throw new Error('Could not find input box');
editor.focus();
document.execCommand('insertText', false, ${JSON.stringify(message)});
})()
`);
await sleep(500);
await page.pressKey('Enter');
return;
}
// Get the bounding box of the Lexical editor for a physical mouse click
const rect = await page.evaluate(`
(() => {
const container = document.getElementById('antigravity.agentSidePanelInputBox');
if (!container) throw new Error('Could not find antigravity.agentSidePanelInputBox');
const editor = container.querySelector('[data-lexical-editor="true"]');
if (!editor) throw new Error('Could not find Antigravity input box');
const r = editor.getBoundingClientRect();
return JSON.stringify({ x: r.left + r.width / 2, y: r.top + r.height / 2 });
})()
`);
const { x, y } = JSON.parse(String(rect));
// Physical mouse click to give the element real browser focus
await bridge.send('Input.dispatchMouseEvent', { type: 'mousePressed', x, y, button: 'left', clickCount: 1 });
await sleep(50);
await bridge.send('Input.dispatchMouseEvent', { type: 'mouseReleased', x, y, button: 'left', clickCount: 1 });
await sleep(200);
// Inject text at the CDP level (no deprecated execCommand)
await bridge.send('Input.insertText', { text: message });
await sleep(300);
// Send Enter via native CDP key event
await bridge.send('Input.dispatchKeyEvent', { type: 'keyDown', key: 'Enter', code: 'Enter', windowsVirtualKeyCode: 13, nativeVirtualKeyCode: 13 });
await sleep(50);
await bridge.send('Input.dispatchKeyEvent', { type: 'keyUp', key: 'Enter', code: 'Enter', windowsVirtualKeyCode: 13, nativeVirtualKeyCode: 13 });
}
async function waitForReply(
page: IPage,
beforeText: string,
opts: { timeout?: number; pollInterval?: number } = {},
): Promise<void> {
const timeout = opts.timeout ?? 120_000; // 2 minutes max
const pollInterval = opts.pollInterval ?? 500; // 500ms polling
const deadline = Date.now() + timeout;
// Wait a bit to ensure the UI transitions to "generating" state after we hit Enter
await sleep(1000);
let hasStartedGenerating = false;
let lastText = beforeText;
let stableCount = 0;
const stableThreshold = 4; // 4 * 500ms = 2s of stability fallback
while (Date.now() < deadline) {
const generating = await isGenerating(page);
const currentText = await getConversationText(page);
const textChanged = currentText !== beforeText && currentText.length > 0;
if (generating) {
hasStartedGenerating = true;
stableCount = 0; // Reset stability while generating
} else {
if (hasStartedGenerating) {
// It actively generated and now it stopped -> DONE
// Provide a small buffer to let React render the final message fully
await sleep(500);
return;
}
// Fallback: If it never showed "Generating/Cancel", but text changed and is stable
if (textChanged) {
if (currentText === lastText) {
stableCount++;
if (stableCount >= stableThreshold) {
return; // Text has been stable for 2 seconds -> DONE
}
} else {
stableCount = 0;
lastText = currentText;
}
}
}
await sleep(pollInterval);
}
throw new Error('Timeout waiting for Antigravity reply');
}
// ─── Request Handlers ────────────────────────────────────────────────
async function handleMessages(
body: AnthropicRequest,
page: IPage,
bridge?: CDPBridge,
): Promise<AnthropicResponse> {
// Extract the last user message
const userMessages = body.messages.filter(m => m.role === 'user');
if (userMessages.length === 0) {
throw new Error('No user message found in request');
}
const lastUserMsg = userMessages[userMessages.length - 1];
const userText = extractTextContent(lastUserMsg.content);
if (!userText.trim()) {
throw new Error('Empty user message');
}
// Optimization 1: New conversation if this is the first message in the session
if (body.messages.length === 1) {
console.error(`[serve] New session detected (1 message). Starting new conversation in UI.`);
await startNewConversation(page);
}
// Optimization 3: Switch model if requested
if (body.model) {
await switchModel(page, body.model);
}
// Get conversation state before sending
const beforeText = await getConversationText(page);
// Send the message
console.error(`[serve] Sending: "${userText.slice(0, 80)}${userText.length > 80 ? '...' : ''}"`);
await sendMessage(page, userText, bridge);
// Poll for reply (change detection)
console.error('[serve] Waiting for reply...');
await waitForReply(page, beforeText);
// Extract the actual reply text precisely from the DOM
const replyText = await getLastAssistantReply(page, userText);
console.error(`[serve] Got reply: "${replyText.slice(0, 80)}${replyText.length > 80 ? '...' : ''}"`);
return {
id: generateMsgId(),
type: 'message',
role: 'assistant',
content: [{ type: 'text', text: replyText }],
model: body.model ?? 'antigravity',
stop_reason: 'end_turn',
stop_sequence: null,
usage: {
input_tokens: estimateTokens(userText),
output_tokens: estimateTokens(replyText),
},
};
}
// ─── Server ──────────────────────────────────────────────────────────
export async function startServe(opts: { port?: number } = {}): Promise<void> {
const port = opts.port ?? 8082;
// Lazy CDP connection — connect when first request comes in
let cdp: CDPBridge | null = null;
let page: IPage | null = null;
let requestInFlight = false;
async function ensureConnected(): Promise<IPage> {
if (page) {
try {
await page.evaluate('1+1');
return page;
} catch {
console.error('[serve] CDP connection lost, reconnecting...');
cdp?.close().catch(() => {});
cdp = null;
page = null;
}
}
const endpoint = await resolveElectronEndpoint('antigravity');
// Note: Antigravity chat panel lives inside editor windows, not in Launchpad.
// If multiple editor windows are open, set OPENCLI_CDP_TARGET to the window title.
if (process.env.OPENCLI_CDP_TARGET) {
console.error(`[serve] Using OPENCLI_CDP_TARGET=${process.env.OPENCLI_CDP_TARGET}`);
}
// List available targets for debugging
try {
const res = await fetch(`${endpoint.replace(/\/$/, '')}/json`);
const targets = await res.json() as Array<{ title?: string; type?: string }>;
const pages = targets.filter(t => t.type === 'page');
console.error(`[serve] Available targets: ${pages.map(t => `"${t.title}"`).join(', ')}`);
} catch { /* ignore */ }
console.error(`[serve] Connecting via CDP (target pattern: "${process.env.OPENCLI_CDP_TARGET}")...`);
cdp = new CDPBridge();
try {
page = await cdp.connect({ timeout: 15_000, cdpEndpoint: endpoint });
} catch (err: unknown) {
cdp = null;
const errMsg = getErrorMessage(err);
const cause = err instanceof Error ? (err.cause as Record<string, unknown> | undefined) : undefined;
const isRefused = cause?.code === 'ECONNREFUSED' || errMsg.includes('ECONNREFUSED');
throw new Error(
isRefused
? `Cannot connect to Antigravity at ${endpoint}.\n` +
' 1. Make sure Antigravity is running\n' +
' 2. Launch with: --remote-debugging-port=9234'
: `CDP connection failed: ${errMsg}`
);
}
console.error('[serve] ✅ CDP connected.');
// Quick verification
const hasUI = await page.evaluate(`
(() => !!document.getElementById('conversation') || !!document.getElementById('antigravity.agentSidePanelInputBox'))()
`);
if (!hasUI) {
console.error('[serve] ⚠️ Warning: chat UI elements not found in this target. Try setting OPENCLI_CDP_TARGET to the correct window title.');
}
return page;
}
const server = createServer(async (req, res) => {
// CORS preflight
if (req.method === 'OPTIONS') {
res.writeHead(204, {
'Access-Control-Allow-Origin': '*',
'Access-Control-Allow-Methods': 'GET, POST, OPTIONS',
'Access-Control-Allow-Headers': 'Content-Type, x-api-key, anthropic-version, Authorization',
});
res.end();
return;
}
const url = req.url ?? '/';
const pathname = url.split('?')[0];
try {
// GET /v1/models — return available models
if (req.method === 'GET' && pathname === '/v1/models') {
jsonResponse(res, 200, {
data: [
{
id: 'antigravity',
object: 'model',
created: Math.floor(Date.now() / 1000),
owned_by: 'antigravity',
},
],
});
return;
}
// POST /v1/messages — main endpoint
if (req.method === 'POST' && pathname === '/v1/messages') {
if (requestInFlight) {
jsonResponse(res, 429, {
type: 'error',
error: {
type: 'rate_limit_error',
message: 'Another request is currently being processed. Antigravity can only handle one request at a time.',
},
});
return;
}
requestInFlight = true;
try {
const rawBody = await readBody(req);
const body = JSON.parse(rawBody) as AnthropicRequest;
if (body.stream) {
jsonResponse(res, 400, {
type: 'error',
error: {
type: 'invalid_request_error',
message: 'Streaming is not supported. Set "stream": false.',
},
});
return;
}
// Lazy connect on first request
const activePage = await ensureConnected();
const response = await handleMessages(body, activePage, cdp ?? undefined);
jsonResponse(res, 200, response);
} finally {
requestInFlight = false;
}
return;
}
// Health check
if (req.method === 'GET' && (pathname === '/' || pathname === '/health')) {
jsonResponse(res, 200, { ok: true, cdpConnected: page !== null });
return;
}
jsonResponse(res, 404, {
type: 'error',
error: { type: 'not_found_error', message: `Not found: ${pathname}` },
});
} catch (err) {
console.error('[serve] Error:', err instanceof Error ? err.message : err);
jsonResponse(res, 500, {
type: 'error',
error: {
type: 'api_error',
message: err instanceof Error ? err.message : 'Internal server error',
},
});
}
});
server.listen(port, '127.0.0.1', () => {
console.error(`\n[serve] ✅ Antigravity API proxy running at http://127.0.0.1:${port}`);
console.error(`[serve] Compatible with Anthropic /v1/messages API`);
console.error(`[serve] CDP connection will be established on first request.`);
console.error(`\n[serve] Usage with Claude Code:`);
console.error(` ANTHROPIC_BASE_URL=http://localhost:${port} claude\n`);
});
// Graceful shutdown
const shutdown = () => {
console.error('\n[serve] Shutting down...');
cdp?.close().catch(() => {});
server.close();
process.exit(EXIT_CODES.SUCCESS);
};
process.on('SIGTERM', shutdown);
process.on('SIGINT', shutdown);
// Keep alive
await new Promise(() => {});
}
+18
View File
@@ -0,0 +1,18 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const statusCommand = cli({
site: 'antigravity',
name: 'status',
description: 'Check Antigravity CDP connection and get current page state',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['status', 'url', 'title'],
func: async (page) => {
return {
status: 'Connected',
url: await page.evaluate('window.location.href'),
title: await page.evaluate('document.title'),
};
},
});
-19
View File
@@ -1,19 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const statusCommand = cli({
site: 'antigravity',
name: 'status',
description: 'Check Antigravity CDP connection and get current page state',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
columns: ['status', 'url', 'title'],
func: async (page) => {
return {
status: 'Connected',
url: await page.evaluate('window.location.href'),
title: await page.evaluate('document.title'),
};
},
});
+41
View File
@@ -0,0 +1,41 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const watchCommand = cli({
site: 'antigravity',
name: 'watch',
description: 'Stream new chat messages from Antigravity in real-time',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
timeoutSeconds: 86400, // Run for up to 24 hours
columns: [], // We use direct stdout streaming
func: async (page) => {
console.log('Watching Antigravity chat... (Press Ctrl+C to stop)');
let lastLength = 0;
// Loop until process gets killed
while (true) {
const text = await page.evaluate(`
async () => {
const container = document.getElementById('conversation');
return container ? container.innerText : '';
}
`);
const currentLength = text.length;
if (currentLength > lastLength) {
// Delta mode
const newSegment = text.substring(lastLength);
if (newSegment.trim().length > 0) {
process.stdout.write(newSegment);
}
lastLength = currentLength;
}
else if (currentLength < lastLength) {
// The conversation was cleared or updated significantly
lastLength = currentLength;
console.log('\\n--- Conversation Cleared/Changed ---\\n');
process.stdout.write(text);
}
await new Promise(resolve => setTimeout(resolve, 500));
}
},
});
-45
View File
@@ -1,45 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
export const watchCommand = cli({
site: 'antigravity',
name: 'watch',
description: 'Stream new chat messages from Antigravity in real-time',
domain: 'localhost',
strategy: Strategy.UI,
browser: true,
args: [],
timeoutSeconds: 86400, // Run for up to 24 hours
columns: [], // We use direct stdout streaming
func: async (page) => {
console.log('Watching Antigravity chat... (Press Ctrl+C to stop)');
let lastLength = 0;
// Loop until process gets killed
while (true) {
const text = await page.evaluate(`
async () => {
const container = document.getElementById('conversation');
return container ? container.innerText : '';
}
`);
const currentLength = text.length;
if (currentLength > lastLength) {
// Delta mode
const newSegment = text.substring(lastLength);
if (newSegment.trim().length > 0) {
process.stdout.write(newSegment);
}
lastLength = currentLength;
} else if (currentLength < lastLength) {
// The conversation was cleared or updated significantly
lastLength = currentLength;
console.log('\\n--- Conversation Cleared/Changed ---\\n');
process.stdout.write(text);
}
await new Promise(resolve => setTimeout(resolve, 500));
}
},
});
+99
View File
@@ -0,0 +1,99 @@
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { getRegistry } from '@jackwener/opencli/registry';
import './search.js';
import './top.js';
describe('apple-podcasts search command', () => {
beforeEach(() => {
vi.restoreAllMocks();
});
it('uses the positional query argument for the iTunes search request', async () => {
const cmd = getRegistry().get('apple-podcasts/search');
expect(cmd?.func).toBeTypeOf('function');
const fetchMock = vi.fn().mockResolvedValue({
ok: true,
json: () => Promise.resolve({
results: [
{
collectionId: 42,
collectionName: 'Machine Learning Guide',
artistName: 'OpenCLI',
trackCount: 12,
primaryGenreName: 'Technology',
},
],
}),
});
vi.stubGlobal('fetch', fetchMock);
const result = await cmd.func(null, {
query: 'machine learning',
keyword: 'sports',
limit: 5,
});
expect(fetchMock).toHaveBeenCalledWith('https://itunes.apple.com/search?term=machine%20learning&media=podcast&limit=5');
expect(result).toEqual([
expect.objectContaining({
id: 42,
title: 'Machine Learning Guide',
author: 'OpenCLI',
episodes: 12,
genre: 'Technology',
url: '',
}),
]);
});
});
describe('apple-podcasts top command', () => {
beforeEach(() => {
vi.restoreAllMocks();
});
it('adds a timeout signal to chart fetches', async () => {
const cmd = getRegistry().get('apple-podcasts/top');
expect(cmd?.func).toBeTypeOf('function');
const fetchMock = vi.fn().mockResolvedValue({
ok: true,
json: () => Promise.resolve({
feed: {
results: [
{ id: '100', name: 'Top Show', artistName: 'Host A' },
],
},
}),
});
vi.stubGlobal('fetch', fetchMock);
await cmd.func(null, { country: 'US', limit: 1 });
const [, options] = fetchMock.mock.calls[0] ?? [];
expect(options).toBeDefined();
expect(options.signal).toBeDefined();
expect(options.signal).toHaveProperty('aborted', false);
});
it('uses the canonical Apple charts host and maps ranked results', async () => {
const cmd = getRegistry().get('apple-podcasts/top');
expect(cmd?.func).toBeTypeOf('function');
const fetchMock = vi.fn().mockResolvedValue({
ok: true,
json: () => Promise.resolve({
feed: {
results: [
{ id: '100', name: 'Top Show', artistName: 'Host A' },
{ id: '101', name: 'Second Show', artistName: 'Host B' },
],
},
}),
});
vi.stubGlobal('fetch', fetchMock);
const result = await cmd.func(null, { country: 'US', limit: 2 });
expect(fetchMock).toHaveBeenCalledWith('https://rss.marketingtools.apple.com/api/v2/us/podcasts/top/2/podcasts.json', expect.objectContaining({
signal: expect.any(Object),
}));
expect(result).toEqual([
{ rank: 1, title: 'Top Show', author: 'Host A', id: '100' },
{ rank: 2, title: 'Second Show', author: 'Host B', id: '101' },
]);
});
it('normalizes network failures into CliError output', async () => {
const cmd = getRegistry().get('apple-podcasts/top');
expect(cmd?.func).toBeTypeOf('function');
vi.stubGlobal('fetch', vi.fn().mockRejectedValue(new Error('socket hang up')));
await expect(cmd.func(null, { country: 'us', limit: 3 })).rejects.toThrow('Unable to reach Apple Podcasts charts for US');
});
});
-123
View File
@@ -1,123 +0,0 @@
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { getRegistry } from '@jackwener/opencli/registry';
import './search.js';
import './top.js';
describe('apple-podcasts search command', () => {
beforeEach(() => {
vi.restoreAllMocks();
});
it('uses the positional query argument for the iTunes search request', async () => {
const cmd = getRegistry().get('apple-podcasts/search');
expect(cmd?.func).toBeTypeOf('function');
const fetchMock = vi.fn().mockResolvedValue({
ok: true,
json: () => Promise.resolve({
results: [
{
collectionId: 42,
collectionName: 'Machine Learning Guide',
artistName: 'OpenCLI',
trackCount: 12,
primaryGenreName: 'Technology',
},
],
}),
});
vi.stubGlobal('fetch', fetchMock);
const result = await cmd!.func!(null as any, {
query: 'machine learning',
keyword: 'sports',
limit: 5,
});
expect(fetchMock).toHaveBeenCalledWith(
'https://itunes.apple.com/search?term=machine%20learning&media=podcast&limit=5',
);
expect(result).toEqual([
expect.objectContaining({
id: 42,
title: 'Machine Learning Guide',
author: 'OpenCLI',
episodes: 12,
genre: 'Technology',
url: '',
}),
]);
});
});
describe('apple-podcasts top command', () => {
beforeEach(() => {
vi.restoreAllMocks();
});
it('adds a timeout signal to chart fetches', async () => {
const cmd = getRegistry().get('apple-podcasts/top');
expect(cmd?.func).toBeTypeOf('function');
const fetchMock = vi.fn().mockResolvedValue({
ok: true,
json: () => Promise.resolve({
feed: {
results: [
{ id: '100', name: 'Top Show', artistName: 'Host A' },
],
},
}),
});
vi.stubGlobal('fetch', fetchMock);
await cmd!.func!(null as any, { country: 'US', limit: 1 });
const [, options] = fetchMock.mock.calls[0] ?? [];
expect(options).toBeDefined();
expect(options.signal).toBeDefined();
expect(options.signal).toHaveProperty('aborted', false);
});
it('uses the canonical Apple charts host and maps ranked results', async () => {
const cmd = getRegistry().get('apple-podcasts/top');
expect(cmd?.func).toBeTypeOf('function');
const fetchMock = vi.fn().mockResolvedValue({
ok: true,
json: () => Promise.resolve({
feed: {
results: [
{ id: '100', name: 'Top Show', artistName: 'Host A' },
{ id: '101', name: 'Second Show', artistName: 'Host B' },
],
},
}),
});
vi.stubGlobal('fetch', fetchMock);
const result = await cmd!.func!(null as any, { country: 'US', limit: 2 });
expect(fetchMock).toHaveBeenCalledWith(
'https://rss.marketingtools.apple.com/api/v2/us/podcasts/top/2/podcasts.json',
expect.objectContaining({
signal: expect.any(Object),
}),
);
expect(result).toEqual([
{ rank: 1, title: 'Top Show', author: 'Host A', id: '100' },
{ rank: 2, title: 'Second Show', author: 'Host B', id: '101' },
]);
});
it('normalizes network failures into CliError output', async () => {
const cmd = getRegistry().get('apple-podcasts/top');
expect(cmd?.func).toBeTypeOf('function');
vi.stubGlobal('fetch', vi.fn().mockRejectedValue(new Error('socket hang up')));
await expect(cmd!.func!(null as any, { country: 'us', limit: 3 })).rejects.toThrow(
'Unable to reach Apple Podcasts charts for US',
);
});
});
+28
View File
@@ -0,0 +1,28 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
import { itunesFetch, formatDuration, formatDate } from './utils.js';
cli({
site: 'apple-podcasts',
name: 'episodes',
description: 'List recent episodes of an Apple Podcast (use ID from search)',
strategy: Strategy.PUBLIC,
browser: false,
args: [
{ name: 'id', positional: true, required: true, help: 'Podcast ID (collectionId from search output)' },
{ name: 'limit', type: 'int', default: 15, help: 'Max episodes to show' },
],
columns: ['title', 'duration', 'date'],
func: async (_page, args) => {
const limit = Math.max(1, Math.min(Number(args.limit), 200));
// results[0] is the podcast itself; the rest are episodes
const data = await itunesFetch(`/lookup?id=${args.id}&entity=podcastEpisode&limit=${limit + 1}`);
const episodes = (data.results ?? []).filter((r) => r.kind === 'podcast-episode');
if (!episodes.length)
throw new CliError('NOT_FOUND', 'No episodes found', 'Check the podcast ID from: opencli apple-podcasts search <keyword>');
return episodes.slice(0, limit).map((ep) => ({
title: ep.trackName,
duration: formatDuration(ep.trackTimeMillis),
date: formatDate(ep.releaseDate),
}));
},
});
-28
View File
@@ -1,28 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
import { itunesFetch, formatDuration, formatDate } from './utils.js';
cli({
site: 'apple-podcasts',
name: 'episodes',
description: 'List recent episodes of an Apple Podcast (use ID from search)',
strategy: Strategy.PUBLIC,
browser: false,
args: [
{ name: 'id', positional: true, required: true, help: 'Podcast ID (collectionId from search output)' },
{ name: 'limit', type: 'int', default: 15, help: 'Max episodes to show' },
],
columns: ['title', 'duration', 'date'],
func: async (_page, args) => {
const limit = Math.max(1, Math.min(Number(args.limit), 200));
// results[0] is the podcast itself; the rest are episodes
const data = await itunesFetch(`/lookup?id=${args.id}&entity=podcastEpisode&limit=${limit + 1}`);
const episodes = (data.results ?? []).filter((r: any) => r.kind === 'podcast-episode');
if (!episodes.length) throw new CliError('NOT_FOUND', 'No episodes found', 'Check the podcast ID from: opencli apple-podcasts search <keyword>');
return episodes.slice(0, limit).map((ep: any) => ({
title: ep.trackName,
duration: formatDuration(ep.trackTimeMillis),
date: formatDate(ep.releaseDate),
}));
},
});
+30
View File
@@ -0,0 +1,30 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
import { itunesFetch } from './utils.js';
cli({
site: 'apple-podcasts',
name: 'search',
description: 'Search Apple Podcasts',
strategy: Strategy.PUBLIC,
browser: false,
args: [
{ name: 'query', positional: true, required: true, help: 'Search keyword' },
{ name: 'limit', type: 'int', default: 10, help: 'Max results' },
],
columns: ['id', 'title', 'author', 'episodes', 'genre', 'url'],
func: async (_page, args) => {
const term = encodeURIComponent(args.query);
const limit = Math.max(1, Math.min(Number(args.limit), 25));
const data = await itunesFetch(`/search?term=${term}&media=podcast&limit=${limit}`);
if (!data.results?.length)
throw new CliError('NOT_FOUND', 'No podcasts found', `Try a different keyword`);
return data.results.map((p) => ({
id: p.collectionId,
title: p.collectionName,
author: p.artistName,
episodes: p.trackCount ?? '-',
genre: p.primaryGenreName ?? '-',
url: p.collectionViewUrl || '',
}));
},
});
-30
View File
@@ -1,30 +0,0 @@
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CliError } from '@jackwener/opencli/errors';
import { itunesFetch } from './utils.js';
cli({
site: 'apple-podcasts',
name: 'search',
description: 'Search Apple Podcasts',
strategy: Strategy.PUBLIC,
browser: false,
args: [
{ name: 'query', positional: true, required: true, help: 'Search keyword' },
{ name: 'limit', type: 'int', default: 10, help: 'Max results' },
],
columns: ['id', 'title', 'author', 'episodes', 'genre', 'url'],
func: async (_page, args) => {
const term = encodeURIComponent(args.query);
const limit = Math.max(1, Math.min(Number(args.limit), 25));
const data = await itunesFetch(`/search?term=${term}&media=podcast&limit=${limit}`);
if (!data.results?.length) throw new CliError('NOT_FOUND', 'No podcasts found', `Try a different keyword`);
return data.results.map((p: any) => ({
id: p.collectionId,
title: p.collectionName,
author: p.artistName,
episodes: p.trackCount ?? '-',
genre: p.primaryGenreName ?? '-',
url: p.collectionViewUrl || '',
}));
},
});

Some files were not shown because too many files have changed in this diff Show More