发布

  • [OPIK-6137] [QA] feat: Test Suites + Prompt Playground smokes (SDK + UI variants) (#6868)

    frostbyte_neo 发布于 2026-05-26 19:05:35 +00:00

    • [OPIK-6137] [QA] feat: add POST /test-suites and POST /test-suites/run bridge routes

    Wraps client.create_test_suite() + suite.insert() for seeding, and
    opik.run_tests() with a deterministic task function for running. The judge
    model is configurable per request (defaults to the SDK's gpt-4o-mini path).

    Tautological assertions in the test fixture combined with task_output='PASS'
    make the LLM judge's verdict mechanically predictable across runs while still
    exercising the entire judging code path.

    • feat(sdk): add createTestSuite and runTestSuite on PythonSdkClient

    Snake_case field names preserved per the OPIK-6128 transport-faithful
    convention. judge_model is passed through to opik.run_tests() so callers can
    pick a cheap/stable model (e.g. anthropic/claude-haiku-4-5) instead of the
    SDK's gpt-4o-mini default.

    • feat(backend-client): add findTestSuiteByName, listTestSuitesWithPrefix, getTestSuiteItems

    Test suites share storage with datasets, discriminated by DatasetPublic.type
    ('evaluation_suite' vs 'dataset'). The new methods filter on type so callers
    get back only suites, never plain datasets that happen to share a name.

    • feat(fixtures): add testSuite fixture extending experiment.fixture

    Seeds 3 items + 2 tautological LLM-judged assertions with runs_per_item=1.
    Explicit teardown via backendClient.deleteDataset() because test suites
    share storage with datasets and don't cascade with project deletion.

    • feat(fe): add data-testids for Playground, AI Providers, and Dataset detail page

    These attributes give the E2E POMs stable selectors for elements that
    previously required brittle ancestor scoping or accessible-name fallbacks.

    Playground (in PlaygroundPage/, LLMPromptMessages/):

    • playground-run-button (with data-mode='experiment-trigger'|'run'|'re-run')
    • playground-loaded-source-pill (with data-source-type)
    • playground-add-variant-button
    • playground-results-table
    • playground-variant-card (with data-variant-index)
    • playground-message-row (with data-role='system'|'user'|'assistant')

    RunExperimentDialog:

    • run-experiment-dialog
    • run-experiment-dialog-source-dataset
    • run-experiment-dialog-source-suite

    AI Providers config + Add provider dialog:

    • ai-providers-tabpanel
    • ai-provider-row-cell (with data-provider)
    • add-provider-dialog
    • add-provider-dialog-option (with data-provider)

    Dataset/Test-suite detail header:

    • dataset-detail-version-label
    • dataset-detail-global-assertions-pill (with data-count)
    • feat(bridge): add POST /test-suites/insert-items (idempotent get-or-create)

    Wraps client.get_or_create_test_suite + suite.insert. Lets the CUJ test
    seed items into a suite created via the UI without colliding on the name.

    • feat(poms): add TestSuites, TestSuiteItems, Playground, Configuration POMs

    TestSuitesPage: list view + create-suite wizard (Advanced settings expand
    to set Pass criteria + Global assertions). 3-step dialog ends on a
    'Test suite created!' confirmation screen with 'Go to test suite' nav.

    TestSuiteItemsPage: per-suite items table + 'Use test suite' → 'Open in
    Playground' menu path (handles both the confirmation dialog and the
    direct-load case when playground state is empty).

    PlaygroundPage: variant configuration with multi-role messages
    (System/Assistant/User), {{column}} templating, model picker, the
    'Run experiment' dialog (Dataset/Test suite tabs), and a one-shot
    runSimplePromptAndAwaitResponse helper for provider-sanity tests.

    ConfigurationPage: AI Providers tab with idempotent
    ensureProviderConfigured(provider, apiKey) for UI self-provisioning of
    provider keys from env vars.

    • feat(tests): add Test Suites + Playground CUJ smokes (+ provider sanity)

    @t1-smoke @test-suites:

    • Test A: SDK-create + SDK-run (tautological LLM-judged assertions,
      task returns 'PASS' so judge verdict is mechanically predictable).
      Items + assertion pill render correctly on the suite items page.
    • Test B: UI-create suite via the 3-step wizard + SDK-verify name +
      description, seed items via insert-items bridge route (idempotent),
      open in Prompt Playground, configure System+User messages with
      {{question}} templating, run, assert outputs land.

    @t1-smoke @playground:

    • Single case: load dataset into playground, run, SDK-verify an
      experiment landed (Playground runs auto-create experiments
      server-side — no separate Save step).

    @provider-sanity @playground (NOT in @t1-smoke):

    • Per provider × model from data/playground-models.yaml. Provisions
      keys from env vars via the AI Providers config UI (idempotent),
      then runs a simple prompt and asserts a non-empty response.

    js-yaml added as a runtime dep for the YAML config loader.

    • fix(qa): address baz-reviewer findings on test-suites bridge + POM
    • Bridge: TestSuiteRunRequest and TestSuiteInsertItemsRequest now require
      project_name and pass it through to get_or_create_test_suite. Without
      it, same-named suites across projects could collide on resolution.
    • backendClient.findTestSuiteByName: require type === 'evaluation_suite'
      strictly (no longer treats a missing type field as a suite). Plain
      datasets that share a name with a suite now return null cleanly.
    • PlaygroundPage.waitForRunsComplete: actually use the expectedRows arg.
      Previously only waited for 'No runs yet' to disappear, which could
      return early during table re-renders. Now also gates on row count
      matching expectedRows.
    • Playground T1 test: poll for the auto-created experiment instead of
      reading once. The experiment record lands shortly after rows render.
    下载附件