feature/python-runner-refactor
2383 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
54a30571aa |
Dotnet - Add support for Foundry Adaptive evals (#6267)
* .NET: feat(evals): RubricScore type + EvalScoreResult.Dimensions Adds the core rubric-evaluator surface that mirrors the Python work in PR #6101 (commit e45b934cc). Provider-agnostic types only — no Foundry coupling. Subsequent commits will wire these into FoundryEvals. - RubricScore: per-dimension score record (Id, Score?, Applicable, Weight, Reason). - EvalScoreResult.Dimensions: optional init-only list of RubricScore. Null for non-rubric (built-in) evaluators. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(evals): GeneratedEvaluatorRef + assertion helpers Adds the provider-agnostic surface for referencing a pre-existing rubric evaluator and gating CI on per-item / per-dimension thresholds. Mirrors Python PR #6101 commits e5830dd7f (ref type) and 4bc60462d (asserts). - GeneratedEvaluatorRef: name + optional version/display-name, plus a Latest(name) factory for versionless refs (discouraged for CI; consumers should warn at run time). - AgentEvaluationResults.AssertScoreAtLeast: walks DetailedItems[].Scores, optionally filtered by evaluator name, recurses into SubResults. - AgentEvaluationResults.AssertDimensionScoreAtLeast: walks each score's Dimensions list, skips non-applicable dimensions by default, supports requireApplicable to flip that, recurses into SubResults. - AgentEvaluationResults.AssertNoFailedItems: walks DetailedItems for fail/error statuses, recurses into SubResults. All helpers throw InvalidOperationException (matches existing AssertAllPassed). Truncates offender lists to the first 5 with a '+N more' suffix to keep CI output readable, mirroring the Python helpers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(foundry-evals): accept GeneratedEvaluatorRef in evaluators= Adds FoundryEvaluatorSpec, a readonly-struct union with implicit conversions from both string and GeneratedEvaluatorRef so call sites can mix built-in evaluator names with rubric evaluator references: var evals = new FoundryEvals( projectClient, model, new GeneratedEvaluatorRef("policy-rubric", "3"), FoundryEvals.Relevance, FoundryEvals.Coherence); FoundryEvals constructors (3 overloads), EvaluateTracesAsync, and EvaluateFoundryTargetAsync now take FoundryEvaluatorSpec[]/params instead of string[]/params. Existing call sites using string literals or string[] keep working unchanged via implicit conversion. FoundryEvalConverter.BuildTestingCriteria emits the documented Foundry wire format for rubric refs: { "type": "azure_ai_evaluator", "name": <DisplayName ?? Name>, "evaluator_name": <Name>, "evaluator_version": <Version>, // omitted when null "initialization_parameters": { "deployment_name": <model> }, "data_mapping": { conversation arrays, optional tool_definitions } } WireTestingCriterion gains an optional EvaluatorVersion field. Rubric refs are preserved through FilterToolEvaluators (tool-aware but not tool-required) and ignored by FindMissingGroundTruthEvaluators. A versionless ref emits a Trace.TraceWarning at criterion-build time so CI authors notice the floating version (mirrors the Python warning). Adds 6 new Foundry unit tests (3 BuildTestingCriteria rubric paths, 1 FindMissingGroundTruthEvaluators, 1 FilterToolEvaluators preservation, 1 mixed-order). 369/369 Foundry tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(foundry-evals): parse rubric dimension_scores into RubricScore Adds FoundryEvals.ParseRubricScores, called per result inside ParseDetailedItem. Each EvalScoreResult now populates Dimensions when the evaluator's sample carries a rubric breakdown. Accepts three shapes for forward compatibility with provider SDK iterations: 1. sample.properties.dimension_scores (canonical Foundry runtime shape) 2. sample.properties.rubric_scores (preview/legacy key) 3. top-level sample.dimension_scores / sample.rubric_scores (defensive fallback) Entries missing 'id', 'weight', or 'applicable' are skipped without invalidating well-formed siblings. Non-applicable dimensions may omit 'score' (parsed as null). Adds 6 unit tests covering canonical and legacy keys, top-level fallback, no-match returns null, malformed-entry skipping, and the non-applicable null-score path. 375/375 Foundry tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(samples): Evaluation_FoundryRubric end-to-end sample Adds dotnet/samples/05-end-to-end/Evaluation/Evaluation_FoundryRubric mirroring the Python evaluate_with_rubric_sample.py: - Fetches a pre-existing Foundry agent via AgentAdministrationClient (GetAgentAsync for latest, GetAgentVersionAsync when FOUNDRY_AGENT_VERSION is pinned). - References a rubric evaluator by GeneratedEvaluatorRef(name, version); falls back to GeneratedEvaluatorRef.Latest(name) with the documented floating-version warning. - Mixes the rubric with FoundryEvals.Relevance and FoundryEvals.Coherence in a single FoundryEvals run (implicit string-and-ref conversion). - Prints per-dimension breakdowns from EvalScoreResult.Dimensions for each item. - Demonstrates a CI quality gate with AssertDimensionScoreAtLeast("general_quality", 3.0). Documents the FOUNDRY_PROJECT_ENDPOINT footgun (must be project-scoped URL .../api/projects/<project>, not the bare Azure OpenAI endpoint) and the Eval-Definition-vs-Rubric-Evaluator distinction in the README. Ships a .env.example with the FOUNDRY_* variables. Registers the project in agent-framework-dotnet.slnx and cross-links from the sibling Evaluation_Multimodal / Evaluation_ExpectedOutputs READMEs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(foundry-evals): harden FoundryEvals public surface for review Address PR #6267 review comments on the .NET FoundryEvals integration: - Add source-compat overloads accepting `string[] evaluators` for `FoundryEvals` ctor, `EvaluateTracesAsync`, and `EvaluateFoundryTargetAsync` so existing callers passing string arrays keep compiling unchanged. New overloads forward via a private `ToSpecs` helper that wraps each name through the implicit `string -> FoundryEvaluatorSpec` conversion. - Guard against `default(FoundryEvaluatorSpec)` entries (both `BuiltinName` and `GeneratedRef` null) that would NRE the downstream converter. Adds `FoundryEvaluatorSpec.IsValid` / `EnsureValid` plus an internal `EnsureAllSpecsValid` helper, wired into the main ctor and both static evaluation entry points. - Add 6 unit tests covering the new validation surface. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sample): set ExitCode=1 when rubric dimension gate trips PR #6267 review comment: the FoundryRubric sample swallowed the AssertDimensionScoreAtLeast failure, so a CI run that included it as a quality gate would still exit 0 even when the rubric regressed. Set `System.Environment.ExitCode = 1` in the catch so CI fails while still letting the rest of the sample's logging complete cleanly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(foundry-evals): search typed Sample directly for rubric scores PR #6267 review comment: `_extract_rubric_scores` only searched the `properties` dict when the sample exposed one. When the Azure AI Projects typed SDK returns a Sample object that puts `dimension_scores` / `rubric_scores` directly on the instance (no `properties` wrapper), we missed them and surfaced no per-dimension scores. Add an `else: containers.append(sample)` branch so non-dict typed samples are also inspected for the score keys. Covered by two new tests: one with `dimension_scores` directly on a typed Sample without a `properties` wrapper, and one with the legacy `rubric_scores` key in the same shape. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * test(evals): cover assert_score_at_least and assert_no_failed_items PR #6267 review comments: both assertion helpers shipped without unit tests. Add `TestAssertScoreAtLeast` (above threshold, below w/ offenders, evaluator filter, sub_results recursion) and `TestAssertNoFailedItems` (all passing, failed/errored statuses, sub_results recursion) with a shared `_score_results` fixture builder. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(samples): remove dead rubric-evaluator doc link from FoundryRubric sample The Azure AI Foundry rubric evaluator concept doc page has not yet been published, so the link in the sample README and Program.cs comment 404s. Drop the references until the upstream doc is live. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Address PR 6267 review nits Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Ben Thomas <25218250+alliscode@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
dc445592ed |
Python: [BREAKING] Port FileMemoryProvider and integrate FileMemoryProvider & FileAccess into the harness agent (#6547)
* Port FileMemoryProvider to python and integrate it and FileAccessProvider into the harness * Address PR comments * Address PR comments * Create FileSystemAgentFileStore root lazily on first write Construction no longer calls mkdir, so building a store (and therefore a default create_harness_agent, which wires default file-memory and file-access stores under the CWD) performs no filesystem writes and does not fail in read-only working directories. The root directory is created on the first write_file / create_directory call; all read/list/search operations already tolerate a missing root. Updates docstrings and adds a regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix typing * Fixing typing errors --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
92823e9e61 |
.NET: [BREAKING] Require approval for FileAccessProvider tools with auto-approval rules (#6521)
* Require approval for file-access * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Rename DisableToolApproval to DisableToolAutoApproval for clarity * Fix broken suggestion. * Address PR comments and fix build issue. * Update dotnet/src/Microsoft.Agents.AI.Harness/HarnessAgentOptions.cs Co-authored-by: Roger Barreto <19890735+rogerbarreto@users.noreply.github.com> --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Roger Barreto <19890735+rogerbarreto@users.noreply.github.com> |
||
|
|
7a491f8e76 |
Python: Add hosting channel ADRs and spec (#6578)
* Add Python hosting channel ADRs Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Add Python hosting implementation spec Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
015e3bcd3b |
.NET: Enabling sequential orchestration to pass entire conversation or only previous output. (#6554)
* Fix sequential workflow input forwarding Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Make sequential workflow context configurable Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Clarify sequential chain-only behavior Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Clarify sequential output messaging Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Ben Thomas <25218250+alliscode@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
6e95517659 |
Python: Split type checkers by target (pyright source, 5 checkers on tests/samples) (#6443)
* Python: Split type checkers by target (pyright source, 5 checkers on tests/samples) Rework the typing setup along the lines of the 'too many type checkers' approach: - Pyright (strict) is now the sole source-code type checker; mypy is removed from source and its [tool.mypy] block becomes a relaxed profile used only for tests/samples. - Tests are checked by all five checkers (pyright relaxed, mypy, pyrefly, ty, zuban); samples by pyright, pyrefly, and ty. All run in a relaxed/ basic profile so authors aren't forced into over-annotation. - Add pyrightconfig.tests.json and bump sample pyright configs to basic. - Unify test/sample typing onto the same parallel fan-out used by source pyright via run_command_items in task_runner.py. - Make version-conditional imports symmetric: keep or drop the '# type: ignore' on both branches so results match across interpreter versions (local vs CI). - Update SKILL.md, DEV_SETUP.md, and CODING_STANDARD.md for the five gating checkers and pyright on source+tests+samples. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix merge regressions from main (typing + runtime) Merging main into the type-checker split branch surfaced regressions that the new five-checker test suite and unit tests caught: Runtime fixes: - anthropic: restore the dropped `cache_read_input_token_count` mapping in _parse_usage_from_anthropic (lost during merge conflict resolution). - gemini: _get_function_calling_mode test helper returned str(enum) ('FunctionCallingConfigMode.AUTO') instead of the enum value ('AUTO'). - openai: _response_id_from_token test helper was an infinite self-recursion; return token['response_id']. - orchestrations: reset output_events per approval iteration so the terminal output assertion counts only the final run. - core: drop a stale duplicate harness test whose message ('non-negative') contradicted the source ('positive'). - purview: import PolicyLocation/PolicyScope/ProtectionScopeActivities/ ExecutionMode used by the processor tests. Type-checker fixes (tests, relaxed profile): - core: pyright/mypy/pyrefly/ty/zuban green-ups across the harness, MCP, observability and types tests. - anthropic/openai: route provider-namespaced UsageDetails keys through a dict cast (extra_items TypedDict unsupported by mypy/ty). - purview: typed model constructors and cache-mock casts. - ag-ui: annotate WorkflowContext[Any, Any] so yield_output accepts test payloads, guard Optional forwarded_props, and ty-ignore intentional bad args. Source pyright (sole source checker) flagged unnecessary ignores newly introduced by merged code in core _tools.py and declarative _declarative_base.py. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Isolate per-package mypy cache in test-typing fan-out The parallel test-typing fan-out runs many mypy processes concurrently, all defaulting to a single shared ./.mypy_cache. Concurrent writes corrupt the cache and mypy aborts with INTERNAL ERROR (intermittently, depending on worker timing) -- which is why CI's Test Typing job failed on a shifting set of packages while a single-package run was fine. Give each mypy invocation an isolated cache dir keyed by its target paths so incremental caching still works per package without races. Other checkers (zuban/pyrefly/ty/pyright) maintain their own caches and are unaffected. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Make lab pyright-only on source (drop source mypy) Lab was the last package still running mypy on its source code, requiring mypy-only `# type: ignore` comments that pyright (the sole source checker everywhere else) flags as unnecessary. Align lab with the rest of the monorepo: - Remove the lab source mypy poe tasks (mypy-gaia/lightning/tau2) and the now-dead strict [tool.mypy] config block. - Drop the 'Run lab mypy' CI step; lab source is type-checked by pyright only. Lab tests remain covered by the workspace test-typing fan-out (mypy, pyrefly, ty, zuban, pyright over tests using the relaxed root config). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix test-typing regressions from latest main merge A fresh merge from main brought in new test code never run under the five-checker test-typing suite. Green up across the affected packages: - core: narrow Optional span.attributes with 'and' guards in span filters and assert+cast the json.loads(...attributes[...]) reads (test_observability); match the existing as_agent ignore on the protocol-typed fixture (test_clients). - openai: align new streaming tests with the established chat_options dict pattern (ChatOptions TypedDict isn't assignable to dict), route Optional .annotations[0] access through a small _first_annotation helper (mirrors the file's assert-not-None convention), and annotate a mapped ResponseStream. - foundry_hosting: annotate error: dict[str, Any] = body.get(...) or {} (zuban needs the annotation). - foundry: narrow ignores for the live AIProjectClient credential arg (pyrefly) and connections.get_default (zuban) SDK type gaps. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * updated pyright version * pyright fix * Python: Fix source typing for pyright 1.1.410 Pyright 1.1.410 tightened several checks. Apply the same source fixes as upstream PR #6275: - anthropic: import AsyncAnthropicBedrock from anthropic.lib.bedrock and AsyncAnthropicVertex from anthropic.lib.vertex (no longer re-exported from the anthropic top-level package -> reportPrivateImportUsage). - core _types.py: cast the transform-hook result to UpdateT (reportAssignmentType). - core _workflows/_events.py: annotate the @contextmanager helper as Generator[None] instead of Iterator[None] (reportDeprecated). - redis: build the combined filter expression with an explicit loop instead of reduce(and_, ...), which pyright could no longer fully type (drops the now unused functools.reduce / operator.and_ imports). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Accept plain-text body in Azure Functions workflow/run endpoint The workflow_orchestrator already accepts plain strings as well as JSON objects via context.get_input(), but the start_workflow_orchestration HTTP handler only accepted JSON and returned 400 for any non-JSON body. This made the functions integration tests that POST text/plain to /api/workflow/run (e.g. test_09_workflow_shared_state) fail consistently with 400 != 202. Fall back to the raw request body (decoded as UTF-8) when the body is not JSON, rejecting only a truly empty body. The JSON path is unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
97bb1d588a | Migrate to using issue type bug instead of label bug. (#6595) | ||
|
|
1fc57c45ee |
.NET: Bump Azure.AI.Projects to 2.1.0-beta.3 (#6542)
* Bump Azure.AI.Projects to 2.1.0-beta.3 Updates Azure.AI.Projects from 2.1.0-beta.2 to 2.1.0-beta.3, together with the transitive Azure.Core (1.56.0 to 1.57.0) and System.ClientModel (1.12.0 to 1.13.0) pins that beta.3 requires (beta.3 forces System.ClientModel 1.13.0.0 via Azure.Core 1.57.0). Migrates the affected samples and integration test to the beta.3 surface: * MemorySearch sample: MemorySearchToolCallResponseItem renamed to MemorySearchToolCall, Results renamed to Memories, MemoryItem indirection removed. * AgentSkills sample: skill provisioning/download API redesigned to a version based model (CreateSkillVersionFromFiles, GetSkillContent which now downloads and unzips), removing manual ZIP handling. * Session files integration test: GetSessionFilesAsync now returns an async collection of SessionDirectoryEntry and renames the sessionId parameter to agentSessionId. * Stream session file listing and short-circuit in integration test Avoids materializing the entire session directory listing into a List. The test now streams GetSessionFilesAsync and breaks as soon as the expected entry is found, then asserts it was located. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
b3f8aaa9d7 |
Python: adjust coverage report handoff (#6576)
* Adjust coverage report handoff Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Simplify coverage report handoff check Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
c22fc8d653 |
Build(deps): Bump esbuild, @tailwindcss/vite, @vitejs/plugin-react and vite (#6503)
Removes [esbuild](https://github.com/evanw/esbuild). It's no longer used after updating ancestor dependencies [esbuild](https://github.com/evanw/esbuild), [@tailwindcss/vite](https://github.com/tailwindlabs/tailwindcss/tree/HEAD/packages/@tailwindcss-vite), [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) and [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite). These dependencies need to be updated together. Removes `esbuild` Updates `@tailwindcss/vite` from 4.1.12 to 4.3.1 - [Release notes](https://github.com/tailwindlabs/tailwindcss/releases) - [Changelog](https://github.com/tailwindlabs/tailwindcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/tailwindlabs/tailwindcss/commits/v4.3.1/packages/@tailwindcss-vite) Updates `@vitejs/plugin-react` from 5.0.1 to 5.2.0 - [Release notes](https://github.com/vitejs/vite-plugin-react/releases) - [Changelog](https://github.com/vitejs/vite-plugin-react/blob/plugin-react@5.2.0/packages/plugin-react/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite-plugin-react/commits/plugin-react@5.2.0/packages/plugin-react) Updates `vite` from 7.3.2 to 8.0.16 - [Release notes](https://github.com/vitejs/vite/releases) - [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite/commits/v8.0.16/packages/vite) --- updated-dependencies: - dependency-name: esbuild dependency-version: dependency-type: indirect - dependency-name: "@tailwindcss/vite" dependency-version: 4.3.1 dependency-type: direct:production - dependency-name: "@vitejs/plugin-react" dependency-version: 5.2.0 dependency-type: direct:development - dependency-name: vite dependency-version: 8.0.16 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
a3131b8130 |
Build(deps): Bump esbuild, @vitejs/plugin-react and vite (#6501)
Removes [esbuild](https://github.com/evanw/esbuild). It's no longer used after updating ancestor dependencies [esbuild](https://github.com/evanw/esbuild), [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) and [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite). These dependencies need to be updated together. Removes `esbuild` Updates `@vitejs/plugin-react` from 4.7.0 to 6.0.2 - [Release notes](https://github.com/vitejs/vite-plugin-react/releases) - [Changelog](https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite-plugin-react/commits/plugin-react@6.0.2/packages/plugin-react) Updates `vite` from 7.3.2 to 8.0.16 - [Release notes](https://github.com/vitejs/vite/releases) - [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite/commits/v8.0.16/packages/vite) --- updated-dependencies: - dependency-name: esbuild dependency-version: dependency-type: indirect - dependency-name: "@vitejs/plugin-react" dependency-version: 6.0.2 dependency-type: direct:development - dependency-name: vite dependency-version: 8.0.16 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
3d46595111 |
Python: Bump prek from 0.4.3 to 0.4.5 in /python (#6527)
* Bump prek from 0.4.3 to 0.4.5 in /python Bumps [prek](https://github.com/j178/prek) from 0.4.3 to 0.4.5. - [Release notes](https://github.com/j178/prek/releases) - [Changelog](https://github.com/j178/prek/blob/master/CHANGELOG.md) - [Commits](https://github.com/j178/prek/compare/v0.4.3...v0.4.5) --- updated-dependencies: - dependency-name: prek dependency-version: 0.4.5 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> * fix python workspace prek pin mismatch --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> |
||
|
|
205f7bcca8 |
Python: Bump pytest from 9.0.3 to 9.1.0 across /python workspace (#6524)
* Bump pytest from 9.0.3 to 9.1.0 in /python Bumps [pytest](https://github.com/pytest-dev/pytest) from 9.0.3 to 9.1.0. - [Release notes](https://github.com/pytest-dev/pytest/releases) - [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst) - [Commits](https://github.com/pytest-dev/pytest/compare/9.0.3...9.1.0) --- updated-dependencies: - dependency-name: pytest dependency-version: 9.1.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Fix Python workspace pytest pin mismatch --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> |
||
|
|
2048289fb0 |
Build(deps): Bump pydantic-monty from 0.0.17 to 0.0.18 in /python (#6392)
Bumps [pydantic-monty](https://github.com/pydantic/monty) from 0.0.17 to 0.0.18. - [Release notes](https://github.com/pydantic/monty/releases) - [Commits](https://github.com/pydantic/monty/compare/v0.0.17...v0.0.18) --- updated-dependencies: - dependency-name: pydantic-monty dependency-version: 0.0.18 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
699916d639 |
Python: Add WebSearchDisplayObserver to harness console (#6572)
* Adding an observer to the python harness for web search tools * Escape dynamic strings with rich.markup.escape() in WebSearchDisplayObserver Apply rich.markup.escape() to all user/tool-provided strings (queries, URLs, titles, patterns) before interpolation into Rich-markup-enabled output. This prevents characters like '['/']' from being interpreted as Rich markup tags. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
1ba5cd3f44 |
.NET: Scope argument-based standing approvals correctly in ToolApprovalAgent (#6486) (#6487)
Ensure an argument-scoped standing approval (the "always approve with exact arguments" path) records an empty argument set rather than null when the approved call has no arguments, so it matches only future no-argument calls. null remains reserved exclusively for tool-level approvals, keeping the two scopes distinct. This aligns the .NET behavior with the existing Python harness. Adds regression tests covering the no-argument standing-approval flow, the MatchesRule argument-scoping semantics, and empty-arguments rule serialization. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
1519e50f2f |
Harden archive extraction guard so path containment is statically recognized (#6564) (#6565)
The Hosted-AgentSkills sample and its mirrored unit-test helper gated ZIP extraction on `StartsWith(destinationRoot)` OR `Equals(destinationRoot)`. The second branch left an acceptance path not covered by the containment check, so static analysis could not prove the extraction sink stays within the destination. Make the single resolved-path StartsWith check the only gate to extraction in both files and add a nested-entry regression test. Closes #6564 |
||
|
|
b55992bb67 |
Bump Python package versions for 1.9.0 release (#6583)
Selective, CHANGELOG-driven version bumps for the 2026-06-18 release. Released tier: agent-framework-core and the root agent-framework go to 1.9.0 (minor). Core ships new public APIs (agent-loop middleware, tool-approval middleware and harness integration, shell-tool harness integration, AG-UI thread snapshot persistence, context-provider telemetry) plus two behavioral breaking changes on evolving surfaces: MCP sampling now denies server-initiated requests by default, and the FileAccess tools were aligned with the .NET implementation. These are treated as within-1.x changes because every package caps core at <2; a major bump would require rewriting those caps. The foundry and openai packages go to 1.8.2 (patch, bug fixes only). The root agent-framework-core[all] pin was moved to 1.9.0 in lockstep with core. Release-candidate tier: ag-ui to 1.0.0rc5 and declarative to 1.0.0rc2 for their respective changes. orchestrations is promoted to stable 1.0.0; PACKAGE_STATUS and the README install hint were updated accordingly. Prerelease tier (new Pacific date stamp 260618): anthropic (beta), azure-contentunderstanding (alpha) and foundry-hosting (alpha). No beta cohort bump was applied; only packages with changes this cycle were stamped. Dependency floors: following the established convention, the core floor was raised to >=1.9.0 on every non-core package bumped this cycle, preserving the existing <2 upper bound. Also resolves two pre-existing failures in the dependency-bounds validator that are unrelated to the version bumps. Hosted-environment detection now catches a bare ImportError so optional Foundry hosting probing cannot crash user-agent setup. The harness shell-tool integration, which lazily imports the separate agent-framework-tools package to avoid a circular runtime dependency, is now type-checked and tested in isolated environments via a core dev dependency-group, with the shell-tool tests guarded to skip when that package is absent.python-1.9.0 |
||
|
|
e8cec71ed8 |
Use issue type for triage workflow (#6577)
* Use issue type for triage workflow * Disable blank issue reports * Revert "Disable blank issue reports" This reverts commit 222c8444a7b3b5768e01b9d562195a27d1a29f1a. |
||
|
|
d7e63d7d0e |
Fix Foundry aiohttp dependency (#6567)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
f59d5c67d8 |
Python: Adopt azure-ai-contentunderstanding to_llm_input in CU context provider (#5796)
* Refactor DocumentEntry model and update result handling - Changed the type of `result` in DocumentEntry from dict to str to store LLM-ready text. - Introduced `search_payload` in DocumentEntry for optional alternate rendering. - Updated FileSearchConfig to include `include_fields` option for vector store uploads. - Modified tests to reflect changes in DocumentEntry and FileSearchConfig. - Adjusted integration tests to validate new result structure and rendering. - Removed legacy format_result tests as rendering is now handled by the SDK. * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Add test to ensure page markers are preserved in LLM input Co-authored-by: Copilot <copilot@github.com> * fix(cu-context-provider): scope LLMStats telemetry filter to rai_warnings block Address PR #5796 review comment: the previous defensive scrubber ran a global regex substitution over the full rendered string, so any markdown body bullet shaped like '- LLMStats: ...' would also be silently deleted. Add a _strip_rai_telemetry helper that confines the substitution to the front-matter rai_warnings: YAML sub-block, leaving the body verbatim. Cover the new behavior with three tests (scoped strip, body preservation, and no-op branches). * Sync uv.lock with azure-ai-contentunderstanding>=1.2.0b1 dependency bump * Python: Drop search_payload/include_fields, single to_llm_input rendering (CU context provider) Address PR #5796 review: remove the redundant search_payload field and _render_search_payload helper, drop the include_fields opt-in (already covered by output_sections), rename _resolve_pending_tokens -> _resolve_pending_analysis, and have _upload_to_vector_store read entry['result'] directly. * Python: Adopt SDK 1.2.0b2 LLMStats filtering, drop local workaround (CU context provider) azure-ai-contentunderstanding 1.2.0b2 filters LLMStats telemetry from rai_warnings and emits InputPageNumber page markers in to_llm_input, so the provider's local defense is redundant. - Bump dependency to azure-ai-contentunderstanding>=1.2.0b2 (re-lock uv.lock) - Remove _strip_rai_telemetry and its two regexes; _render_for_llm now returns to_llm_input(...) directly - Delete 4 workaround unit tests for the removed helper --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Copilot <copilot@github.com> Co-authored-by: changjian-wang <v-changjwang@microsoft.com> Co-authored-by: aluneth <wangchangjian1130@163.com> |
||
|
|
26a0a7e8be |
.NET: (Durable): bind MCP threadId to the current agent and guard cross-agent session dispatch (#6531)
* scope MCP threadId to the current agent * Fix Async suffix on test methods and add CHANGELOG entries - Rename three test methods to include Async suffix (IDE1006 fix) - Add CHANGELOG entries for DurableTask and Hosting.AzureFunctions Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
616315339e | .NET: fix fan-in checkpoint edge state (#6491) | ||
|
|
fcc5576b04 |
.NET: feat(dotnet): Add LocalCodeAct package for local Python execution (#6105)
* feat(dotnet): Add LocalCodeAct package scaffold Create Microsoft.Agents.AI.LocalCodeAct package with: - Project file with embedded Python resources - ExecutionMode enum (Subprocess only) - ProcessExecutionLimits record - FileMount record and FileMountMode enum - README.md documentation - Embedded Python runner and validator scripts This is the .NET equivalent of the Python agent-framework-local-codeact package. Next: Implement process bridge and tool integration. * feat(dotnet): Add embedded Python runner and validator Copy Python runner and validator scripts from the Python implementation as embedded resources for the .NET package. * feat(dotnet): Add CodeValidator wrapper Implement CodeValidator.cs that: - Extracts embedded Python validator script to temp file - Invokes Python validator with JSON request - Passes custom allow/block lists - Throws CodeValidationException on failures - Cleans up temp files Uses the embedded Resources/validator.py for AST validation. * feat(dotnet): Add LocalExecuteCodeFunction Implement LocalExecuteCodeFunction as AIFunction: - Accepts Python executable path (required) - Registers host tools for code to call - Validates code via CodeValidator if custom lists provided - Executes via ProcessBridge - Converts result dict to ChatMessage list - Builds dynamic description including available tools Matches Python LocalExecuteCodeTool functionality. * feat(dotnet): Add LocalCodeActProvider Implement AIContextProvider that: - Injects execute_code tool into context - Adds CodeAct instructions - Enforces single-provider-per-agent via StateKeys - Wraps LocalExecuteCodeFunction lifecycle Minimal provider implementation matching Python LocalCodeActProvider. * feat(dotnet): Add tests and sample for LocalCodeAct Add unit tests: - LocalExecuteCodeFunctionTests (4 tests) - ProcessExecutionLimitsTests (2 tests) - FileMountTests (2 tests) Add sample: - LocalCodeAct/Program.cs - Demonstrates provider and function usage - LocalCodeAct/README.md - Documentation and safety warnings Tests verify basic construction, metadata, and disposal. Sample shows provider creation, function setup, and configuration. Note: Build requires .NET 10 SDK per global.json. * feat(dotnet): Add LocalCodeAct sample project Add sample demonstrating: - LocalCodeActProvider creation and configuration - LocalExecuteCodeFunction direct usage - Execution modes and file mount configuration - Safety warnings and prerequisites Includes project file and README with security guidance. * feat(dotnet): Add file mount support and integration tests - Added FileMountHelper.cs for file mount normalization, snapshot, and capture - Updated LocalExecuteCodeFunction to support file mounts parameter - Added file snapshot before/after execution with capture logic - Updated LocalCodeActProvider to pass file mounts through - Created comprehensive IntegrationTests.cs with 10 test cases: - Simple code execution - Timeout handling - Syntax error handling - Blocked import validation - Blocked builtin validation - Custom allowed imports - File mount read/write with capture - Stdout capture - Provider tool injection All features from Python implementation now ported to .NET. * Rewrite .NET LocalCodeAct to address all PR review comments Complete rewrite that follows the Hyperlight package conventions (see Microsoft.Agents.AI.Hyperlight) and addresses all 24 review comments on PR #6105: Architectural fixes: * LocalCodeActProvider now uses options-class constructor pattern matching HyperlightCodeActProvider. * Override of ProvideAIContextAsync uses the correct (InvokingContext, CancellationToken) signature returning ValueTask<AIContext>. * ExecuteCodeFunction follows the AIFunction Name/Description/JsonSchema property pattern with InvokeCoreAsync override. * Provider exposes AddTools/GetTools/RemoveTools/ClearTools and AddFileMounts/GetFileMounts/RemoveFileMounts/ClearFileMounts CRUD methods, with snapshot-at-invocation semantics under a lock. Runtime/security fixes: * Subprocess IPC uses JsonObject/JsonNode end-to-end (no Dictionary<string, object?> casts that broke under JsonElement deserialization). * Validator runs in its own subprocess with a dedicated timeout (ProcessExecutionLimits.ValidationTimeoutSeconds), never reuses the runner script. * Validation enabled by default; can be opt-ed out via ValidationEnabled = false. * validator.py has a __main__ entrypoint that reads JSON from stdin and exits with structured errors. * validator.py is now compatible with Python 3.9+ (Match nodes added conditionally). * call_id parsed as long to match Python id(kwargs) range. Other: * README rewritten with valid C# syntax (options-class, FileMount constructor) and accurate descriptions of validator and file capture behavior. * Added integration tests that exercise the real subprocess and validator (skipped gracefully when python3 is not on PATH). * All 18 tests pass (15 unit + 3 integration) across net8/net9/net10. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Sync embedded validator.py with Python package allow-list enforcement The embedded Python validator script used by the .NET LocalCodeAct package now enforces the builtin allow-list, matching the latest behavior of agent_framework_local_codeact._validator. Names that are real Python builtins must appear in the allow-list, while unknown names (user-defined functions, registered tools) remain allowed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Add Hosted-LocalCodeAct foundry hosted-agent sample Mirrors the Python foundry_hosted_agent.py sample for the local-codeact package: registers compute and fetch_data as sandbox-only host tools on LocalCodeActProvider so the model only sees execute_code and reaches them via await call_tool(...). Includes the standard hosted-agent supporting files (agent.yaml, agent.manifest.yaml, Dockerfile, Dockerfile.contributor, .env.example, README.md) and installs python3 in the container images so the embedded runner and validator can execute. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(local-codeact-dotnet): sync validator os.* allow-list with Python Mirror the Python package change: the embedded validator.py invoked by the .NET ProcessBridge replaces the os.* deny-list with an allow-list of {environ, path}. Add allowed_os_attrs parameter to validate_code and _CodeValidator, and surface it via the stdin JSON request schema so the .NET host can opt in to a broader allow-list when needed. Default behavior tightens to match the documented contract: any os.* attribute outside {environ, path} (for example os.listdir, os.open, os.getcwd) is rejected. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(local-codeact-dotnet): address review + tighten validator - validator.py: enforce os.* allow-list on `from os import X` so names like `system`, `getcwd` cannot bypass the visit_Attribute restriction. - ProcessBridge.ConfigureEnvironment: document that null Environment inherits the parent env (matching real behavior) and update the public LocalCodeActProviderOptions.Environment doc to describe the explicit empty-dictionary opt-in for a scrubbed environment. - Tests: * FileMountHelperTests covers per-file, per-mount, and total capture-limit branches that return TextContent omissions. * Integration tests cover unknown-tool dispatch error, tool throwing exception, and CodeValidator timeout that kills the process and raises CodeValidationException. - Sample: drop unused `Microsoft.Agents.AI.Foundry` using in Hosted-LocalCodeAct/Program.cs to satisfy IDE0005 check-format. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * chore(local-codeact-dotnet): remove stale orphan sample The dotnet/samples/LocalCodeAct/ scaffolding sample referenced APIs that don't exist in the current package (`ExecutionMode`, FileMount object-initializer syntax, the old LocalExecuteCodeFunction constructor signature, function.Metadata.*), produced a long list of check-format violations (CHARSET, IMPORTS, IDE0073 header, IDE0005 unused using, IDE1006 Async suffix, RCS1037 trailing whitespace), and did not match any of the documented sample layouts. The hosted-agent example at dotnet/samples/04-hosting/FoundryHostedAgents/responses/Hosted-LocalCodeAct is the supported entry-point sample for this package. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * style(local-codeact-dotnet): satisfy check-format rules - Add UTF-8 BOM to source files (CHARSET) - Remove unused using directives (IDE0005) - Simplify type names (IDE0001/IDE0002/IDE0090) - Rename static field JsonOptions -> s_jsonOptions (IDE1006) - Rename static field SyncRoot -> s_syncRoot (IDE1006) - Add missing this. qualifications in ProcessBridge (IDE0009) - Remove unused _options field from LocalCodeActProvider (IDE0052) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(local-codeact-dotnet): wire hosted sample into solution Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(local-codeact-dotnet): sync embedded Python scripts Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(local-codeact-dotnet): exercise Python integration on Windows Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Address LocalCodeAct API review feedback Move the required Python executable path to LocalCodeAct constructors, invert the validation flag default, and apply small project/file mount cleanup suggestions. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Address LocalCodeAct concurrency review Surface unauthorized mount traversal errors and use concurrent provider registries for LocalCodeAct tool and file mount CRUD operations. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Simplify LocalCodeAct function wrappers Use AIFunctionFactory-created inner functions for LocalCodeAct execute_code wrappers and remove redundant script cache and JsonNode cloning logic. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Update LocalCodeAct factory result tests Handle JsonElement result values produced by AIFunctionFactory delegation in LocalCodeAct execute_code integration tests. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
6cc7ddb73e |
.NET: Integrate LoopAgent into HarnessAgent with TodoCompletionLoopEvaluator (#6544)
* Add LoopAgent to Harness with TodoEvaluator sample * Address PR comments * Fix build error |
||
|
|
39f4b5ec72 | Align function tool names for BackgroundAgent and FileMemory between python and .net (#6550) | ||
|
|
02eb9435bf |
Bump litellm from 1.83.14 to 1.84.0 in /python (#6559)
Bumps [litellm](https://github.com/BerriAI/litellm) from 1.83.14 to 1.84.0. - [Release notes](https://github.com/BerriAI/litellm/releases) - [Commits](https://github.com/BerriAI/litellm/commits/v1.84.0) --- updated-dependencies: - dependency-name: litellm dependency-version: 1.84.0 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
4ff952e100 |
Python: Capture context provider instructions in agent telemetry (#6515)
* Fix agent instructions telemetry Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Simplify agent instructions telemetry guard Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix observability mypy cast Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
7bf2d2a6d0 |
Python: Fix harness console rendering one streamed tool call many times (#6549)
* Fix render issue for tools that are streamed in parts.
* Address PR review: missing call_id fallback, empty-mapping args, _is_complete perf
- Print call_id-less function calls as-is instead of merging under a name-derived
key (which could drop distinct unnamed calls).
- Preserve an empty {} mapping rather than coercing it to None.
- Add a structural bracket-balance gate before json.loads in _is_complete to
avoid O(n^2) re-parsing of growing streamed arguments.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
||
|
|
1a280ae7c5 |
Bugfix for Declarative Workflow (#6530)
* Declarative workflow bugfix * Fix PR comments |
||
|
|
4d492614a9 |
.NET samples: structural alignment changes (#6485)
* Rescope dotnet provider samples cleanup Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix provider samples README link Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
8e10c0399a |
Python: Remove unsupported as_agent function_invocation_configuration (#6520)
* Remove unsupported as_agent config parameter Fixes #6313 Remove the unsupported function_invocation_configuration parameter from BaseChatClient.as_agent(), which currently forwards an invalid kwarg into Agent.__init__(). This keeps the existing TypeError behavior for callers but changes the error source to the public API boundary, which we do not consider a breaking change. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix sample --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
bce2757477 | Foundry hosted agent responses emit failed events (#6502) | ||
|
|
106e065774 |
.NET: Rebuild Hyperlight sandbox after tool registry updates (#6523)
* .NET: Rebuild Hyperlight sandbox after tool registry updates Track provider tool registry updates in Hyperlight run snapshots so subsequent executions rebuild the sandbox after AddTools replaces registered tools. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Strengthen Hyperlight registry replacement test Add provider-level coverage that same-name AddTools replacement changes the captured execute_code snapshot fingerprint. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Use Guid for Hyperlight registry version Use a Guid token for Hyperlight tool registry generations to avoid overflow concerns in long-lived providers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Fix Hyperlight Guid test import Add the missing System import required by the Guid-based Hyperlight fingerprint test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
0db9305625 |
Python: Integrate tool approval into the harness (#6522)
* Integrate auto tool-approval feature into harness * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Rename disable_tool_approval to disable_tool_auto_approval Addresses PR review feedback that the parameter name was unclear. The flag toggles the auto/standing tool-approval middleware. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
571cae426c |
Python: Fix Azure AI Search citation URLs (#6453)
* Fix Azure AI Search citation URLs Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Enrich MCP search citation metadata Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix Azure AI Search citation enrichment follow-ups Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR #6453 review comments Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * updated filter for paths * also updated python paths * reverted dotnet-format change --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: westey <164392973+westey-m@users.noreply.github.com> |
||
|
|
9fb16e034f |
.NET: Allow custom argument marshaling for skill scripts (#6498)
Add an optional Func<JsonElement?, AIFunctionArguments> argument marshaler to inline and class-based skills so callers can customize how raw JSON tool-call arguments are converted into AIFunctionArguments before delegate invocation. This enables handling backends (e.g. vLLM) that send tool-call arguments as a JSON string instead of a JSON object. The marshaler can be supplied at the script, inline-skill, or class-skill level; when omitted, the existing strict JSON-object behavior is preserved unchanged. Co-authored-by: SergeyMenshykh <SergeMenshikh@outlook.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
e07cfba0f3 | Disable Anthropic tests by not providing environment vars until 404 failure is resolved (#6539) | ||
|
|
40a2dd5cd0 |
.NET: Restore ambient client-header scope between non-streaming ClientHeadersAgent runs (#6517)
* Restore ambient client-header scope between non-streaming runs (#6516) Make ClientHeadersAgent.RunCoreAsync async + await so the per-run ClientHeadersScope is unwound on return, matching the streaming path. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Assert per-run on wire instead of brittle exact request count Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
7e9c043c4c |
Python: Improve PR template and breaking-change label automation (#6473)
* Improve PR template and breaking-change label automation - Add a structured "Related Issue" section using GitHub closing keywords - Add a Review Guide prompt (major changes, impact, reviewer focus) with a note that the focus item is for human reviewers only - Add checklist items for issue linkage / no duplicate PRs and invert the breaking-change item (checked = not breaking) - Extend label-title-prefix to prepend [BREAKING] when the "breaking change" label is added - Add label-breaking-change workflow to apply the "breaking change" label when a PR title contains [BREAKING] Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Add pull-requests agent skill with dotnet/python links - Add root .github/skills/pull-requests/SKILL.md covering PR description authoring (following the PR template) and the review-comment workflow (review -> plan -> user review -> implement -> reply to all -> resolve) - Symlink the skill from python/.github/skills and dotnet/.github/skills - Reference the skill from python/AGENTS.md and dotnet/AGENTS.md Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fold breaking-change labeling into label-pr workflow Move the title -> 'breaking change' label logic into the existing label-pr workflow (which already applies the python/.NET labels) and drop the separate label-breaking-change workflow. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR title prefix review feedback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Pin patched MessagePack for .NET restore Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Revert MessagePack central pin Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Move title prefix tests out of tracked GitHub tests Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Exclude skill docs from CI path filters Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Match skill symlinks in CI path exclusions Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Exclude AGENTS docs from CI path filters Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Scope title-prefix normalization to a real prefix The normalization branch in addTitlePrefix matched ^Python (no colon), so titles like "Python samples improvements" or "Pythonic refactor" were treated as already-prefixed and only re-cased, never receiving the "Python: " prefix. Scope the match to ^<prefix>:\s* so only an actual existing prefix is normalized; otherwise the prefix is prepended. Same fix applies to the .NET prefix (e.g. ".NETStandard bump"). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
d7e8d2206d |
Python: Fix Python OTel usage detail attributes (#6493)
* fix python otel usage detail attributes Map cached/read/reasoning usage detail fields to standard OTel GenAI attributes while preserving provider-specific legacy keys. Add focused coverage for direct response spans, aggregated agent spans, and provider usage parsing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * address usage detail review feedback Omit missing OpenAI Responses usage detail counts while preserving zero-valued counts. Record zero-valued token usage in OTel histograms and add regression coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
d7027fc1f9 |
Python: [BREAKING] Align FileAccess tools with .NET — directory discovery and recursive search (#6476)
* Align FileAccess tools with .Net; add directory discovery and recursive search * Fix choices field description: spacing, line length, grammar Addresses PR review: separate concatenated string literals with proper spacing/newlines, wrap lines under the 120-char Ruff limit, and fix "doesn't" -> "don't". Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR comments --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
df0bd4da82 |
Python: Fix ollama_chat_client.py sample: pass tools via options dict (#6480)
* Fix ollama_chat_client.py sample: pass tools via options dict The sample was passing tools as a direct keyword argument to get_response(), which caused a TypeError. The tools parameter must be passed inside the options dict per the SupportsChatGetResponse protocol. Fixes #6411 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Wrap tools in a list as expected by OllamaChatClient _prepare_tools_for_ollama iterates the tools value, so it must be a list rather than a bare FunctionTool instance. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
ed4ff188fc |
Python: [Breaking] Additional bug fix for declarative workflows (#6489)
* Fix declarative object parsing bug * Remove unnecessary comment * Address PR comments * Address PR comments. * Fix CI failures. * declarative action approval bugfix * Address PR comments * Inlined single use variables. |
||
|
|
0f483fa968 |
Set ApplicationName on CosmosClientOptions for UserAgent telemetry (#6481)
Added CosmosOptionsHelper (in Microsoft.Agents.AI.CosmosNoSql namespace)
that sets CosmosClientOptions.ApplicationName per component, producing
wire-visible UserAgent suffixes:
- CosmosChatHistoryProvider: Microsoft.Agents.CosmosNoSql.ChatHistory/{version}
- CosmosCheckpointStore: Microsoft.Agents.CosmosNoSql.Checkpoint/{version}
This ensures Cosmos DB requests from the Agent Framework are identifiable
in telemetry, enabling usage tracking and diagnostics queries that can
distinguish between chat history and checkpoint workloads.
Addressed review feedback:
- Truncates ApplicationName to 64 chars (Cosmos SDK max length)
- Moved helper to Microsoft.Agents.AI.CosmosNoSql namespace (scoped ownership)
- Uses StringComparison.Ordinal for IndexOf call
When users provide their own CosmosClient instance, the ApplicationName
is not overridden - users retain full control.
Co-authored-by: TheovanKraay <TheovanKraay@users.noreply.github.com>
|
||
|
|
5e830f4dc9 |
.NET: Only use the output from the last message for structured output (#6499)
* Only use the output from the last message for structured output * Address PR comments * Address PR comment * Address PR comments |
||
|
|
1acd242550 |
Python: Add AgentLoopMiddleware for re-running agents in a loop (#6174)
* Python: Add AgentLoopMiddleware for re-running agents in a loop Add `AgentLoopMiddleware`, an `AgentMiddleware` that re-runs the wrapped agent in a loop. A single configurable class covers three common patterns, each with a convenience classmethod factory: - Ralph loop (`.ralph(...)`): no exit criteria, with feedback tracking (`record_feedback`/`progress`), progress injection (`inject_progress`), optional fresh context per iteration (`fresh_context`), and an early-stop completion signal (`is_complete`). - Predicate (`.with_predicate(...)`): loop while a `should_continue` callable returns True (e.g. paired with `todos_remaining`/`background_tasks_running`). - Judge (`.with_judge(...)`): a second chat client decides whether the original request was answered, using a `JudgeVerdict` structured-output response. The loop also auto-resolves pending function-approval / user-input requests via an `on_approval_request` callable (bounded by `max_approval_rounds`), and the next iteration's input is controlled by `next_message`. Supports both streaming and non-streaming runs. Exports `AgentLoopMiddleware`, `JudgeVerdict`, `todos_remaining`, and `background_tasks_running`. Adds tests, a sample, and docs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Refine AgentLoopMiddleware API and sample - with_judge: add criteria list with {{criteria}} templating into judge instructions plus an agent-side instruction; add fresh_context, additional judge feedback relay; default judge max_iterations. - should_continue is now required and positional; supports (bool, str|None) feedback tuples surfaced to next_message/record_feedback via feedback kwarg. - Judge forwards full multi-modal request and response messages. - Default max_iterations=10 (explicit None = unbounded); removed is_complete and Ralph terminology; ShouldContinueResult is a real TypeAlias. - Sample: stream all loops, print iteration counts via injected user-block boundaries (robust to function calling), <role>: content formatting, per-method expected output, and a looping todo sample. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix CI checks for AgentLoopMiddleware - Resolve pyright errors in _loop.py: drop the always-true final_result None check (the while loop always assigns it) and cast finish_reason to the AgentResponse constructor's expected type. - Apply pyupgrade --py310-plus: import TypeAlias from typing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Resolve mypy/pyright disagreement on finish_reason pyright infers AgentResponse.finish_reason as including str and rejects the direct assignment, while mypy considers a cast redundant. Drop the cast and suppress only pyright with a targeted reportArgumentType ignore, satisfying both type checkers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Add todo+judge AgentLoopMiddleware sample Add a second AgentLoopMiddleware sample that composes two criteria in one should_continue predicate: a TodoProvider check (evaluated first) and a report-style judge chat client (evaluated once todos are complete) that grades the assembled report against shared requirements. Register it in the middleware samples README. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Compose todo+judge loops as two middleware Rework the todo+judge sample to compose two AgentLoopMiddleware on the agent itself (middleware=[judge_loop, todo_loop]) instead of a single hand-written predicate. The inner todos_remaining loop drafts the report todo-by-todo and the outer with_judge loop re-runs it until an editor chat client judges the report publication-ready, reusing the built-in helpers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Reset session for fresh_context loops via snapshot/restore AgentLoopMiddleware.fresh_context previously only reset context.messages, so with an attached session each iteration still reloaded the local transcript or re-threaded the service-side conversation id and the model saw the accumulated history. Snapshot the session once before the loop (via to_dict) and restore it (from_dict + field copy) between iterations, so every pass starts from the pre-loop baseline. The final iteration's pass is persisted (no restore after the terminating iteration), so a subsequent agent.run continues from there. Removed the obsolete warning, updated docstrings and core AGENTS.md, and added tests: a snapshot/restore round-trip, a session-reset streaming x fresh_context x inject_progress x store matrix across multiple runs and loop iterations, and response_format parsing across the loop. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Updated samples and docstrings --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
3f77c555cf |
.NET: [BREAKING] Align FileAccess tools with Python; add directory discovery and recursive search (#6474)
* Align FileAccess with python and improve functionality * Addressing PR comments |
||
|
|
cd512da731 |
.NET: Updating MessagePack to latest version (#6497)
* Updating MessagePack to latest version * Remove MessagePack from package directly, since CentralPackageTransitivePinningEnabled is true |
||
|
|
76b2b1bf39 |
Python: Add opt-in AG-UI thread snapshot persistence and hydration (#6471)
* feat(ag-ui): add thread snapshot store primitives Key decisions:\n- Introduce an AGUIThreadSnapshot model limited to replayable messages, optional Shared State, and optional interrupt state.\n- Define AGUIThreadSnapshotStore as an async protocol keyed by explicit Snapshot Scope and AG-UI Thread id.\n- Add InMemoryAGUIThreadSnapshotStore as memory-only, latest-only, bounded local/demo/test storage; no file-backed store is introduced.\n- Require snapshot_scope_resolver whenever an endpoint is configured with a snapshot store, including pre-wrapped runners, so thread ids are not authorization boundaries.\n\nFiles changed:\n- packages/ag-ui/agent_framework_ag_ui/_snapshots.py\n- packages/ag-ui/agent_framework_ag_ui/__init__.py\n- packages/ag-ui/agent_framework_ag_ui/_agent.py\n- packages/ag-ui/agent_framework_ag_ui/_workflow.py\n- packages/ag-ui/agent_framework_ag_ui/_endpoint.py\n- packages/core/agent_framework/ag_ui/__init__.py\n- packages/core/agent_framework/ag_ui/__init__.pyi\n- packages/ag-ui/tests/ag_ui/test_snapshots.py\n- packages/ag-ui/tests/ag_ui/test_endpoint.py\n- packages/ag-ui/tests/ag_ui/test_public_exports.py\n- packages/ag-ui/AGENTS.md\n\nVerification:\n- uv run pytest packages/ag-ui/tests/ag_ui/test_snapshots.py packages/ag-ui/tests/ag_ui/test_public_exports.py packages/ag-ui/tests/ag_ui/test_endpoint.py::test_endpoint_requires_snapshot_scope_resolver_when_store_configured packages/ag-ui/tests/ag_ui/test_endpoint.py::test_endpoint_accepts_snapshot_store_with_scope_resolver -q\n- uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_endpoint_requires_snapshot_scope_resolver_when_store_configured packages/ag-ui/tests/ag_ui/test_endpoint.py::test_endpoint_requires_snapshot_scope_resolver_when_wrapped_runner_has_store packages/ag-ui/tests/ag_ui/test_endpoint.py::test_endpoint_accepts_snapshot_store_with_scope_resolver -q\n- uv run poe syntax -P ag-ui -C\n- uv run poe pyright -P ag-ui\n- uv run poe syntax -P core -C\n- uv run poe pyright -P core\n- uv run poe typing -P ag-ui\n- uv run poe typing -P core\n- uv run poe test -P ag-ui\n- uv run poe check -P ag-ui\n- git diff --check\n- git diff --cached --check\n\nBlockers / next iteration:\n- No blockers. Next slice can use the store contract to capture and hydrate agent snapshots.\n- uv repeatedly refreshed azure-ai-projects in uv.lock during local runs; reverted the generated lockfile churn because this change does not alter dependencies.\n- The poe-check commit hook was skipped after manual verification because it reformatted unrelated core MCP files outside this task. * feat(ag-ui): hydrate agent threads from snapshots Key decisions: - Resolve Snapshot Scope per endpoint request and pass it to the AG-UI runner only when snapshot storage is active. - Treat empty messages with no resume payload as an agent Hydrate Request when a scoped snapshot store is configured, replaying stored Shared State and message snapshots without invoking the wrapped agent. - Save the latest replayable agent message snapshot and Shared State at normal completion under Snapshot Scope plus AG-UI Thread id; no durable or file-backed store is introduced. Files changed: - packages/ag-ui/agent_framework_ag_ui/_agent_run.py - packages/ag-ui/agent_framework_ag_ui/_endpoint.py - packages/ag-ui/agent_framework_ag_ui/_snapshots.py - packages/ag-ui/tests/ag_ui/test_endpoint.py Verification: - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_stored_thread_snapshot_without_invoking_agent -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_stored_thread_snapshot_without_invoking_agent packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_snapshots_by_scope_and_thread -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_endpoint_empty_messages packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_stored_thread_snapshot_without_invoking_agent packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_snapshots_by_scope_and_thread -q - uv run poe syntax -P ag-ui -C - uv run poe pyright -P ag-ui - uv run poe typing -P ag-ui - uv run poe test -P ag-ui - uv run poe check -P ag-ui - git diff --check - git diff --cached --check Blockers / next iteration: - No blockers. Next slice can reconstruct normal new-user agent turns from stored snapshots. - uv repeatedly refreshed azure-ai-projects in uv.lock during local runs; reverted the generated lockfile churn because this change does not alter dependencies. - The poe-check commit hook was skipped after manual verification because it refreshed unrelated uv.lock dependency resolution. * feat(ag-ui): reconstruct agent turns from snapshots Key decisions: - Load scoped thread snapshots for non-hydrate agent requests only when snapshot storage is active and no resume payload is present. - Rebuild prior AG-UI history from stored snapshot messages, preserving the incoming new user suffix and treating stored snapshot content as authoritative over conflicting prior client history. - Merge stored Shared State with request state overrides before schema defaults and existing state-context injection. Files changed: - packages/ag-ui/agent_framework_ag_ui/_agent_run.py - packages/ag-ui/tests/ag_ui/test_endpoint.py Verification: - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_prepends_stored_snapshot_for_new_user_turn -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_deduplicates_full_history_and_merges_fresh_state -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_endpoint_empty_messages packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_stored_thread_snapshot_without_invoking_agent packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_snapshots_by_scope_and_thread packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_prepends_stored_snapshot_for_new_user_turn packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_deduplicates_full_history_and_merges_fresh_state -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py -q - uv run poe syntax -P ag-ui -C - uv run poe pyright -P ag-ui - uv run poe test -P ag-ui - uv run poe check -P ag-ui - uv run poe typing -P ag-ui - git diff --check - git diff --cached --check Blockers / next iteration: - No blockers. Next slice can enable workflow AG-UI Thread Snapshot persistence and hydration. - uv repeatedly refreshed azure-ai-projects in uv.lock during local runs; reverted the generated lockfile churn because this change does not alter dependencies. - The poe-check commit hook was skipped after manual verification because it refreshes unrelated uv.lock dependency resolution. * feat(ag-ui): hydrate workflow threads from snapshots Key decisions: - Handle workflow Hydrate Requests before resolving or invoking the wrapped workflow when snapshot storage and Snapshot Scope are active. - Capture only replayable workflow protocol data: workflow-emitted state snapshots, workflow-emitted message snapshots, and synthesized messages from text/tool output. - Keep workflow snapshot capture inactive without configured persistence, and skip saving snapshots when the workflow stream emits RUN_ERROR. Files changed: - packages/ag-ui/agent_framework_ag_ui/_workflow.py - packages/ag-ui/tests/ag_ui/test_endpoint.py Verification: - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_workflow_endpoint_hydrates_emitted_snapshots_without_invoking_workflow packages/ag-ui/tests/ag_ui/test_endpoint.py::test_workflow_endpoint_hydrates_synthesized_text_and_tool_snapshot -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py -q - uv run pytest packages/ag-ui/tests/ag_ui/golden/test_scenario_workflow.py -q - uv run poe syntax -P ag-ui -C - uv run poe pyright -P ag-ui - uv run poe test -P ag-ui - uv run poe typing -P ag-ui - uv run poe check -P ag-ui - git diff --check - git diff --cached --check Blockers / next iteration: - No blockers. Next slice can preserve interruption state and protect snapshots on errors across agent and workflow endpoints. - uv repeatedly refreshed azure-ai-projects in uv.lock during local runs; reverted the generated lockfile churn because this change does not alter dependencies. - The poe-check commit hook was skipped after manual verification because it refreshes unrelated uv.lock dependency resolution. * feat(ag-ui): preserve interrupted thread snapshots Key decisions: - Capture workflow RUN_FINISHED interrupt metadata in replayable AG-UI Thread Snapshots so Hydrate Requests can restore pending workflow actions without invoking or resuming the workflow. - Keep failed agent and workflow runs from replacing the last good snapshot; RUN_ERROR streams leave the previous snapshot available for hydration. - Verify interruption hydration through endpoint-level AG-UI streams for both agent and workflow wrappers, including Shared State replay and no wrapped runner invocation. Files changed: - packages/ag-ui/agent_framework_ag_ui/_workflow.py - packages/ag-ui/tests/ag_ui/test_endpoint.py Verification: - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_workflow_endpoint_hydrates_interrupted_thread_without_invoking_workflow -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_hydrates_interrupted_thread_without_invoking_agent packages/ag-ui/tests/ag_ui/test_endpoint.py::test_agent_endpoint_run_error_does_not_overwrite_previous_snapshot packages/ag-ui/tests/ag_ui/test_endpoint.py::test_workflow_endpoint_hydrates_interrupted_thread_without_invoking_workflow packages/ag-ui/tests/ag_ui/test_endpoint.py::test_workflow_endpoint_run_error_does_not_overwrite_previous_snapshot -q - uv run pytest packages/ag-ui/tests/ag_ui/test_endpoint.py -q - uv run pytest packages/ag-ui/tests/ag_ui/golden/test_scenario_workflow.py -q - uv run poe syntax -P ag-ui -C - uv run poe pyright -P ag-ui - uv run poe test -P ag-ui - uv run poe typing -P ag-ui - uv run poe check -P ag-ui - git diff --check - git diff --cached --check Blockers / next iteration: - No blockers. Next slice can document AG-UI Thread Snapshot security and usage. - uv repeatedly refreshed azure-ai-projects in uv.lock during local runs; reverted the generated lockfile churn because this change does not alter dependencies. - The poe-check commit hook was skipped after manual verification because it refreshes unrelated uv.lock dependency resolution. * docs(ag-ui): document thread snapshot security Key decisions: - Document AG-UI Thread Snapshot persistence as opt-in and disabled unless a snapshot_store is configured. - Place Snapshot Scope guidance next to endpoint authentication guidance, making clear that AG-UI Thread ids identify threads but do not authorize snapshot access. - Describe built-in storage as in-memory only, process-local, latest-only, and not durable production storage; durable stores remain app-owned implementations of AGUIThreadSnapshotStore. - Call out snapshot confidentiality impact and that no file-backed AG-UI snapshot store is provided. Files changed: - packages/ag-ui/README.md Verification: - uv run python scripts/check_md_code_blocks.py packages/ag-ui/README.md --no-glob - git diff --check - git diff --cached --check - commit hook without SKIP ran changed-package lint/format and AG-UI README markdown-code-lint successfully before stopping because uv.lock was modified - uv run poe markdown-code-lint (failed due existing unrelated packages/mistral/README.md missing agent_framework_mistral import resolution; changed AG-UI README blocks passed) Blockers / next iteration: - No blockers. Local issue/PRD planning artifacts remain uncommitted. - uv refreshed azure-ai-projects in uv.lock during markdown lint and the commit hook; reverted the generated lockfile churn because this documentation change does not alter dependencies. - The poe-check commit hook was skipped after manual verification because it refreshes unrelated uv.lock dependency resolution. * fix(ag-ui): harden thread snapshot persistence edge cases - Persist the completed confirm_changes turn with interrupt=None so hydration no longer replays a stale pending interrupt after the user responds; resume requests prepend stored history so the persisted thread is not truncated. - Defer endpoint default_state application to the runners when snapshot persistence is active, filling only keys missing from both the stored snapshot state and the request state so defaults never reset persisted Shared State. - Always fold the turn's output into the persisted messages snapshot even when the outbound MESSAGES_SNAPSHOT event is suppressed for predictive tools without confirmation. - Load the stored snapshot on workflow follow-up turns, reconstruct full thread history into the run input, and seed the snapshot builder with merged state so saving a new turn no longer replaces prior history. - Move snapshot message reconstruction helpers to _run_common for reuse by the workflow runner; load stored agent snapshots on resume turns for state merge. - Add endpoint regression tests for all four scenarios. * fix(ag-ui): protect snapshot history on resume and harden suffix trust - Prepend stored thread history when persisting snapshots for resume runs on both the agent and workflow paths, so a resumed interrupt no longer overwrites the stored thread with just the resume turn's output. - Filter the incoming message suffix during thread reconstruction: only user turns and tool results answering backend-issued tool calls (stored tool calls or pending interrupts) may extend authoritative history. Client-forged assistant and tool messages are dropped and logged instead of being persisted and replayed. - Close the workflow snapshot builder's tool-call group when a tool result or text message lands, so synthesized transcripts keep tool results adjacent to their tool_calls message and stay valid as provider replay history. - Export DEFAULT_MAX_THREAD_SNAPSHOTS from agent_framework_ag_ui and expose SnapshotScopeResolver through the core ag_ui facade and stub. - Add regression tests for agent and workflow resume history preservation, forged suffix rejection, builder tool-call grouping, and the export surface. * fix(ag-ui): tolerate snapshot save failures and scope workflow cache - Wrap snapshot_store.save() on both the agent and workflow paths so a transient store failure (timeout, connection refused) is logged instead of propagating. Previously a failing save converted an already-streamed successful run into RUN_ERROR, and on the workflow path emitted RUN_ERROR after RUN_FINISHED, violating the single-terminal-event invariant. The previous snapshot stays available for hydration. - Key the workflow_factory instance cache by (snapshot_scope, thread_id). The Snapshot Scope is the authorization boundary, so the same thread id under different scopes no longer shares an in-memory workflow instance. clear_thread_workflow accepts an optional snapshot_scope and clears all scopes for the thread when omitted. - Add tests: save-failure tolerance for agent and workflow endpoints, scope-isolated workflow cache, async snapshot_scope_resolver support, and in-memory store key validation errors. * fix(ci): ignore all dotnet.microsoft.com links in linkspector The existing ignore pattern only matched https://dotnet.microsoft.com/download, but Microsoft sites insert a locale segment between host and path (e.g. /en-us/download/dotnet/10.0), so localized links slip past the pattern and get checked. dotnet.microsoft.com bot-blocks CI link checkers with intermittent 403s across the whole site, which fails markdown-link-check on unrelated pull requests since linkspector scans the entire repository. Ignore the domain wholesale, matching how platform.openai.com is already handled for the same reason. A 403 from bot blocking is indistinguishable from a removed page, so the checker cannot produce a meaningful signal for this domain either way. * ag-ui: simplify raw_messages assignment and drop OrderedDict - Replace list(cast(...)) with a typed annotation for raw_messages (_agent_run.py:866) per review suggestion - Replace OrderedDict with a plain dict in InMemoryAGUIThreadSnapshotStore (_snapshots.py:136); regular dicts are insertion-order-safe since Python 3.7, so OrderedDict is unnecessary. Update _evict_oldest to use next(iter(...)) for FIFO removal instead of popitem(last=False). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review feedback for #2458: review comment fixes --------- Co-authored-by: Copilot <copilot@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |