Commit Graph

1428 Commits

Author SHA1 Message Date
Theo van Kraay a057cd505c Python: Add agent-framework-azure-cosmos-memory context provider (#6719)
* Add agent-framework-azure-cosmos-memory context provider (draft)

Introduces CosmosMemoryContextProvider, a ContextProvider that wraps the azure-cosmos-agent-memory toolkit to give agents long-term, Cosmos DB-backed memory (fact/procedural recall + user summaries). Includes package scaffolding, unit tests (mocked client), live Azure integration tests (marked), samples, README, and AGENTS.md.

Draft: uv.lock is intentionally left unchanged. This package depends on azure-cosmos-agent-memory (requires Python >=3.11), which is unsatisfiable against the workspace's current >=3.10 floor, so adding it to the shared lock requires a workspace decision (raise floor to 3.11 or exclude from workspace). Test coverage to be expanded.

* ci: exclude azure-cosmos-memory from uv workspace resolution

The package depends on azure-cosmos-agent-memory which requires Python
>=3.11 and a prompty pre-release (>=2.0.0a9). Both are unsatisfiable
against the workspace's >=3.10 floor and pre-release policy, causing
uv sync to fail in every Python CI job. Exclude the package from the
shared workspace so it is resolved and tested as a standalone package.

* ci: fix code-quality failures for azure-cosmos-memory

- Strip trailing whitespace from package files (pre-commit trailing-whitespace hook)
- Exclude the package README from markdown-code-lint: the package is excluded
  from the uv workspace, so its README snippets import a module that is not
  installed in the workspace env and Pyright cannot resolve it

* Exclude azure-cosmos-memory README from markdown-code-lint task

* Address PR review comments on cosmos-memory context provider

- Wire credential into Cosmos and AI Foundry clients; let toolkit own
  DefaultAzureCredential when none supplied (remove dead import).
- Honor auto_extract=False by zeroing extraction/summary cadence thresholds.
- Skip whitespace-only conversation turns and store stripped content.
- Show confidence 0.0 and coerce confidence to float in _format_memories.
- Register both 'integration' and 'azure' pytest markers accurately.
- Fix duplicated install block in README.
- Update and extend unit tests for new credential wiring and fixes.

* Include azure-cosmos-memory in the uv workspace

Follow the github_copilot pattern for a package with a Python 3.11-only
dependency: lower requires-python to >=3.10 and gate azure-cosmos-agent-memory
behind a python_version >= '3.11' marker. Add a direct, gated prompty
pre-release dependency so the workspace's if-necessary-or-explicit prerelease
policy permits the toolkit's transitive prompty requirement. Guard the test
modules with pytest.importorskip so the 3.10 CI leg skips cleanly. Remove the
workspace exclude and the markdown-code-lint exclude, and regenerate uv.lock.

* Address review feedback on cosmos-memory provider

Rename provider parameters to match Agent Framework conventions:
foundry_endpoint (was ai_foundry_endpoint) and embedding_model/chat_model
(were *_deployment_name). Move DEFAULT_* to module-level constants, type
memory_types as a Literal, use DEFAULT_CONTEXT_PROMPT as the default value,
and add ProcessorConfig/CosmosMemorySettings TypedDicts. Resolve connection
settings via agent_framework load_settings with required-field validation,
replacing the manual getenv/raise blocks. Scope user_id/thread_id to the
provider state and drop the unpreventable first-turn warning.

Rewrite the samples around Agent (not raw SessionContext), provider-scoped
state, and session-id threading; use PEP 723 inline dependencies instead of a
samples dependency group; use a plain input() loop; remove the dead custom
processor stub. Update README/AGENTS for the renamed parameters and env vars.
Add a samples ruff per-file-ignores entry now that the package is linted in CI.

* Add emulator-backed vector search integration test

Bump azure-cosmos-agent-memory to >=0.2.0b2 (adds the embeddings/chat client
injection seam) and add tests/test_emulator.py: an integration (not azure)
suite that exercises real Cosmos vector search with a quantizedFlat index
against a local Cosmos DB emulator, using deterministic in-memory fakes for
embeddings and chat so no Azure AI Foundry account or LLM is required.

To run on a stock emulator the fixture strips the toolkit's full-text index
(the provider only does pure vector search) and requests provisioned autoscale
throughput instead of serverless. The suite skips cleanly when no emulator is
reachable.

* Fix CI typing and package checks for azure-cosmos-memory

The package recently joined the uv workspace, so its source and tests are now covered by the Test Typing Checks and Package Checks gates for the first time.

tests: rename stale constructor kwargs to the current provider API (foundry_endpoint/embedding_model/chat_model); use a typed _STUB_AGENT for the unused agent param so pyright/pyrefly/ty/zuban all accept it; make processor_config values ints; assert non-None memory_client in the emulator tests.

source: relax reportUnknown*/reportOptional* for this package only (the toolkit ships no py.typed; mirrors the hosting-telegram precedent); decouple the conditional toolkit import from the annotation type; use settings.get(); fix memory_types list invariance; drop a redundant None guard; read role via getattr.

* Apply pyupgrade: single-arg AsyncGenerator in test_integration

* Make Cosmos memory extraction drain transparently on provider exit

The provider now drains in-flight background memory extraction in __aexit__, so applications no longer need to call flush() in their own control flow; the client's close() would otherwise cancel pending extraction tasks. flush() is hardened against clients that expose no usable background-task registry.

sample: interactive_chat reads input via asyncio.to_thread so the event loop stays free and background extraction runs during the session; removes the manual flush now that the provider drains on exit.

tests: add explicit transparent-extraction integration tests (emulator: after_run schedules extraction and __aexit__ drains it; live Azure: a fact is extracted and recalled in a later session with no manual flush). Emulator tests reuse a single fixed database to avoid exhausting the emulator's partition budget across runs.

* Add custom extraction-prompt seam and sample to cosmos-memory provider

Adds a prompts_dir option to CosmosMemoryContextProvider that points the Agent Memory Toolkit pipeline at a caller-supplied directory of Prompty templates, so callers can override extract_memories.prompty to control what the extraction LLM produces. The toolkit exposes no public prompts-directory seam, so the provider contains the one internal touch (swapping the pipeline's template loader after the store connects); applies to both provider-built and supplied clients.

sample: interactive_chat_custom_extraction.py - the interactive chat wired with a custom coding-assistant extraction rubric. It derives a complete prompts directory at runtime (copies the bundled templates and augments extract_memories.prompty) so it stays schema-compatible with the installed toolkit.

tests: unit tests assert the provider redirects the pipeline loader only when prompts_dir is set; an emulator integration test proves end to end that a unique marker in a custom extract_memories.prompty reaches the extraction LLM call.

* docs: document prompts_dir custom-extraction seam in cosmos-memory README

Replaces the stale, non-functional CustomMemoryProcessor snippet with the working prompts_dir approach, lists the new interactive_chat_custom_extraction.py sample, and corrects the interactive-sample feature list.

* Address review: rename _new_session, drop defensive toolkit import guard

Sample (comment): rename _new_thread to _new_session in both interactive samples (a new session is the new thread).

Provider (comment): replace the _memory_toolkit_available flag + __init__ ImportError guard with a plain guarded import that re-raises a clear ImportError, matching the github_copilot package's pattern for its 3.11-only SDK. Kept requires-python >=3.10 (bumping this one workspace member to 3.11 would force the entire uv workspace lock floor to 3.11). Tests now run importorskip before importing the package, mirroring github_copilot.

* Pass cadence via cadence_thresholds instead of mutating os.environ

* Mark package alpha and drop private naming in samples

* Require Python 3.11 and inject user summary as untrusted context

* CI: exclude azure-cosmos-memory from uv sync on Python 3.10

* Re-trigger CI (flaky external link check)

* Require chat/embedding models instead of silent defaults

* Fix pyright: narrow resolved chat/embedding models to str

---------

Co-authored-by: Theo van Kraay <thvankra@microsoft.com>
2026-07-20 09:44:11 +00:00
Evan Mattson c66bb39ea2 Normalize durable workflow inputs (#7205) 2026-07-20 08:33:20 +00:00
Yufeng He 7c6b1e975f Python: fix compaction token count inflating non-ASCII text (#7124)
TokenBudgetComposedStrategy estimates tokens by feeding a JSON-serialized
message to the tokenizer, but _serialize_message() used ensure_ascii=True.
That escapes non-ASCII text into \uXXXX sequences, so CJK and other
non-Latin content is token-counted as the escape sequences rather than the
characters the model actually sees, inflating the estimate (~1.6x for mixed
Japanese, more for pure CJK) and skewing compaction/token-budget decisions.

Serialize with ensure_ascii=False, matching the ensure_ascii=False already
used elsewhere in this module. Only affects token estimation; the serialized
string is never stored or transmitted.

Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
2026-07-19 12:44:40 +00:00
Yufeng He 0d2925037d Python: preserve explicit null arguments in auto function calling (#7108)
* Python: preserve explicit null arguments in auto function calling

FunctionTool.invoke dumped validated arguments with model_dump(exclude_none=True),
which strips any argument the model set to null. A required nullable parameter
(e.g. unit: Literal["C","F"] | None) that the model deliberately sets to null was
therefore dropped, and the function failed to invoke on the missing argument.

Use exclude_unset instead: keep the arguments the model actually provided (null
included) and omit only the ones it left out, so the function's own defaults still
apply. Because the input model is generated from the function signature, its field
defaults match the signature defaults, so omitted optionals are unchanged.

Fixes #5934

Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>

* Python: extend null-arg fix to the auto function-calling path

The earlier change fixed FunctionTool.invoke, but _auto_invoke_function
(the path a model-emitted function_call actually takes) still ran
model_dump(exclude_none=True), so an explicit null for a required
nullable argument was still dropped and the call failed with a missing
argument. Switch it to exclude_unset to match invoke, and add a
regression test that drives _auto_invoke_function with an explicit null.

---------

Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
2026-07-19 12:43:41 +00:00
Eduard van Valkenburg b5e635ed4d Python: isolate hosted session snapshots (#7141)
* Python: isolate hosted session snapshots

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 75d26ffd-7dc7-46b3-9966-9aaebb7b6bc3

* Python: avoid duplicate conversation snapshots

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 75d26ffd-7dc7-46b3-9966-9aaebb7b6bc3

* added some notes in the docstring
2026-07-18 11:59:24 +00:00
Eduard van Valkenburg 1036fa7438 Python: docs: add self-hosting sample snippets (#7104)
* docs: add self-hosting sample snippets

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 636b7fb2-381a-4d62-b5cf-d029efa3ad20

* docs: use stable hosting sample ranges

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 636b7fb2-381a-4d62-b5cf-d029efa3ad20

* docs: address hosting sample review feedback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 636b7fb2-381a-4d62-b5cf-d029efa3ad20

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-17 22:10:28 +00:00
Eduard van Valkenburg 62da382082 Python: Optimize shared serialization paths (#7165)
* Optimize core serialization paths

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: cda42f21-2500-4f78-a527-d6eeaeef922b

* Streamline AG-UI serialization

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: cda42f21-2500-4f78-a527-d6eeaeef922b

* Document shared serialization guidance

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: cda42f21-2500-4f78-a527-d6eeaeef922b

* Bound serialization protocol cache

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: cda42f21-2500-4f78-a527-d6eeaeef922b
2026-07-17 22:09:53 +00:00
Eduard van Valkenburg 3604ba70f6 Python: Normalize chat finish reasons (#7105)
* Python: Normalize chat finish reasons

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 45b65bfa-8e36-47b0-99d9-ec58ec60ace1

* Preserve Copilot finish reasons

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 45b65bfa-8e36-47b0-99d9-ec58ec60ace1

* fix claude finish reasons
2026-07-17 22:05:36 +00:00
Eduard van Valkenburg bc59c72170 Python: Add A2A hosting helpers (#7050)
* Python: Add A2A hosting helpers

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5d6987cd-1b67-4ba1-8b54-3c50da6e7607

* Python: Preserve final A2A streaming output

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5d6987cd-1b67-4ba1-8b54-3c50da6e7607

* Python: Clarify A2A conversion boundary

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 5d6987cd-1b67-4ba1-8b54-3c50da6e7607

* Python: Document A2A sample auth boundary

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 5d6987cd-1b67-4ba1-8b54-3c50da6e7607

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-17 16:58:39 +00:00
VectorPeak 6afae2f9b4 Python: raise ValueError for malformed data URIs (#6916)
* Python: raise ValueError for malformed data URIs

* Address malformed data URI review feedback

---------

Co-authored-by: VectorPeak <VectorPeak@users.noreply.github.com>
Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
2026-07-17 11:40:20 +00:00
Benke Qu cad81923e3 Python: fix: concurrent_agents sample incorrectly treats output as list[Message] (#6548)
* Python: fix: concurrent_agents sample treats output as list[Message] but it is AgentResponse

The default ConcurrentBuilder aggregator yields AgentResponse, not
list[Message]. The sample incorrectly cast the output to list[Message]
and iterated it directly. Fix by checking isinstance(output, AgentResponse)
and iterating output.messages instead.

* fix: remove warning print per review feedback

* fix: address review - remove unused Message import, update docs and sample output

- Remove unused Message import
- Update docstring: default aggregator yields AgentResponse objects, not list[Message]
- Fix sample output: remove user prompt entry (aggregator returns only assistant messages)
- Renumber sample output entries (researcher=01, marketer=02, legal=03)

---------

Co-authored-by: Benke Qu <bequ@microsoft.com>
Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
2026-07-17 11:38:04 +00:00
Ahmed Muhsin 5ab8877ba5 Python: HITL respond-URL addressing from inside workflows (#7001)
* feat(durabletask): surface HITL respond-URL addressing to workflow executors

Let a workflow notify a human reviewer (e.g. email an approval link) from inside the
graph, without the caller threading the instanceId/requestId by hand.

- durabletask: the orchestrator injects host_context {instance_id, workflow_name,
  request_path_prefix} into each activity input; CapturingRunnerContext surfaces it as
  host_metadata. No new core API.
- azurefunctions: WorkflowHitlContext.from_context(ctx) builds the canonical
  respond/status URLs (returns None in-process so callers degrade gracefully).
  Re-exported through the agent_framework.azure lazy namespace.
- Nested sub-workflows: the address context (root instance + workflow name +
  accumulated {executor}~{ordinal}~ prefix) propagates down call_sub_orchestrator via a
  new SUBWORKFLOW_ADDRESS_KEY marker, so an executor at any depth builds a URL that
  targets the addressable top-level instance with a qualified request id. The per-child
  ordinal matches the read-side enumerate() index used by the status/respond endpoints.
  The marker is stripped from untrusted input alongside SUBWORKFLOW_INPUT_KEY
  (confused-deputy / info-leak guard).
- Samples 12 and 13 reworked into the retry-safe two-step notify pattern: the emitter
  generates an explicit request id and a downstream NotifyExecutor builds the URL and
  notifies, so failed upstream retries never produce a dead link.

Tests: unit coverage for the metadata round-trip, address/ordinal agreement (fan-out at
depth and nested prefix accumulation), marker stripping, and URL building; integration
tests assert the helper-built URL equals the server respondUrl and resumes the run, for
both the flat (12) and nested (13) samples.

* refactor(durabletask): read back request_info id instead of generating one in samples

Add WorkflowHitlContext.pending_request_id(ctx), an async helper that returns the
id request_info just generated (read from the runner context's pending request-info
events). This works on any host via the core RunnerContext protocol method, so it
needs no core change.

Samples 12 and 13 now call request_info() and read the id back to forward to the
NotifyExecutor, instead of minting a uuid by hand and passing request_id=. The
read-back happens in the same activity execution that generated the id, so the
pending request event and the notify message still commit together with the same id
(retry-safe; failed upstream retries notify no one).

* docs(azurefunctions): document request_info id read-back and notify safety

Tighten pending_request_id docstring to require calling it immediately after request_info, and explain why that is safe on the durable host (each executor runs in its own activity with its own runner context, so the pending set only holds this executor's requests and the newest is the one just emitted). Document the two-step notify pattern in the 12 and 13 sample READMEs, including the downstream-notifier retry safety and the nested address-prefix propagation.

* fix(python): resolve ty typing error and address PR review comments

- test_subworkflow_orchestration: replace the mypy-only type:ignore[arg-type] with a cast so the ty checker passes too (the other four checkers already honored the ignore).

- samples 12/13 README: guard the notify snippet against None before build_respond_url to match the documented graceful-degradation behavior.

- integration tests 12/13: reword comments that implied request_info now generates an explicit uuid4; it generates the id internally by default.

* fix(python): honor configurable Functions route prefix and address HITL PR review

- Resolve the route prefix from host.json (extensions.http.routePrefix, default api) in a new azurefunctions _routes module, used by both the server endpoints and WorkflowHitlContext, so a custom or empty routePrefix no longer 404s respond/status URLs (was hardcoded /api/ in four places).

- Extract respond/status URL construction into one shared builder called from _app.py and _hitl_context.py, removing the sync-by-test duplication.

- Broaden loopback detection (localhost, 127.0.0.0/8, 0.0.0.0, ::1, [::1]) via a _is_loopback helper so local links use http.

- Pin the host_context key names as shared constants in durabletask so producer and azurefunctions consumer cannot drift.

- Reword a stale base_url comment to reference WEBSITE_HOSTNAME.

- Add unit tests for the route module and loopback handling.

* refactor(azurefunctions): derive server-side HITL URLs from the request URL

The run and status endpoints now derive the base URL and route prefix from the incoming request URL (the value the host actually routed) via split_request_url, so the caller-visible respond/status URLs no longer depend on reading host.json on the server. The in-workflow helper keeps reading host.json since it has no request context. Replaces strip_route_prefix and updates its tests.

* test(durabletask): enforce sub-workflow ordinal and read-index agreement

Extract the read-side subworkflows grouping into a shared _index_subworkflows helper (used by the orchestrator) and add test_readside_index_matches_dispatch_ordinal, which round-trips a fan-out through that helper and asserts subworkflows[executor][ordinal] resolves to the child the dispatch stamped that ordinal onto. Turns the previously comment-only write-ordinal / read-index invariant into a shared, CI-enforced one.

* fix(azurefunctions): suppress bandit B104 on loopback host set
2026-07-16 22:06:25 +00:00
westey a376577263 Gradudate FileMemoryProvider (#7113) 2026-07-15 22:57:21 +00:00
Tao Chen b2549337ff Python: Best effort to serialize tool def to Json for observability (#7029)
* Best effort to serialize tool def to Json

* Fix formatting

* Fix tests

* Optimize json serialization

* Best effort: Add secret filtering

* Remove frozen set and only convert required fields

* Fix tests

* Address comments

* Fix typing
2026-07-15 20:43:36 +00:00
Chris Brown b5300fe0c0 Python: fix per-run additional_beta_flags leaking into Anthropic request kwargs (#7060)
* Python: fix per-run additional_beta_flags leaking into Anthropic request kwargs

_prepare_options copied every key from the caller-supplied options dict
into run_options except "instructions" and "response_format". A
per-run additional_beta_flags value is correctly folded into the betas
set by _prepare_betas, but the raw key was never excluded, so it
survived into run_options and was forwarded straight through to
AsyncMessages.create(), which rejects it with TypeError: got an
unexpected keyword argument 'additional_beta_flags'. Add it to the
exclusion set alongside the other framework-level keys.

Fixes #5764

* Exclude additional_beta_flags from filtered_kwargs too

Copilot's review on the original fix pointed out the exclusion only
covered the options-dict copy, not kwargs passed directly to
_prepare_options — so additional_beta_flags supplied as a raw kwarg
would still leak through and reproduce the same TypeError. Add the
same exclusion to filtered_kwargs for consistency, with a regression
test covering the kwarg path.

---------

Co-authored-by: Chris Brown <albatrossflyon1@gmail.com>
2026-07-14 23:02:55 +00:00
Evan Mattson c35a63ed8d Python: feat: cross-session origin attribution on context messages (#7041)
* Python: feat: cross-session origin attribution on context messages

Add an optional origin_session_id parameter to SessionContext.extend_messages
that propagates into the existing _attribution payload on
Message.additional_properties. Downstream context observers can use it to
detect when a provider injects content stored under a different session than
the requesting one.

Populate the field from the harness memory consolidation pipeline
(_harness/_memory.py) when injected topics include contributions from
sessions other than the current one. Add a self-contained sample observer
under samples/02-agents/context_providers/cross_session_observer.py
demonstrating how to subscribe to the signal.

Backward-compatible: omitting the parameter preserves the existing
attribution shape exactly. Tests added in test_sessions.py and
test_harness_memory.py cover the new parameter, the harness cross-session
case, and the same-session case.

Motivated by Dai et al., Stateful Agent Backdoor (arXiv:2605.06158, May
2026), which specifically surveys MAF in section 6.1 / Table 10.
See #5914 for design discussion.

Surfaced during independent audit conducted by @finnoybu (Ken Tannenbaum, AEGIS Initiative); [MEDIUM, python/packages/core].

* Address cross-session attribution review feedback

* Address follow-up review feedback

* Address follow-up review comments

* Address origin attribution review feedback

---------

Co-authored-by: finnoybu <21694570+finnoybu@users.noreply.github.com>
2026-07-14 17:34:52 +00:00
Giles Odigwe 56e9a8f74c Python: Make foundry toolbox MCP skills sample self-contained (#7099)
* Python: Make foundry toolbox MCP skills sample self-contained

Rework sample 12 (foundry_toolbox_mcp_skills) so users can build it from
zero with azd, mirroring samples 04 and 09:

- Bundle two single-file SKILL.md skills (support-style, escalation-policy)
  and a skills-only toolbox.yaml (with one connectionless code_interpreter
  tool, required by `azd ai toolbox create`).
- Rewrite the README as an azd-native, from-zero guide (create skills ->
  create toolbox -> set TOOLBOX_ENDPOINT -> run) and fix the stale
  MCPSkillsSource API description to match main.py.
- Switch config from TOOLBOX_NAME to the versioned TOOLBOX_ENDPOINT
  (.env.example, agent.yaml, agent.manifest.yaml); add .azdignore.

Also enable the sample to run unattended behind ResponsesHostServer:

- Forward disable_load_skill_approval / disable_read_skill_resource_approval
  / disable_run_skill_script_approval from FoundryToolbox.as_skills_provider()
  to the underlying SkillsProvider, so load_skill needs no approval round-trip
  (the Responses host runs without an AgentSession, which the default approval
  flow requires). main.py now uses as_skills_provider(disable_load_skill_approval=True).
- Add unit tests covering the default and overridden approval behaviour.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 4f14f83d-1868-45c1-be1a-12f49a58ac36

* Python: Address PR review on toolbox MCP skills sample

- Remove the unused parameters section from agent.manifest.yaml (TOOLBOX_ENDPOINT
  is supplied via environment_variables, matching sample 04).
- README: state the sample is self-contained directly instead of contrasting
  with the C# sample.
- README: describe skill discovery behaviour without naming the internal
  MCPSkillsSource class.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 4f14f83d-1868-45c1-be1a-12f49a58ac36

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-14 17:03:26 +00:00
westey ba0ad2d1d2 Gradudate ToolApprovalMiddleware (#7106) 2026-07-14 11:16:10 +00:00
Evan Mattson b123480b65 Python: Fix AG-UI workflow handoff replay results (#7102)
* Fix AG-UI workflow handoff replay results

Decisions:
- Reconcile finalized function results only for call IDs exposed in the current run, and skip results already emitted or never exposed.
- Treat message-derived function results as workflow responses only when their IDs match pending interrupts.

Files:
- Updated _workflow_run.py reconciliation and resume filtering.
- Added runner and public two-turn handoff acceptance coverage.
- Expanded finalized-response call-ID, privacy, and deduplication tests.

Verification:
- uv run poe test -P ag-ui
- uv run poe syntax -P ag-ui -C
- uv run poe pyright -P ag-ui
- uv run poe test-typing -P ag-ui

Notes:
- No blockers. The handoff sample remains unchanged; local PRD and issue files are not included.

* Prevent duplicate AG-UI workflow tool results

* Preserve finalized AG-UI tool results

* Use AG-UI text emission controls
2026-07-14 09:29:46 +00:00
t-anjan 1c0082721c Python: Fix structured value parsing for split text chunks (#6990)
* Fix structured value parsing for split text chunks

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Fix structured value parsing for split text chunks

---------

Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-07-14 08:39:11 +00:00
S3rj 6c0950adeb fix: clarify require_confirmation docstring to reflect confirm_changes HITL gating (#6884)
Co-authored-by: Sergey Borisov <sergey.borisov@dataimpact.io>
Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
2026-07-14 08:01:58 +00:00
Hasan Ghomi f1ba16e3fd Python: Fix Magentic manager duplicating conversation history (#6297)
* Python: Fix Magentic manager duplicating conversation history

_complete() reused one persistent AgentSession, so the default history provider
re-injected prior turns on top of the full prompt the manager already rebuilds
each call — duplicating task/facts/plan and compounding every round. Use a fresh
session per call; keep self._session only for
checkpointing. GroupChatOrchestrator is unaffected. Add a regression test and
update the session-propagation test.

* Python: Clean up Magentic manager per Copilot review (drop dead _session)

---------

Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
2026-07-14 08:01:18 +00:00
Giles Odigwe cba77e3cd0 Python: quiet A2AExecutor logging for unmapped content types (#7034)
* Python: quiet A2AExecutor logging for unmapped content types

Tool-use responses include function_call/function_result content that the A2A executor does not surface, causing a WARNING per tool call. Log these at DEBUG and skip instead, matching the outbound content-conversion convention used across the Python chat clients.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: ee74cc55-44df-4fcf-b38f-1f79f2600dfc

* Address PR review: assert debug log args and drop redundant cast

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: ee74cc55-44df-4fcf-b38f-1f79f2600dfc
2026-07-14 07:44:31 +00:00
Eduard van Valkenburg 7ca8bb55b6 Python: Add Telegram hosting helpers and samples (#7047)
* Python: Add Telegram hosting helpers

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 5d6987cd-1b67-4ba1-8b54-3c50da6e7607

* Python: Exclude Telegram samples from aggregate typing

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 5d6987cd-1b67-4ba1-8b54-3c50da6e7607

* Python: Address Telegram helper review feedback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 5d6987cd-1b67-4ba1-8b54-3c50da6e7607

* Python: Serialize Telegram webhook sessions

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: d76a9c32-d170-426d-a64f-b70958b08b12
2026-07-14 07:40:12 +00:00
Nick Brady 54617557e6 Update Foundry branding (#6999)
Replace user-facing Azure AI Foundry branding with Microsoft Foundry across docs, samples, comments, and display text while preserving technical identifiers.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
2026-07-14 06:44:26 +00:00
Syed Osama Ali Shah df198005fd Python: fix: preserve function-call name when merging streaming deltas (#6809)
* fix: preserve function-call name when merging streaming deltas

`Content._add_function_call_content` built the merged name with
`getattr(self, "name", getattr(other, "name", None))`. Because
`Content.__init__` always sets `self.name` (defaulting to `None`), the
attribute is never missing, so the `getattr` default is never consulted
and `other.name` is ignored. When two function_call contents are merged
and only the second carries the name -- e.g. a streaming delta where the
function name arrives after the first chunk -- the name was silently
dropped.

Use the same "either side" pattern already used for the sibling
`exception` field on the next line: `getattr(self, "name", None) or
getattr(other, "name", None)`. Extend the existing merge test to cover
the late-name and both-None cases.

* test: construct nameless function-call deltas via Content(...) directly

Per review (pyright `reportArgumentType`): `Content.from_function_call`
annotates `name: str`, so passing `name=None` to model a streaming delta
with no name yet tripped the typing gate. Build those nameless deltas
with the `Content("function_call", ...)` constructor instead (its `name`
param is `str | None`) — the factory just wraps that same constructor, so
the runtime objects and the merge assertions are unchanged.

---------

Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
2026-07-14 06:36:57 +00:00
westey 0ceca9a76a Add name collision warnings for auto-approvals (#7090)
Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
2026-07-14 06:35:38 +00:00
Evan Mattson 18b03ea487 Python: adjust checkpoint encoding handling (#6579)
* Adjust checkpoint encoding handling

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Refine checkpoint encoding handling

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Adjust checkpoint dict encoding

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Clarify checkpoint unpickler blocked globals

* Preserve typed HITL requests in durable activities

* test: skip flaky durabletask multi-turn integration test

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-14 15:21:12 +09:00
Evan Mattson 56c4425db2 Python: bridge AG-UI request state and session continuity (#7084)
* Python: bridge AG-UI request state into sessions

Decisions:
- Project resolved AG-UI Shared State into the per-run AgentSession without typed restoration.
- Preserve existing local/service session identifiers and keep AG-UI state out of provider metadata.

Files changed:
- packages/ag-ui/agent_framework_ag_ui/_agent_run.py
- packages/ag-ui/tests/ag_ui/test_endpoint.py

Verification:
- uv run poe check -P ag-ui
- uv run poe test -P ag-ui (912 passed)

Notes:
- Scoped cross-run Session Continuation State remains for the next dependent issue.

* Python: persist scoped AG-UI session continuity

Decisions:
- Store private Session Continuation State atomically in scoped thread snapshots and restore it through the core AgentSession contract.
- Exclude Shared State keys, all HistoryProvider buckets, and tool approval state; request overlays evict colliding private values.
- Finalize interrupted response streams before snapshotting so provider after_run mutations are included.

Files changed:
- packages/ag-ui/agent_framework_ag_ui/_agent_run.py
- packages/ag-ui/agent_framework_ag_ui/_snapshots.py
- packages/ag-ui/tests/ag_ui/test_endpoint.py
- packages/ag-ui/tests/ag_ui/test_snapshots.py
- packages/ag-ui/AGENTS.md

Verification:
- uv run poe test -P ag-ui (921 passed)
- uv run poe check -P ag-ui
- uv run poe typing -P ag-ui

Notes:
- Lifecycle, isolation, and broader storage guidance remain for the next dependent issue.

* Python: document AG-UI session continuity lifecycle

Decisions:
- Keep scoped thread snapshots as the single reset and continuity boundary, with missing request Shared State preserving private continuation.
- Document trusted typed-restoration storage, State Authorities, custom-store round trips, and one-active-run last-writer-wins consistency.
- Verify failure, hydration privacy, scope/thread isolation, and reset mechanics through public endpoint and store seams.

Files changed:
- packages/ag-ui/README.md
- packages/ag-ui/tests/ag_ui/test_endpoint.py
- packages/ag-ui/tests/ag_ui/test_snapshots.py

Verification:
- uv run pytest -q <focused lifecycle tests> (6 passed)
- uv run poe test -P ag-ui
- uv run poe syntax -P ag-ui -C
- uv run poe typing -P ag-ui
- uv run poe check -P ag-ui
- uv run poe markdown-code-lint

Notes:
- No runtime capability probe, secondary state store, locking, or configuration flag was added.
- No blockers remain for this lifecycle and guidance slice.

* Python: harden AG-UI session continuity

* Python: isolate AG-UI request state
2026-07-14 04:55:46 +00:00
Evan Mattson 13066cdf96 Python: Refine DevUI request logging (#7083)
* Refine DevUI request logging

* Address DevUI logging review feedback
2026-07-14 03:40:18 +00:00
westey f11cfd9d76 Switch FileAcessProvider on Harness to opt-in (#7094) 2026-07-14 02:26:59 +00:00
Peter Ibekwe 4bac2c2c05 Python: Promote python declarative workflows to stable version (#7065)
* Promote python declarative workflows to stable version

* Updated changelog with PR detail.

* Updated to address pr comments.

* Remove changelog update
2026-07-13 22:30:21 +00:00
pratik wayase 43568f1ef2 Fix: coalesce reasoning deltas into single block when content.id is None (#6804)
python/packages/ag-ui/tests/ag_ui/test_run_commoclear
:wq
2026-07-13 21:11:22 +00:00
VectorPeak 52005ff17d Python: accept AG-UI state data URI parameters (#6905)
* Python: Accept AG-UI state data URI parameters

* Python: Handle invalid AG-UI state base64

---------

Co-authored-by: VectorPeak <VectorPeak@users.noreply.github.com>
Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
2026-07-13 21:05:27 +00:00
pratik wayase a4e4a5a51c Python: Fix: Ollama parallel tool calls collide on same call_id (#6822)
* Fix: Ollama parallel tool calls collide on same call_id

* fix(ollama): use uuid4 for tool call IDs and support colons in tool names

---------

Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
2026-07-13 21:02:24 +00:00
Alireza Afzali c8fb491644 Python: docs: add env example files for durabletask samples (#5948)
* docs: add env example files for durabletask samples

* docs: clarify env example values and comments

* docs: set default Redis URL in streaming sample env example
2026-07-13 21:01:19 +00:00
westey c9b19e831f Gradudate mode and todo providers (#7053) 2026-07-13 13:47:53 +00:00
westey b3d523ee50 Python: [BREAKING] Fix harness before-strategy compaction under per-service-call persistence (#7055)
* Fix middleware ordering to ensure compaction runs

* Address PR comments

* Fix build issue
2026-07-13 09:51:00 +00:00
Tao Chen 8e74360d52 Python: Add Microsoft OpenTelemetry Distro sample (#5632)
* Add Microsoft OpenTelemetry Distro sample

* Verify and add README

* Add dependency header

* Add maf dependency in PEP 723 block
2026-07-13 01:51:15 +00:00
ByteWise 7f4cc296fd Python: preserve tool span context for parallel calls (#6512)
* Python: preserve tool span context for parallel calls

* Python: address parallel tool span review feedback

* Python: fix parallel tool span test checks

---------

Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
2026-07-13 00:07:53 +00:00
Benke Qu f3057ef20c Python: fix: clear service_session_id in _agent_wrapper when propagate_session=True (#5875)
* fix: clear service_session_id in _agent_wrapper when propagate_session=True

When propagate_session=True, the child agent inherits the parent's
service_session_id. After the parent's first LLM call, MAF auto-populates
this from the Responses API conversation_id. The child sends it as
previous_response_id which the server rejects because the parent's
tool_call is still pending (400 error).

This fix saves and clears service_session_id before calling the child
agent and restores it in a finally block, preserving session.state
sharing while isolating the server-side conversation pointer.

Fixes #5874

* refactor: use child session copy instead of in-place mutation

Address Copilot review comments:
- Create a child AgentSession with shared state dict but isolated
  service_session_id, avoiding race conditions under concurrent
  asyncio.gather tool invocations.
- Update tests to verify child gets a separate session object and
  that child-set service_session_id does not leak to parent.

* fix: update test_chat_agent_as_tool_propagate_session_true for child session isolation

The existing test asserted captured_session is parent_session, but since
we now create a separate child AgentSession (to avoid racing under concurrent
asyncio.gather), the child is a different object. Updated assertions to verify:
- child is NOT the parent object (isolation)
- child shares the same session_id and state dict (by reference)
- child's service_session_id is None (isolated)

* fix: add type narrowing asserts for captured_session

Add 'assert captured_session is not None' before attribute access to
satisfy mypy/pyright type checking on Optional values.

* Python: Fix test typing checks

---------

Co-authored-by: Benke Qu <bequ@microsoft.com>
Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
Co-authored-by: Evan Mattson <evan.mattson@microsoft.com>
2026-07-12 23:37:53 +00:00
Eduard van Valkenburg 68136ee081 Python: Clean up dependency groups and compatibility (#7046)
* Python: Clean up dependency management

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2f7b1c89-f3ff-418d-ab4e-4f014fda308f

* Python: Harden Mistral SDK import fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2f7b1c89-f3ff-418d-ab4e-4f014fda308f
2026-07-10 22:40:42 +00:00
Evan Mattson 87af313119 Python: [BREAKING]: Emit TOOL_CALL events for workflow participant tool calls in AG-UI (#7039)
* Python: emit participant tool calls in AG-UI workflows

Decisions:
- Pass function call, function result, and approval request content from streaming agent updates regardless of role.
- Preserve the assistant-role gate for text and reuse the shared AG-UI content emitters without dual custom-event emission.

Files changed:
- packages/ag-ui/agent_framework_ag_ui/_workflow_run.py
- packages/ag-ui/tests/ag_ui/test_workflow_run.py

Verification:
- uv run poe test -P ag-ui
- uv run poe pyright -P ag-ui
- uv run poe test-typing -P ag-ui
- uv run poe syntax -P ag-ui -C

Notes:
- Existing workflow golden scenarios do not exercise participant tool calls, so no snapshot changed.
- No blockers.

* Python: guard participant tool call duplication

Decisions:
- Assert the workflow stream emits one TOOL_CALL_START when a streamed call is also present in final conversation history.
- Keep production flow unchanged because latest-assistant final-response conversion prevents duplication.

Files changed:
- packages/ag-ui/tests/ag_ui/test_workflow_run.py

Verification:
- uv run pytest packages/ag-ui/tests/ag_ui/test_workflow_run.py -k 'participant_tool_call or repeat_tool_call' -q
- uv run poe test -P ag-ui
- uv run poe pyright -P ag-ui
- uv run poe test-typing -P ag-ui
- uv run poe syntax -P ag-ui -C
- git diff --check

Notes:
- No blockers; no call-id guard was required.

* Python: scope workflow tool content bypass to resumable tool calls

- Exclude approval request content from the role bypass. Workflow approvals
  resume through request_info pending state, so an approval interrupt emitted
  from streamed content would have no pending request to resume against.
- Admit mcp_server_tool_call and mcp_server_tool_result so provider-hosted MCP
  tool calls from workflow participants emit standard tool call events.
- Add unit tests for MCP passthrough, approval exclusion, and mixed
  text-plus-tool content in non-assistant updates.
2026-07-10 16:58:07 +00:00
Evan Mattson 9ac548ad15 Python: keep attachments close (#7038)
* Python: keep attachments close

* Python: close attachment edge cases
2026-07-10 16:57:29 +00:00
HaoJun 32a547a1a7 Python: bind AG-UI tool arguments to call IDs (#6342)
Co-authored-by: White-Mouse <15983334+White-Mouse@users.noreply.github.com>
2026-07-10 06:54:53 +00:00
Evan Mattson 7464a59228 Python: Bump Python package versions for 1.11.0 release (#7035)
* Bump Python package versions for 1.11.0 release

Bump the CHANGELOG-selected packages for the 1.11.0 release: core and the root package move to 1.11.0 for the new stable APIs, Foundry and OpenAI receive patch bumps, changed prerelease packages receive the 260709 stamp or next RC counter, and Monty joins the bump set for corrected published dependency metadata. No beta cohort bump was applied. Raise core floors conservatively on every package publishing this cycle and correct dependency floors exposed by lower-bound validation.

Copilot-Session: ee33d338-c1fc-4182-9106-0345ccf26b8e

* Fix Gemini streaming type suppression

Move the targeted Pyright suppression to the SDK contents argument, where the google-genai invariant content-list alias produces the compatibility diagnostic, and remove the now-unnecessary member suppression.

Copilot-Session: ee33d338-c1fc-4182-9106-0345ccf26b8e

* Raise Monty core dependency floor

Align Monty with the conservative release policy by requiring agent-framework-core 1.11.0 or later for the package version published in this cycle.

Copilot-Session: ee33d338-c1fc-4182-9106-0345ccf26b8e
2026-07-10 12:15:26 +09:00
Ethan qu 01ec3b7bcf Fix AG-UI approval thread aliases (#6908)
Co-authored-by: godququ5-code <256881196+godququ5-code@users.noreply.github.com>
2026-07-10 00:51:49 +00:00
westey ce96fd4b72 Python: Integrate message injection into harness agent (#7027)
* Integrate message injection into harness agent and sample console

* Add agents.md update.

* Address PR comments
2026-07-10 00:33:31 +00:00
Evan Mattson 52237b8eff Python: consolidate dependency updates (#7033) 2026-07-10 00:25:30 +00:00
Giles Odigwe e6cc2c09af Python: Fix read_skill_resource instruction dropping .md extension (#7031)
The RESOURCE_INSTRUCTIONS example told the model to use

eferences/FAQ instead of 
eferences/FAQ.md, contradicting the
actual exact-match resource lookup (which lists and matches names
including the extension). This caused read_skill_resource to fail with
'Resource not found'. Align the example with the .NET original.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-10 00:06:09 +00:00