* Python: Split type checkers by target (pyright source, 5 checkers on tests/samples) Rework the typing setup along the lines of the 'too many type checkers' approach: - Pyright (strict) is now the sole source-code type checker; mypy is removed from source and its [tool.mypy] block becomes a relaxed profile used only for tests/samples. - Tests are checked by all five checkers (pyright relaxed, mypy, pyrefly, ty, zuban); samples by pyright, pyrefly, and ty. All run in a relaxed/ basic profile so authors aren't forced into over-annotation. - Add pyrightconfig.tests.json and bump sample pyright configs to basic. - Unify test/sample typing onto the same parallel fan-out used by source pyright via run_command_items in task_runner.py. - Make version-conditional imports symmetric: keep or drop the '# type: ignore' on both branches so results match across interpreter versions (local vs CI). - Update SKILL.md, DEV_SETUP.md, and CODING_STANDARD.md for the five gating checkers and pyright on source+tests+samples. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix merge regressions from main (typing + runtime) Merging main into the type-checker split branch surfaced regressions that the new five-checker test suite and unit tests caught: Runtime fixes: - anthropic: restore the dropped `cache_read_input_token_count` mapping in _parse_usage_from_anthropic (lost during merge conflict resolution). - gemini: _get_function_calling_mode test helper returned str(enum) ('FunctionCallingConfigMode.AUTO') instead of the enum value ('AUTO'). - openai: _response_id_from_token test helper was an infinite self-recursion; return token['response_id']. - orchestrations: reset output_events per approval iteration so the terminal output assertion counts only the final run. - core: drop a stale duplicate harness test whose message ('non-negative') contradicted the source ('positive'). - purview: import PolicyLocation/PolicyScope/ProtectionScopeActivities/ ExecutionMode used by the processor tests. Type-checker fixes (tests, relaxed profile): - core: pyright/mypy/pyrefly/ty/zuban green-ups across the harness, MCP, observability and types tests. - anthropic/openai: route provider-namespaced UsageDetails keys through a dict cast (extra_items TypedDict unsupported by mypy/ty). - purview: typed model constructors and cache-mock casts. - ag-ui: annotate WorkflowContext[Any, Any] so yield_output accepts test payloads, guard Optional forwarded_props, and ty-ignore intentional bad args. Source pyright (sole source checker) flagged unnecessary ignores newly introduced by merged code in core _tools.py and declarative _declarative_base.py. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Isolate per-package mypy cache in test-typing fan-out The parallel test-typing fan-out runs many mypy processes concurrently, all defaulting to a single shared ./.mypy_cache. Concurrent writes corrupt the cache and mypy aborts with INTERNAL ERROR (intermittently, depending on worker timing) -- which is why CI's Test Typing job failed on a shifting set of packages while a single-package run was fine. Give each mypy invocation an isolated cache dir keyed by its target paths so incremental caching still works per package without races. Other checkers (zuban/pyrefly/ty/pyright) maintain their own caches and are unaffected. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Make lab pyright-only on source (drop source mypy) Lab was the last package still running mypy on its source code, requiring mypy-only `# type: ignore` comments that pyright (the sole source checker everywhere else) flags as unnecessary. Align lab with the rest of the monorepo: - Remove the lab source mypy poe tasks (mypy-gaia/lightning/tau2) and the now-dead strict [tool.mypy] config block. - Drop the 'Run lab mypy' CI step; lab source is type-checked by pyright only. Lab tests remain covered by the workspace test-typing fan-out (mypy, pyrefly, ty, zuban, pyright over tests using the relaxed root config). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix test-typing regressions from latest main merge A fresh merge from main brought in new test code never run under the five-checker test-typing suite. Green up across the affected packages: - core: narrow Optional span.attributes with 'and' guards in span filters and assert+cast the json.loads(...attributes[...]) reads (test_observability); match the existing as_agent ignore on the protocol-typed fixture (test_clients). - openai: align new streaming tests with the established chat_options dict pattern (ChatOptions TypedDict isn't assignable to dict), route Optional .annotations[0] access through a small _first_annotation helper (mirrors the file's assert-not-None convention), and annotate a mapped ResponseStream. - foundry_hosting: annotate error: dict[str, Any] = body.get(...) or {} (zuban needs the annotation). - foundry: narrow ignores for the live AIProjectClient credential arg (pyrefly) and connections.get_default (zuban) SDK type gaps. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * updated pyright version * pyright fix * Python: Fix source typing for pyright 1.1.410 Pyright 1.1.410 tightened several checks. Apply the same source fixes as upstream PR #6275: - anthropic: import AsyncAnthropicBedrock from anthropic.lib.bedrock and AsyncAnthropicVertex from anthropic.lib.vertex (no longer re-exported from the anthropic top-level package -> reportPrivateImportUsage). - core _types.py: cast the transform-hook result to UpdateT (reportAssignmentType). - core _workflows/_events.py: annotate the @contextmanager helper as Generator[None] instead of Iterator[None] (reportDeprecated). - redis: build the combined filter expression with an explicit loop instead of reduce(and_, ...), which pyright could no longer fully type (drops the now unused functools.reduce / operator.and_ imports). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Accept plain-text body in Azure Functions workflow/run endpoint The workflow_orchestrator already accepts plain strings as well as JSON objects via context.get_input(), but the start_workflow_orchestration HTTP handler only accepted JSON and returned 400 for any non-JSON body. This made the functions integration tests that POST text/plain to /api/workflow/run (e.g. test_09_workflow_shared_state) fail consistently with 400 != 202. Fall back to the raw request body (decoded as UTF-8) when the body is not JSON, rejecting only a truly empty body. The JSON path is unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
agent-framework-tools
Alpha built-in tools for the Microsoft Agent Framework. A home for first-party
Python tools that plug into any chat client's shell / function surface. The
first tool is LocalShellTool.
Installation
pip install agent-framework-tools --pre
LocalShellTool quick start
import asyncio
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient
from agent_framework_tools.shell import LocalShellTool
async def main() -> None:
client = OpenAIChatClient(model="gpt-5.4-nano")
async with LocalShellTool() as shell:
agent = Agent(
client=client,
instructions="You are a helpful assistant that can run shell commands.",
tools=[client.get_shell_tool(func=shell.as_function())],
)
result = await agent.run("Print the current working directory.")
print(result.text)
asyncio.run(main())
Modes
- Persistent (default): a single long-lived shell session.
cd,export, and shell functions persist across tool invocations. - Stateless (
mode="stateless"): each command runs in a fresh subprocess.
Safety
LocalShellToolis not a sandbox. It runs commands directly on the host with the agent process's privileges. The actual security boundary is approval-in-the-loop. For untrusted input use a sandboxed executor — seeagent-framework-hyperlight.
Defenses (in priority order):
- Approval-in-the-loop — every command surfaces as a
user_input_request; nothing runs without consent. Disabling this requiresacknowledge_unsafe=True. - Process-tree termination on timeout via
psutil, so child processes (make, watchers, network tools) cannot survive the timeout. - Output truncation to 64 KiB (head + tail with marker).
- Audit hook (
on_command=…) for SIEM / append-only logs. - Optional command-pattern filter via
ShellPolicy(denylist=[...], allowlist=[...]). Empty by default. This is a UX pre-filter, not a security boundary — operators are expected to supply patterns that match their workload (and they can be defeated by trivial obfuscation such as\rm -rf /,${RM:=rm} -rf /,python -c "…", encoded payloads, or PowerShell-native equivalents). Real isolation comes from approval gating and the sandbox tier (DockerShellTool). Seetests/test_security.pyfor the documented residual risk surface.
Override with ShellPolicy:
from agent_framework_tools.shell import LocalShellTool, ShellPolicy
shell = LocalShellTool(
policy=ShellPolicy(allowlist=[r"^ls\b", r"^cat\b", r"^git status$"]),
approval_mode="never_require",
acknowledge_unsafe=True, # required to bypass approval
)
Cross-OS
- Windows:
pwsh -NoProfile -Command -(falls back topowershell.exe). - Linux / macOS:
/bin/bash --noprofile --norc(falls back to/bin/sh). - Override via the
shell=constructor argument or theAGENT_FRAMEWORK_SHELLenvironment variable.
ShellEnvironmentProvider — context provider
A model talking to a PowerShell session will sometimes default to bash
syntax (export FOO=bar, ls -la, > /dev/null) and vice versa.
ShellEnvironmentProvider is an AIContextProvider that probes the live
shell once per session — family, version, OS, working directory, and a
configurable list of CLI tools (git, node, python, docker by
default) — and injects a system-prompt block describing the shell idiom
to use and the available CLIs.
from agent_framework_tools.shell import (
LocalShellTool,
ShellEnvironmentProvider,
ShellEnvironmentProviderOptions,
)
shell = LocalShellTool()
provider = ShellEnvironmentProvider(
shell,
ShellEnvironmentProviderOptions(probe_tools=("git", "uv", "node")),
)
agent = Agent(
client=client,
tools=[client.get_shell_tool(func=shell.as_function())],
context_providers=[provider],
)
Probe failures from expected error types (timeouts, policy rejections,
spawn failures) are recorded as None fields in the snapshot rather
than raised; a missing CLI never fails the agent. A failed first probe
does not poison the cache — the next call retries.
DockerShellTool — sandboxed tier
When commands originate from untrusted input (e.g. the model is acting on
prompt-injected document content), prefer DockerShellTool. With the
default isolation flags and a trusted container runtime, the container
is the intended security boundary and approval gating becomes optional.
import asyncio
from agent_framework_tools.shell import DockerShellTool
async def main() -> None:
async with DockerShellTool(
image="mcr.microsoft.com/azurelinux/base/core:3.0",
approval_mode="never_require", # container is the boundary
) as shell:
result = await shell.run("uname -a && id")
print(result.stdout)
asyncio.run(main())
Defaults applied to every container:
--network none— no host or external network.--user 65534:65534— runs asnobody:nogroup.--read-onlyroot filesystem; only mounted host paths are writable.--cap-drop ALLand--security-opt no-new-privileges.--memory 512m,--pids-limit 256, ephemeraltmpfs /tmp.
To expose a host directory, pass host_workdir="/path" (mounted
read-only by default; mount_readonly=False to allow writes). Swap the
container runtime with docker_binary="podman".
Sandbox tiers at a glance
| Use case | Tool | Sandbox |
|---|---|---|
| Run code (untrusted) | HyperlightCodeActProvider.execute_code (agent-framework-hyperlight) |
Hyperlight WASM microVM |
| Run shell (untrusted) | DockerShellTool |
OCI container (network-off, non-root, capabilities dropped) |
| Run shell (trusted dev) | LocalShellTool |
Approval-in-the-loop |
Relationship to agent-framework-hyperlight
agent-framework-hyperlight is a code sandbox (a single WASM guest
loaded into a microVM, called via a hostcall ABI — there is no kernel,
userland, or shell binary inside). It is the right tier for executing
generated code. For sandboxing shell commands, the realistic tier is
OCI, which DockerShellTool provides.