* Python: Add hosted agent sample for the agent harness * Disable file providers and fix call_server usage in hosted harness sample Addresses PR review: disable the harness file-memory and file-access providers so the headless sample doesn't expose file tools or write outside storage/, and correct the app.py docstring to match call_server.py (which takes no prompt argument). * Python: update hosted harness sample for current APIs --------- Co-authored-by: Evan Mattson <evan.mattson@microsoft.com>
local_responses_harness — hosting a harness agent behind Responses routes
The sibling of local_responses/, with one change: the
hosted target is a batteries-included harness agent built with
create_harness_agent instead of a plain
Agent.
Everything else is the same helper-first Responses hosting shape: one native
FastAPI route, a small SessionStore via AgentState, and the Responses helper
functions:
responses_to_run(...)responses_session_id(...)create_response_id(...)responses_from_run(...)responses_from_streaming_run(...)
The takeaway is that a harness agent is just an Agent, so it drops straight
into the same AgentState / Responses-helper seam as any other target. The
harness supplies the function-invocation loop, per-service-call history
persistence, context-window compaction, todo management, and heuristic tool
approval on top of a single @tool.
What the sample does with the harness:
- Turns off the interactive-only features (plan/execute mode and the Textual console) because a one-shot HTTP request has no console to drive.
- Turns off web search to keep the sample self-contained.
- Keeps todo management and compaction enabled, so the target is a genuine
harness agent and not just a relabelled plain
Agent. - Registers
lookup_weatherwithapproval_mode="never_require"so a headless run never blocks waiting for a human to approve a tool call.
What the route demonstrates (identical to local_responses/):
- Uses an explicit request-option allowlist. This sample only allows
max_tokensandreasoning; all other caller-supplied options, includingmodel,temperature,store,tools, andtool_choice, are denied by default. Your app decides the exact allowed, altered, and denied options. - Produces the AF messages, options, and session id that the route passes to
agent.run(...). - Stores each newly minted response id for response-keyed continuation, via
state.set_session(response_id, session)afteragent.run(...)has updated the session. OpenAI'sprevious_response_idrotates every turn by design — it lets a caller continue from any earlier response, not just the latest one — so every response id needs to stay independently resolvable. - Treats an unknown
conversation_idas a request to create a new local session. Your app can choose a stricter policy. - Explicitly advances a supplied
conversation_idafter each completed run. A conversation id is a mutable head, so production apps should serialize writers or use optimistic concurrency; the sample stores the updated session only under that stable conversation id. - Treats each
previous_response_idas an immutable snapshot. Multiple callers can branch from the same response concurrently because each receives a session copy and stores its result under a newly minted response id.
app:app is a module-level FastAPI ASGI app; recommended local launch is
Hypercorn.
Production readiness
This is not a full-fledged production deployment. Before exposing this pattern to callers, add authentication and authorization at the infrastructure layer, the FastAPI app layer, or inside the route body.
Session continuation deserves particular care: treat previous_response_id and
conversation_id as untrusted request values, authorize the caller before
loading or storing a session for those ids, and partition any durable session
store by tenant/user as appropriate for your application.
Run
export FOUNDRY_PROJECT_ENDPOINT=https://<your-project>.services.ai.azure.com
export FOUNDRY_MODEL=gpt-5-nano
az login
uv sync
uv run hypercorn app:app --bind 0.0.0.0:8000
Single-process for quick iteration:
uv run python app.py
Call locally
uv sync --group dev
# Plain OpenAI SDK call:
uv run python call_server.py
The client intentionally omits model; the app chooses the backing deployment
from FOUNDRY_MODEL. The script then sends two more turns, each continuing from
the previous turn's response.id as previous_response_id. The third turn asks
about the first turn's city, so it only succeeds if the harness agent behind the
route still remembers that far back in the chain.
This sample is local-only — no Dockerfile, no Foundry packaging.