Which agent runtime checks tool permissions, OpenAI or Claude?
AI Architect
Key takeaways
- OpenAI's Agents guide says the Agents API runs the Codex harness and manages orchestration plus durable session state, while function tools still run through your application.
- Anthropic's permission policy docs say the auto policy evaluates each agent or MCP tool call and can allow it, deny it, or pause for approval.
- Anthropic's permission policies do not cover custom tools, which still execute in your application after an `agent.custom_tool_use` event.
- OpenAI's architecture guide says a self-hosted environment remains your lifecycle problem, including provisioning, reconnection, shutdown, and file preservation.
- Anthropic's defaults differ by toolset: the agent toolset defaults to `always_allow`, while MCP toolsets default to `always_ask`.
OpenAI and Anthropic both sell managed agent runtimes, but the published trust boundary is different. In the pages cited here, OpenAI says it runs the harness and session state, while Anthropic says its server can evaluate each server-run tool call, so a team buying “managed agents” is choosing between two different approval models rather than one generic category.
What does OpenAI say about tool permissions in its agent runtime?
OpenAI’s Agents guide defines the hosted offer in control-plane terms. OpenAI runs a managed Codex harness, manages orchestration, handles automatic context compaction, and saves session configuration, turns, and items between tasks. The architecture page sharpens the line: the harness is the OpenAI-hosted Codex instance that runs the model and tool loop, while your application server submits tasks, receives events, and handles function tools. That matters because a function tool call still crosses back into code you operate before any side effect happens. The cited OpenAI pages describe where the runtime runs, how sessions persist, and who owns an optional self-hosted environment. They do not describe a vendor policy that evaluates each function tool call as allow, ask, or deny before your application sees it. If approvals are the requirement, that omission is the decision point.
OpenAI's documented permission boundary
Show as text
| # | Layer | Note |
|---|---|---|
| 1 | Hosted Codex harness | Runs model and tool loop. |
| 2 | Session orchestration | Compaction, recovery, durable state. |
| 3 | Function tool request | Sent to your app for execution. |
| · | OpenAI boundary (breakpoint) | |
| 4 | Tool approval logic | Your code decides whether it runs. |
| 5 | Self-hosted environment | You own lifecycle if you provide it. |
What does Claude say about tool permissions in its managed runtime?
Anthropic’s
permission policy page
is explicit about the per-call boundary. Permission policies govern
server-executed tools, meaning the pre-built agent toolset and the MCP toolset,
and the auto policy lets “the server evaluate each call and run it, deny it,
or pause for your approval.” The same page says the evaluation considers the
tool, the call input, and the session content up to that point, so two calls to
the same tool can resolve differently. Anthropic’s
10 September 2026 release note
confirms that the evaluation lands on agent.tool_use and agent.mcp_tool_use
events, with evaluation plus evaluated_permission fields that record the
outcome. This is a stronger documented claim than hosted orchestration alone.
It says the managed service can participate directly in the approval decision
for the tool paths it owns, which is the main distinction buyers should care
about when they compare agent runtimes.
Claude's documented permission boundary
Show as text
| # | Layer | Note |
|---|---|---|
| 1 | Server-run agent tools | Prebuilt toolset can auto-run. |
| 2 | Server-run MCP tools | MCP toolset can auto-run or ask. |
| 3 | Per-call evaluation | allow, ask, or deny each call. |
| · | Claude boundary (breakpoint) | |
| 4 | Custom tool execution | Your app decides and returns results. |
| 5 | Human approval path | always_ask or indeterminate auto calls. |
Which tool paths stay outside vendor control on both runtimes?
Neither vendor removes the need for application-side control everywhere.
OpenAI’s architecture guide says function tools return to your application
server, and Anthropic says custom tools are excluded from permission policies
and show up as agent.custom_tool_use for your code to handle. That leaves a
common engineering pattern across both products: hosted orchestration or hosted
evaluation may cover some tool paths, but the highest-consequence paths often
still terminate inside code you deploy. Teams that skip this distinction end up
buying a managed runtime and assuming the approval problem disappeared with it.
It did not. The right mental model is the same one this site used in
the read-versus-write control piece and in
the MCP mechanics explainer: the trust
boundary sits where a call turns into a real side effect. A vendor can host a
lot of orchestration above that point and still leave the decisive action in
your application.
Which OpenAI or Claude default creates the first production surprise?
The first surprise is that “managed” says almost nothing about default
approval posture. Anthropic documents two defaults by toolset: the agent
toolset defaults to always_allow, while MCP toolsets default to
always_ask. That split is rational because a new MCP server can expose tools
your application has never seen before, yet it still means a team can ship
server-run agent tools with no confirmation path unless it changes the default.
OpenAI’s cited pages create a different surprise. Because function tools are
handled by your application, the real default is whatever your handler does the
first time a tool request arrives. A missing review path is therefore not a
vendor mistake or a vendor feature. It is just your application’s behavior.
That is why the earlier OpenAI runtime boundary piece
and the earlier Claude auto policy analysis
are best read together. They describe two different places where the first
meaningful approval decision can happen.
Which agent runtime should you choose when approvals matter?
My view is simple. Choose Claude Managed Agents when a server-side call gate is
part of the product requirement and the relevant tools fit inside Anthropic’s
managed toolsets. Anthropic documents the evaluation path narrowly enough that
an engineer can reason about allow, ask, deny, event fields, and the
cases where the client cannot override the server. Choose OpenAI’s Agents API
when the pain you want removed is operating the harness, session continuity,
compaction, and recovery while your own application keeps the approval logic
for function tools and optional self-hosted compute. That is still a valuable
managed boundary. It just answers a different problem. Treating these runtimes
as if they made the same promise produces design errors early and audit trouble
later. If approvals drive the buying decision, compare the permission boundary
first and every other runtime feature second.
Do this
Pick the right hosted agent permission boundary
Start with the approval owner, because every later runtime decision is just that first choice expressed through tools, sandboxes, and audit logs.
List the tools that can spend money or change production state
Separate read-only tools from side-effecting ones before comparing vendors. The only permission boundary that matters is the one around calls that can hurt you.
Mark which of those tools the vendor runs and which your app runs
OpenAI function tools and Anthropic custom tools both cross back into your application. Vendor-side evaluation cannot govern a path your own code still executes.
Decide whether you need a documented server-side call gate
If the requirement is that the hosted runtime itself can allow, deny, or pause a tool call, Anthropic documents that for the managed toolsets. OpenAI's cited pages document orchestration ownership instead.
Check the defaults before your first session goes live
Anthropic defaults the agent toolset to `always_allow` and MCP toolsets to `always_ask`. OpenAI requires your tool handler to exist at all, which means your own default behavior is the real control point there.
Keep logs at the event that made the decision
On Claude, keep `evaluated_permission` and `evaluation.reason_code` when they exist. On OpenAI, log the moment your application accepted or rejected the function tool request, because that is where the approval decision occurred.
Frequently asked questions
- Does OpenAI's Agents API approve tool calls for me?
- Not in the OpenAI pages cited here. OpenAI documents a hosted harness, orchestration, and session state, while function tools still go back to your application to run and return results. That means your application still owns the approval decision for those calls.
- Does Claude's auto permission policy cover custom tools?
- No. Anthropic's permission policy page says permission policies apply to server-executed tools, meaning the agent toolset and MCP toolset. Custom tools emit `agent.custom_tool_use`, and your application decides whether to execute them.
- Can a client override a Claude tool call that auto denied?
- No. Anthropic says a call denied under `auto` does not run and your client cannot override the denial. The session continues with an error tool result instead.
- Are MCP tools treated the same way on both runtimes?
- No. OpenAI's architecture page says the harness can call remote MCP tools directly, but the cited OpenAI pages do not describe a server-side permission policy for each call. Anthropic documents a separate MCP toolset with its own default policy and the same `auto` evaluation path.
- Which runtime is the better fit for production approvals?
- Claude Managed Agents is the cleaner fit when you want the vendor to document a per-call server evaluation path for server-run tools. OpenAI's Agents API is the cleaner fit when your main goal is hosted orchestration and durable sessions while approvals stay in your own application.