DeepThinking AI

Which agent runtime checks tool permissions, OpenAI or Claude?

AI Architect

Key takeaways

  • OpenAI's Agents guide says the Agents API runs the Codex harness and manages orchestration plus durable session state, while function tools still run through your application.
  • Anthropic's permission policy docs say the auto policy evaluates each agent or MCP tool call and can allow it, deny it, or pause for approval.
  • Anthropic's permission policies do not cover custom tools, which still execute in your application after an `agent.custom_tool_use` event.
  • OpenAI's architecture guide says a self-hosted environment remains your lifecycle problem, including provisioning, reconnection, shutdown, and file preservation.
  • Anthropic's defaults differ by toolset: the agent toolset defaults to `always_allow`, while MCP toolsets default to `always_ask`.

OpenAI and Anthropic both sell managed agent runtimes, but the published trust boundary is different. In the pages cited here, OpenAI says it runs the harness and session state, while Anthropic says its server can evaluate each server-run tool call, so a team buying “managed agents” is choosing between two different approval models rather than one generic category.

What does OpenAI say about tool permissions in its agent runtime?

OpenAI’s Agents guide defines the hosted offer in control-plane terms. OpenAI runs a managed Codex harness, manages orchestration, handles automatic context compaction, and saves session configuration, turns, and items between tasks. The architecture page sharpens the line: the harness is the OpenAI-hosted Codex instance that runs the model and tool loop, while your application server submits tasks, receives events, and handles function tools. That matters because a function tool call still crosses back into code you operate before any side effect happens. The cited OpenAI pages describe where the runtime runs, how sessions persist, and who owns an optional self-hosted environment. They do not describe a vendor policy that evaluates each function tool call as allow, ask, or deny before your application sees it. If approvals are the requirement, that omission is the decision point.

OpenAI's documented permission boundary

OpenAI's documented permission boundaryDiagram: 6 ordered layers. Hosted Codex harness, then Session orchestration, then Function tool request, then OpenAI boundary (breakpoint), then Tool approval logic, then Self-hosted environment.1Hosted Codex harnessRuns model and tool loop.2Session orchestrationCompaction, recovery, durable state.3Function tool requestSent to your app for execution.OpenAI boundary4Tool approval logicYour code decides whether it runs.5Self-hosted environmentYou own lifecycle if you provide it.
Show as text
OpenAI's documented permission boundary. Diagram: 6 ordered layers. Hosted Codex harness, then Session orchestration, then Function tool request, then OpenAI boundary (breakpoint), then Tool approval logic, then Self-hosted environment.
#LayerNote
1Hosted Codex harnessRuns model and tool loop.
2Session orchestrationCompaction, recovery, durable state.
3Function tool requestSent to your app for execution.
·OpenAI boundary (breakpoint)
4Tool approval logicYour code decides whether it runs.
5Self-hosted environmentYou own lifecycle if you provide it.
OpenAI's cited pages define a managed harness boundary, then hand function tools and optional self-hosted compute back to your application. The docs are explicit about the runtime split and silent on a server-side per-call permission policy for those tool paths.

What does Claude say about tool permissions in its managed runtime?

Anthropic’s permission policy page is explicit about the per-call boundary. Permission policies govern server-executed tools, meaning the pre-built agent toolset and the MCP toolset, and the auto policy lets “the server evaluate each call and run it, deny it, or pause for your approval.” The same page says the evaluation considers the tool, the call input, and the session content up to that point, so two calls to the same tool can resolve differently. Anthropic’s 10 September 2026 release note confirms that the evaluation lands on agent.tool_use and agent.mcp_tool_use events, with evaluation plus evaluated_permission fields that record the outcome. This is a stronger documented claim than hosted orchestration alone. It says the managed service can participate directly in the approval decision for the tool paths it owns, which is the main distinction buyers should care about when they compare agent runtimes.

Claude's documented permission boundary

Claude's documented permission boundaryDiagram: 6 ordered layers. Server-run agent tools, then Server-run MCP tools, then Per-call evaluation, then Claude boundary (breakpoint), then Custom tool execution, then Human approval path.1Server-run agent toolsPrebuilt toolset can auto-run.2Server-run MCP toolsMCP toolset can auto-run or ask.3Per-call evaluationallow, ask, or deny each call.Claude boundary4Custom tool executionYour app decides and returns results.5Human approval pathalways_ask or indeterminate auto calls.
Show as text
Claude's documented permission boundary. Diagram: 6 ordered layers. Server-run agent tools, then Server-run MCP tools, then Per-call evaluation, then Claude boundary (breakpoint), then Custom tool execution, then Human approval path.
#LayerNote
1Server-run agent toolsPrebuilt toolset can auto-run.
2Server-run MCP toolsMCP toolset can auto-run or ask.
3Per-call evaluationallow, ask, or deny each call.
·Claude boundary (breakpoint)
4Custom tool executionYour app decides and returns results.
5Human approval pathalways_ask or indeterminate auto calls.
Anthropic's permission policy page draws a tighter server-side boundary for the managed toolsets than OpenAI does in its cited runtime docs. The same page also states where that boundary ends: custom tools stay in your application.

Which tool paths stay outside vendor control on both runtimes?

Neither vendor removes the need for application-side control everywhere. OpenAI’s architecture guide says function tools return to your application server, and Anthropic says custom tools are excluded from permission policies and show up as agent.custom_tool_use for your code to handle. That leaves a common engineering pattern across both products: hosted orchestration or hosted evaluation may cover some tool paths, but the highest-consequence paths often still terminate inside code you deploy. Teams that skip this distinction end up buying a managed runtime and assuming the approval problem disappeared with it. It did not. The right mental model is the same one this site used in the read-versus-write control piece and in the MCP mechanics explainer: the trust boundary sits where a call turns into a real side effect. A vendor can host a lot of orchestration above that point and still leave the decisive action in your application.

Which OpenAI or Claude default creates the first production surprise?

The first surprise is that “managed” says almost nothing about default approval posture. Anthropic documents two defaults by toolset: the agent toolset defaults to always_allow, while MCP toolsets default to always_ask. That split is rational because a new MCP server can expose tools your application has never seen before, yet it still means a team can ship server-run agent tools with no confirmation path unless it changes the default. OpenAI’s cited pages create a different surprise. Because function tools are handled by your application, the real default is whatever your handler does the first time a tool request arrives. A missing review path is therefore not a vendor mistake or a vendor feature. It is just your application’s behavior. That is why the earlier OpenAI runtime boundary piece and the earlier Claude auto policy analysis are best read together. They describe two different places where the first meaningful approval decision can happen.

Which agent runtime should you choose when approvals matter?

My view is simple. Choose Claude Managed Agents when a server-side call gate is part of the product requirement and the relevant tools fit inside Anthropic’s managed toolsets. Anthropic documents the evaluation path narrowly enough that an engineer can reason about allow, ask, deny, event fields, and the cases where the client cannot override the server. Choose OpenAI’s Agents API when the pain you want removed is operating the harness, session continuity, compaction, and recovery while your own application keeps the approval logic for function tools and optional self-hosted compute. That is still a valuable managed boundary. It just answers a different problem. Treating these runtimes as if they made the same promise produces design errors early and audit trouble later. If approvals drive the buying decision, compare the permission boundary first and every other runtime feature second.

Do this

Pick the right hosted agent permission boundary

Start with the approval owner, because every later runtime decision is just that first choice expressed through tools, sandboxes, and audit logs.

  1. List the tools that can spend money or change production state

    Separate read-only tools from side-effecting ones before comparing vendors. The only permission boundary that matters is the one around calls that can hurt you.

  2. Mark which of those tools the vendor runs and which your app runs

    OpenAI function tools and Anthropic custom tools both cross back into your application. Vendor-side evaluation cannot govern a path your own code still executes.

  3. Decide whether you need a documented server-side call gate

    If the requirement is that the hosted runtime itself can allow, deny, or pause a tool call, Anthropic documents that for the managed toolsets. OpenAI's cited pages document orchestration ownership instead.

  4. Check the defaults before your first session goes live

    Anthropic defaults the agent toolset to `always_allow` and MCP toolsets to `always_ask`. OpenAI requires your tool handler to exist at all, which means your own default behavior is the real control point there.

  5. Keep logs at the event that made the decision

    On Claude, keep `evaluated_permission` and `evaluation.reason_code` when they exist. On OpenAI, log the moment your application accepted or rejected the function tool request, because that is where the approval decision occurred.

Frequently asked questions

Does OpenAI's Agents API approve tool calls for me?
Not in the OpenAI pages cited here. OpenAI documents a hosted harness, orchestration, and session state, while function tools still go back to your application to run and return results. That means your application still owns the approval decision for those calls.
Does Claude's auto permission policy cover custom tools?
No. Anthropic's permission policy page says permission policies apply to server-executed tools, meaning the agent toolset and MCP toolset. Custom tools emit `agent.custom_tool_use`, and your application decides whether to execute them.
Can a client override a Claude tool call that auto denied?
No. Anthropic says a call denied under `auto` does not run and your client cannot override the denial. The session continues with an error tool result instead.
Are MCP tools treated the same way on both runtimes?
No. OpenAI's architecture page says the harness can call remote MCP tools directly, but the cited OpenAI pages do not describe a server-side permission policy for each call. Anthropic documents a separate MCP toolset with its own default policy and the same `auto` evaluation path.
Which runtime is the better fit for production approvals?
Claude Managed Agents is the cleaner fit when you want the vendor to document a per-call server evaluation path for server-run tools. OpenAI's Agents API is the cleaner fit when your main goal is hosted orchestration and durable sessions while approvals stay in your own application.

Sources

  1. AgentsOpenAI · 2026-09-24
  2. ArchitectureOpenAI · 2026-09-24
  3. Permission policies, Claude Managed AgentsAnthropic · 2026-09-10
  4. Claude Platform release notesAnthropic · 2026-09-10

openaiclaudeagentspermissionsmcp