---
title: What breaks when you put Jev in an agent workflow?
url: https://deepthinkingai.org/jev-in-workflows/
published: 2026-09-28
author: Shekhar Singh
topic: Agents & Protocols
tags: agents, tool-use, classification, production, evaluation
site: DeepThinking AI
---

# What breaks when you put Jev in an agent workflow?

**Summary:** Two shipped integrations put Jev inside an agent loop, from Pydantic AI and LangChain. Both leave three things to your code: a ledger of actions already executed, because Jev may re-propose a tool that ran; a per-question projection of state, because accuracy falls as irrelevant context grows; and the confidence threshold that escalates to a language model.

## Key takeaways
- Pydantic AI documents that Jev may propose a tool again after it has executed, until the next turn, so the loop must track what already ran.
- A Jev workflow always carries a language model beside it, because a route with a string field escalates rather than being answered.
- Escalation is gated by decision_route_threshold and decision_boolean_threshold, both numbers you choose with no principled default published.
- TypeSafe documents 32k tokens for state plus the longest question and 64k per request, which an accumulating agent transcript outgrows.
- TypeSafe's jaggedness page states that structural invariants across separate questions are not guaranteed by the model.

[Jev](/how-jev-works/) shipped with two integrations that put it inside an agent
loop rather than beside one: a
[model in Pydantic AI](https://pydantic.dev/docs/ai/models/typesafe/) and a
classifier plus middleware in
[LangChain](https://www.langchain.com/blog/building-a-harness-with-jev).

Reading both against TypeSafe's own
[jaggedness page](https://docs.typesafe.ai/model-jaggedness/jev-1.13), the same
pattern appears three times. Work that the language model used to absorb quietly
now has to live in your loop.

## Where does Jev sit in an agent workflow?

At the decision points, with a language model still doing the open-ended work.

LangChain's harness write-up shows two placements. A routing step defines several
model choices with criteria and instructs Jev to choose the least costly model
that can complete the task, selecting on the latest user message. An
`AutoModeMiddleware` uses Jev to check tool calls for risky decisions and block
calls before the tool executes.

Pydantic AI takes the model-shaped route. `TypeSafeModel('jev-latest')` is passed
to an `Agent` with a Pydantic output type, and each field's docstring becomes the
question Jev answers. Confidence per field arrives in
`provider_details['confidence']`.

**One turn of an agent loop with a Jev gate**

```mermaid
sequenceDiagram
    participant Agentloop as Agent loop
    participant Jev as Jev
    Agentloop->>Agentloop: Project state for this decision
    Agentloop->>Jev: POST /v1/systemone
    Jev-->>Agentloop: typed answers with probabilities
    Agentloop->>Agentloop: Compare confidence to your threshold
    Agentloop->>Agentloop: Record the action in the executed ledger
```

- Project state for this decision: Filter in code. Irrelevant context costs accuracy.
- POST /v1/systemone: State plus every question you need about it.
- Compare confidence to your threshold: Below it, hand the decision to a language model.
- Record the action in the executed ledger: Jev may propose the same tool again next call.

Three of the five steps in a turn happen in your code rather than at the API. The projection, the threshold comparison and the ledger are all work the loop owns, and none of them existed when a language model was making the same decision.

Both placements share a property worth naming. The decision is made from a state
you assembled, about an option set you fixed before the call, and the model
returns a probability over that set. Nothing about the loop's history reaches it
unless you put it there.

## Why does a Jev workflow always need a language model?

Because the output space has no strings in it.

Pydantic AI states the consequence directly: Jev does not write text, so a route
with a `str` field escalates to a language model. TypeSafe's jaggedness page is
blunter, saying the model is not trained to generate text and that forcing it
produces poor and slow results.

In an agent loop this bites at the first free-form tool argument. A gate deciding
whether to run `bash` is a bounded decision Jev answers well. Deciding what
command to run is a generation problem, and no amount of question modelling turns
it into a choice among 255 named options unless the command set genuinely is that
small.

The practical shape is two models with a clear division. The language model plans,
writes arguments and explains. Jev answers the bounded questions along the way:
is this risky, which route, how severe. Designing as though one model can cover
both ends with an escalation path you did not plan for.

## What threshold should gate a Jev decision in a workflow?

One you fitted yourself, because nobody has published a defensible default.

Pydantic AI exposes `decision_route_threshold` for handing uncertain decisions to
a fallback language model, and `decision_boolean_threshold` for boolean fields.
LangChain's write-up describes the tool-call gate and the router without naming
any threshold at all.

The reason a default cannot be borrowed is what the number measures. As covered
in [how Jev's question types report confidence](/jev-classification-questions/),
confidence is derived from the probability distribution and describes how
concentrated the mass was. A confident answer is one the model was not torn
about, which leaves correctness open.

So the threshold is a spread cutoff being used as a correctness proxy, and the
mapping between them is a property of your data. Bucket your own decisions by
reported confidence, measure how often each bucket was right, and set the cutoff
from that. Skipping it means every escalation boundary in the workflow sits at an
arbitrary point.

## Why does Jev re-propose a tool the workflow already ran?

Because the model answers a question about the state you sent and holds no
concept of turn progress.

Pydantic AI lists this among its documented failure modes: once a tool executes,
Jev may propose it again until the next turn. That is unsurprising given the
architecture. The question "which action should run next" is answered from the
state, and if the state does not record that the action already ran, the same
answer stays correct by the model's reading.

**What a Jev workflow owns against what Jev owns**

1. Response matches the schema you asked for
   Guaranteed by construction.
2. Every answer carries a probability
   Choice and score add a confidence.
3. All questions answered in one parallel pass
   Adding questions barely moves latency.
--- Jev above, your code below ---
4. Which actions have already executed
   The model has no notion of turn progress.
5. Which fields of state each question needs
   Accuracy falls with irrelevant context.
6. The threshold that escalates to a model
   No defensible default is published.
7. Arithmetic, counting and date ordering
   All three documented as unreliable.

The line moves further down the stack than a first reading of the API suggests. Everything below it was previously handled, badly or well, by the language model that used to make the decision, and swapping in a decision model transfers it to you.

A language model in the same position usually avoids this because the tool result
is in its context and it reads the transcript as a sequence. Jev reads a payload.
The fix is a ledger in your loop: record what executed, then remove those options
from the next question's option set rather than hoping the probability shifts.

This generalises. TypeSafe's jaggedness page also says structural invariants
across separate questions are not guaranteed, so two questions that logically
exclude each other can both come back affirmative.

## How much agent state can a Jev workflow send?

Less than a transcript, and less than the limit suggests.

The hard numbers are 64k tokens per request and 32k for state plus the longest
question, per the Pydantic integration and TypeSafe's
[models page](https://docs.typesafe.ai/models). A loop that appends tool results
crosses 32k quickly on anything involving file contents or search output.

The softer limit binds earlier. The jaggedness page says accuracy falls as the
state grows with content unrelated to the decision, and recommends retrieving and
filtering in code so you send only the fields the question needs. An agent
transcript is the opposite shape, since it grows monotonically and most of it is
irrelevant to any single decision.

That makes state projection a required step rather than an optimisation. For each
question, assemble the minimum payload that could answer it. LangChain's router
selecting on the latest user message is exactly this move, and it keeps the call
small as well as accurate.

## How would you structure a Jev workflow today?

My read is that Jev belongs at gates and routers, with the loop keeping its own
bookkeeping, and that the cost of adoption is mostly in that bookkeeping rather
than in the integration.

Three things move into your code the moment you swap a language model decision
for a Jev one: the executed-action ledger, the per-question state projection, and
the escalation threshold. All three were previously handled by the language model
reading a transcript, imperfectly and for free.

Against that, the gains are real and narrow. A gate that costs
[$0.042 per million input tokens](/jev-use-cases/) with free output and answers in
hundreds of milliseconds changes what you can afford to check. Checking every
tool call becomes cheap enough to be the default rather than a sampled audit.

What I would avoid is treating the integration's defaults as a design. The
threshold, the projection and the ledger are the design, and the API call is the
easy part.

<ReadNext
  href="/mcp-tool-call-retry-safety/"
  kicker="Related"
  title="When is it safe for an agent to retry a tool call?"
  note="The other half of the executed-action problem, and what a loop has to record before a repeat is safe."
/>

## Wire Jev into an agent loop without losing the loop's invariants

The order matters. Each step removes a class of failure that the step after it would otherwise hide, and the first two are the ones teams skip because the API call works without them.

1. **Keep a ledger of actions already executed**: Record every tool call the loop has run this turn and remove those options from the next question's option set. Pydantic AI documents that Jev may propose a tool again once it has executed, and a loop that trusts the answer will call the same tool until something else stops it.
2. **Project the state per question rather than sending the transcript**: Retrieve and filter in code, then send only the fields the question needs. TypeSafe's jaggedness page states that accuracy falls as the state grows with content unrelated to the decision, so an append-only agent history is the worst possible payload shape.
3. **Keep every question independent of every other**: Do not rely on answers cohering across questions. The jaggedness page says structural invariants between separate questions are not guaranteed, so two questions that logically exclude each other can both come back affirmative. Enforce the exclusion in code or model it as one choice question.
4. **Fit the escalation threshold on labelled data before shipping it**: Bucket your own decisions by reported confidence and measure how often each bucket was right. The integrations expose decision_route_threshold and decision_boolean_threshold, and the default value is a convention rather than a measurement of your distribution.
5. **Leave the language model in the loop for anything textual**: Any free-form tool argument, plan or explanation has to come from a model that generates text. Pydantic AI escalates a route with a string field for this reason, and designing as though Jev can be the only model in the loop fails at the first free-form field.
6. **Compute the numbers rather than asking for them**: Arithmetic, counting and date ordering are documented as unreliable. Work them out in your own code and pass the result into the state as a named value or a bucket, which is exact, cheaper and removes a whole class of silent wrong answer.


## Frequently asked questions

### How do you use Jev inside an agent loop?

Through one of the two shipped integrations. Pydantic AI exposes Jev as a model via TypeSafeModel with a Pydantic output type whose field docstrings become the questions. LangChain wraps it as TypeSafeClassifier and uses it inside middleware, including a routing step and a tool-call gate that blocks risky calls before execution.

### Why does Jev repeat a tool call that already ran?

Pydantic AI's documentation lists this as a known failure mode: once a tool executes, Jev may propose it again until the next turn. The model answers a question about the state you handed it and holds no notion of turn progress, so the loop has to record executed actions and filter them out of the option set.

### What confidence threshold should escalate a Jev decision to a language model?

No source publishes a defensible default, and the Pydantic integration's 0.5 is a convention. Confidence describes how concentrated the probability distribution was rather than how likely the answer is to be correct, so the threshold has to be fitted on your own labelled data before it carries routing decisions.

### Can Jev replace the planning model in an agent?

No. Jev is not trained to generate text, so anything requiring a sentence, a plan or a tool argument that is free-form has to come from a language model. The pattern in both integrations is a language model for open-ended work with Jev answering the bounded decisions taken along the way.

### Does an agent transcript fit inside a Jev request?

Often not for long. TypeSafe documents 64k tokens per request and 32k for state plus the longest question, and a loop that appends tool results grows past that. Worse, accuracy is documented to fall as the state fills with content the question does not need, so sending the whole transcript is the wrong move even when it fits.


## Sources
- [TypeSafe (Jev) model integration](https://pydantic.dev/docs/ai/models/typesafe/). Pydantic
- [Building a harness with Jev](https://www.langchain.com/blog/building-a-harness-with-jev). LangChain
- [Jev 1.13 jaggedness](https://docs.typesafe.ai/model-jaggedness/jev-1.13). TypeSafe AI
- [Models](https://docs.typesafe.ai/models). TypeSafe AI
- [API reference](https://docs.typesafe.ai/api). TypeSafe AI

---
Canonical HTML: https://deepthinkingai.org/jev-in-workflows/