DeepThinking AI

What breaks when you put Jev in an agent workflow?

AI Architect

Key takeaways

  • Pydantic AI documents that Jev may propose a tool again after it has executed, until the next turn, so the loop must track what already ran.
  • A Jev workflow always carries a language model beside it, because a route with a string field escalates rather than being answered.
  • Escalation is gated by decision_route_threshold and decision_boolean_threshold, both numbers you choose with no principled default published.
  • TypeSafe documents 32k tokens for state plus the longest question and 64k per request, which an accumulating agent transcript outgrows.
  • TypeSafe's jaggedness page states that structural invariants across separate questions are not guaranteed by the model.

Jev shipped with two integrations that put it inside an agent loop rather than beside one: a model in Pydantic AI and a classifier plus middleware in LangChain.

Reading both against TypeSafe’s own jaggedness page, the same pattern appears three times. Work that the language model used to absorb quietly now has to live in your loop.

Where does Jev sit in an agent workflow?

At the decision points, with a language model still doing the open-ended work.

LangChain’s harness write-up shows two placements. A routing step defines several model choices with criteria and instructs Jev to choose the least costly model that can complete the task, selecting on the latest user message. An AutoModeMiddleware uses Jev to check tool calls for risky decisions and block calls before the tool executes.

Pydantic AI takes the model-shaped route. TypeSafeModel('jev-latest') is passed to an Agent with a Pydantic output type, and each field’s docstring becomes the question Jev answers. Confidence per field arrives in provider_details['confidence'].

One turn of an agent loop with a Jev gate

One turn of an agent loop with a Jev gateSequence diagram between Agent loop and Jev. 1. Agent loop to Agent loop: Project state for this decision. 2. Agent loop to Jev: POST /v1/systemone. 3. Jev to Agent loop: typed answers with probabilities. 4. Agent loop to Agent loop: Compare confidence to your threshold. 5. Agent loop to Agent loop: Record the action in the executed ledger.Agent loopJevProject state for this decisionFilter in code. Irrelevant context costs accuracy.POST /v1/systemoneState plus every question you need about it.typed answers with probabilitiesCompare confidence to your thresholdBelow it, hand the decision to a language model.Record the action in the executed ledgerJev may propose the same tool again next call.
Show as text
One turn of an agent loop with a Jev gate. Sequence diagram between Agent loop and Jev. 1. Agent loop to Agent loop: Project state for this decision. 2. Agent loop to Jev: POST /v1/systemone. 3. Jev to Agent loop: typed answers with probabilities. 4. Agent loop to Agent loop: Compare confidence to your threshold. 5. Agent loop to Agent loop: Record the action in the executed ledger.
#FromToMessage
1Agent loopAgent loopProject state for this decision. Filter in code. Irrelevant context costs accuracy.
2Agent loopJevPOST /v1/systemone. State plus every question you need about it.
3JevAgent looptyped answers with probabilities
4Agent loopAgent loopCompare confidence to your threshold. Below it, hand the decision to a language model.
5Agent loopAgent loopRecord the action in the executed ledger. Jev may propose the same tool again next call.
Three of the five steps in a turn happen in your code rather than at the API. The projection, the threshold comparison and the ledger are all work the loop owns, and none of them existed when a language model was making the same decision.

Both placements share a property worth naming. The decision is made from a state you assembled, about an option set you fixed before the call, and the model returns a probability over that set. Nothing about the loop’s history reaches it unless you put it there.

Why does a Jev workflow always need a language model?

Because the output space has no strings in it.

Pydantic AI states the consequence directly: Jev does not write text, so a route with a str field escalates to a language model. TypeSafe’s jaggedness page is blunter, saying the model is not trained to generate text and that forcing it produces poor and slow results.

In an agent loop this bites at the first free-form tool argument. A gate deciding whether to run bash is a bounded decision Jev answers well. Deciding what command to run is a generation problem, and no amount of question modelling turns it into a choice among 255 named options unless the command set genuinely is that small.

The practical shape is two models with a clear division. The language model plans, writes arguments and explains. Jev answers the bounded questions along the way: is this risky, which route, how severe. Designing as though one model can cover both ends with an escalation path you did not plan for.

What threshold should gate a Jev decision in a workflow?

One you fitted yourself, because nobody has published a defensible default.

Pydantic AI exposes decision_route_threshold for handing uncertain decisions to a fallback language model, and decision_boolean_threshold for boolean fields. LangChain’s write-up describes the tool-call gate and the router without naming any threshold at all.

The reason a default cannot be borrowed is what the number measures. As covered in how Jev’s question types report confidence, confidence is derived from the probability distribution and describes how concentrated the mass was. A confident answer is one the model was not torn about, which leaves correctness open.

So the threshold is a spread cutoff being used as a correctness proxy, and the mapping between them is a property of your data. Bucket your own decisions by reported confidence, measure how often each bucket was right, and set the cutoff from that. Skipping it means every escalation boundary in the workflow sits at an arbitrary point.

Why does Jev re-propose a tool the workflow already ran?

Because the model answers a question about the state you sent and holds no concept of turn progress.

Pydantic AI lists this among its documented failure modes: once a tool executes, Jev may propose it again until the next turn. That is unsurprising given the architecture. The question “which action should run next” is answered from the state, and if the state does not record that the action already ran, the same answer stays correct by the model’s reading.

What a Jev workflow owns against what Jev owns

What a Jev workflow owns against what Jev ownsDiagram: 8 ordered layers. Response matches the schema you asked for, then Every answer carries a probability, then All questions answered in one parallel pass, then Jev above, your code below (breakpoint), then Which actions have already executed, then Which fields of state each question needs, then The threshold that escalates to a model, then Arithmetic, counting and date ordering.1Response matches the schema you asked forGuaranteed by construction.2Every answer carries a probabilityChoice and score add a confidence.3All questions answered in one parallel passAdding questions barely moves latency.Jev above, your code below4Which actions have already executedThe model has no notion of turn progress.5Which fields of state each question needsAccuracy falls with irrelevant context.6The threshold that escalates to a modelNo defensible default is published.7Arithmetic, counting and date orderingAll three documented as unreliable.
Show as text
What a Jev workflow owns against what Jev owns. Diagram: 8 ordered layers. Response matches the schema you asked for, then Every answer carries a probability, then All questions answered in one parallel pass, then Jev above, your code below (breakpoint), then Which actions have already executed, then Which fields of state each question needs, then The threshold that escalates to a model, then Arithmetic, counting and date ordering.
#LayerNote
1Response matches the schema you asked forGuaranteed by construction.
2Every answer carries a probabilityChoice and score add a confidence.
3All questions answered in one parallel passAdding questions barely moves latency.
·Jev above, your code below (breakpoint)
4Which actions have already executedThe model has no notion of turn progress.
5Which fields of state each question needsAccuracy falls with irrelevant context.
6The threshold that escalates to a modelNo defensible default is published.
7Arithmetic, counting and date orderingAll three documented as unreliable.
The line moves further down the stack than a first reading of the API suggests. Everything below it was previously handled, badly or well, by the language model that used to make the decision, and swapping in a decision model transfers it to you.

A language model in the same position usually avoids this because the tool result is in its context and it reads the transcript as a sequence. Jev reads a payload. The fix is a ledger in your loop: record what executed, then remove those options from the next question’s option set rather than hoping the probability shifts.

This generalises. TypeSafe’s jaggedness page also says structural invariants across separate questions are not guaranteed, so two questions that logically exclude each other can both come back affirmative.

How much agent state can a Jev workflow send?

Less than a transcript, and less than the limit suggests.

The hard numbers are 64k tokens per request and 32k for state plus the longest question, per the Pydantic integration and TypeSafe’s models page. A loop that appends tool results crosses 32k quickly on anything involving file contents or search output.

The softer limit binds earlier. The jaggedness page says accuracy falls as the state grows with content unrelated to the decision, and recommends retrieving and filtering in code so you send only the fields the question needs. An agent transcript is the opposite shape, since it grows monotonically and most of it is irrelevant to any single decision.

That makes state projection a required step rather than an optimisation. For each question, assemble the minimum payload that could answer it. LangChain’s router selecting on the latest user message is exactly this move, and it keeps the call small as well as accurate.

How would you structure a Jev workflow today?

My read is that Jev belongs at gates and routers, with the loop keeping its own bookkeeping, and that the cost of adoption is mostly in that bookkeeping rather than in the integration.

Three things move into your code the moment you swap a language model decision for a Jev one: the executed-action ledger, the per-question state projection, and the escalation threshold. All three were previously handled by the language model reading a transcript, imperfectly and for free.

Against that, the gains are real and narrow. A gate that costs $0.042 per million input tokens with free output and answers in hundreds of milliseconds changes what you can afford to check. Checking every tool call becomes cheap enough to be the default rather than a sampled audit.

What I would avoid is treating the integration’s defaults as a design. The threshold, the projection and the ledger are the design, and the API call is the easy part.

Do this

Wire Jev into an agent loop without losing the loop's invariants

The order matters. Each step removes a class of failure that the step after it would otherwise hide, and the first two are the ones teams skip because the API call works without them.

  1. Keep a ledger of actions already executed

    Record every tool call the loop has run this turn and remove those options from the next question's option set. Pydantic AI documents that Jev may propose a tool again once it has executed, and a loop that trusts the answer will call the same tool until something else stops it.

  2. Project the state per question rather than sending the transcript

    Retrieve and filter in code, then send only the fields the question needs. TypeSafe's jaggedness page states that accuracy falls as the state grows with content unrelated to the decision, so an append-only agent history is the worst possible payload shape.

  3. Keep every question independent of every other

    Do not rely on answers cohering across questions. The jaggedness page says structural invariants between separate questions are not guaranteed, so two questions that logically exclude each other can both come back affirmative. Enforce the exclusion in code or model it as one choice question.

  4. Fit the escalation threshold on labelled data before shipping it

    Bucket your own decisions by reported confidence and measure how often each bucket was right. The integrations expose decision_route_threshold and decision_boolean_threshold, and the default value is a convention rather than a measurement of your distribution.

  5. Leave the language model in the loop for anything textual

    Any free-form tool argument, plan or explanation has to come from a model that generates text. Pydantic AI escalates a route with a string field for this reason, and designing as though Jev can be the only model in the loop fails at the first free-form field.

  6. Compute the numbers rather than asking for them

    Arithmetic, counting and date ordering are documented as unreliable. Work them out in your own code and pass the result into the state as a named value or a bucket, which is exact, cheaper and removes a whole class of silent wrong answer.

Frequently asked questions

How do you use Jev inside an agent loop?
Through one of the two shipped integrations. Pydantic AI exposes Jev as a model via TypeSafeModel with a Pydantic output type whose field docstrings become the questions. LangChain wraps it as TypeSafeClassifier and uses it inside middleware, including a routing step and a tool-call gate that blocks risky calls before execution.
Why does Jev repeat a tool call that already ran?
Pydantic AI's documentation lists this as a known failure mode: once a tool executes, Jev may propose it again until the next turn. The model answers a question about the state you handed it and holds no notion of turn progress, so the loop has to record executed actions and filter them out of the option set.
What confidence threshold should escalate a Jev decision to a language model?
No source publishes a defensible default, and the Pydantic integration's 0.5 is a convention. Confidence describes how concentrated the probability distribution was rather than how likely the answer is to be correct, so the threshold has to be fitted on your own labelled data before it carries routing decisions.
Can Jev replace the planning model in an agent?
No. Jev is not trained to generate text, so anything requiring a sentence, a plan or a tool argument that is free-form has to come from a language model. The pattern in both integrations is a language model for open-ended work with Jev answering the bounded decisions taken along the way.
Does an agent transcript fit inside a Jev request?
Often not for long. TypeSafe documents 64k tokens per request and 32k for state plus the longest question, and a loop that appends tool results grows past that. Worse, accuracy is documented to fall as the state fills with content the question does not need, so sending the whole transcript is the wrong move even when it fits.

Sources

  1. TypeSafe (Jev) model integrationPydantic
  2. Building a harness with JevLangChain
  3. Jev 1.13 jaggednessTypeSafe AI
  4. ModelsTypeSafe AI
  5. API referenceTypeSafe AI

agentstool-useclassificationproductionevaluation