What breaks when you put Jev in an agent workflow?
AI Architect
Key takeaways
- Pydantic AI documents that Jev may propose a tool again after it has executed, until the next turn, so the loop must track what already ran.
- A Jev workflow always carries a language model beside it, because a route with a string field escalates rather than being answered.
- Escalation is gated by decision_route_threshold and decision_boolean_threshold, both numbers you choose with no principled default published.
- TypeSafe documents 32k tokens for state plus the longest question and 64k per request, which an accumulating agent transcript outgrows.
- TypeSafe's jaggedness page states that structural invariants across separate questions are not guaranteed by the model.
Jev shipped with two integrations that put it inside an agent loop rather than beside one: a model in Pydantic AI and a classifier plus middleware in LangChain.
Reading both against TypeSafe’s own jaggedness page, the same pattern appears three times. Work that the language model used to absorb quietly now has to live in your loop.
Where does Jev sit in an agent workflow?
At the decision points, with a language model still doing the open-ended work.
LangChain’s harness write-up shows two placements. A routing step defines several
model choices with criteria and instructs Jev to choose the least costly model
that can complete the task, selecting on the latest user message. An
AutoModeMiddleware uses Jev to check tool calls for risky decisions and block
calls before the tool executes.
Pydantic AI takes the model-shaped route. TypeSafeModel('jev-latest') is passed
to an Agent with a Pydantic output type, and each field’s docstring becomes the
question Jev answers. Confidence per field arrives in
provider_details['confidence'].
One turn of an agent loop with a Jev gate
Show as text
| # | From | To | Message |
|---|---|---|---|
| 1 | Agent loop | Agent loop | Project state for this decision. Filter in code. Irrelevant context costs accuracy. |
| 2 | Agent loop | Jev | POST /v1/systemone. State plus every question you need about it. |
| 3 | Jev | Agent loop | typed answers with probabilities |
| 4 | Agent loop | Agent loop | Compare confidence to your threshold. Below it, hand the decision to a language model. |
| 5 | Agent loop | Agent loop | Record the action in the executed ledger. Jev may propose the same tool again next call. |
Both placements share a property worth naming. The decision is made from a state you assembled, about an option set you fixed before the call, and the model returns a probability over that set. Nothing about the loop’s history reaches it unless you put it there.
Why does a Jev workflow always need a language model?
Because the output space has no strings in it.
Pydantic AI states the consequence directly: Jev does not write text, so a route
with a str field escalates to a language model. TypeSafe’s jaggedness page is
blunter, saying the model is not trained to generate text and that forcing it
produces poor and slow results.
In an agent loop this bites at the first free-form tool argument. A gate deciding
whether to run bash is a bounded decision Jev answers well. Deciding what
command to run is a generation problem, and no amount of question modelling turns
it into a choice among 255 named options unless the command set genuinely is that
small.
The practical shape is two models with a clear division. The language model plans, writes arguments and explains. Jev answers the bounded questions along the way: is this risky, which route, how severe. Designing as though one model can cover both ends with an escalation path you did not plan for.
What threshold should gate a Jev decision in a workflow?
One you fitted yourself, because nobody has published a defensible default.
Pydantic AI exposes decision_route_threshold for handing uncertain decisions to
a fallback language model, and decision_boolean_threshold for boolean fields.
LangChain’s write-up describes the tool-call gate and the router without naming
any threshold at all.
The reason a default cannot be borrowed is what the number measures. As covered in how Jev’s question types report confidence, confidence is derived from the probability distribution and describes how concentrated the mass was. A confident answer is one the model was not torn about, which leaves correctness open.
So the threshold is a spread cutoff being used as a correctness proxy, and the mapping between them is a property of your data. Bucket your own decisions by reported confidence, measure how often each bucket was right, and set the cutoff from that. Skipping it means every escalation boundary in the workflow sits at an arbitrary point.
Why does Jev re-propose a tool the workflow already ran?
Because the model answers a question about the state you sent and holds no concept of turn progress.
Pydantic AI lists this among its documented failure modes: once a tool executes, Jev may propose it again until the next turn. That is unsurprising given the architecture. The question “which action should run next” is answered from the state, and if the state does not record that the action already ran, the same answer stays correct by the model’s reading.
What a Jev workflow owns against what Jev owns
Show as text
| # | Layer | Note |
|---|---|---|
| 1 | Response matches the schema you asked for | Guaranteed by construction. |
| 2 | Every answer carries a probability | Choice and score add a confidence. |
| 3 | All questions answered in one parallel pass | Adding questions barely moves latency. |
| · | Jev above, your code below (breakpoint) | |
| 4 | Which actions have already executed | The model has no notion of turn progress. |
| 5 | Which fields of state each question needs | Accuracy falls with irrelevant context. |
| 6 | The threshold that escalates to a model | No defensible default is published. |
| 7 | Arithmetic, counting and date ordering | All three documented as unreliable. |
A language model in the same position usually avoids this because the tool result is in its context and it reads the transcript as a sequence. Jev reads a payload. The fix is a ledger in your loop: record what executed, then remove those options from the next question’s option set rather than hoping the probability shifts.
This generalises. TypeSafe’s jaggedness page also says structural invariants across separate questions are not guaranteed, so two questions that logically exclude each other can both come back affirmative.
How much agent state can a Jev workflow send?
Less than a transcript, and less than the limit suggests.
The hard numbers are 64k tokens per request and 32k for state plus the longest question, per the Pydantic integration and TypeSafe’s models page. A loop that appends tool results crosses 32k quickly on anything involving file contents or search output.
The softer limit binds earlier. The jaggedness page says accuracy falls as the state grows with content unrelated to the decision, and recommends retrieving and filtering in code so you send only the fields the question needs. An agent transcript is the opposite shape, since it grows monotonically and most of it is irrelevant to any single decision.
That makes state projection a required step rather than an optimisation. For each question, assemble the minimum payload that could answer it. LangChain’s router selecting on the latest user message is exactly this move, and it keeps the call small as well as accurate.
How would you structure a Jev workflow today?
My read is that Jev belongs at gates and routers, with the loop keeping its own bookkeeping, and that the cost of adoption is mostly in that bookkeeping rather than in the integration.
Three things move into your code the moment you swap a language model decision for a Jev one: the executed-action ledger, the per-question state projection, and the escalation threshold. All three were previously handled by the language model reading a transcript, imperfectly and for free.
Against that, the gains are real and narrow. A gate that costs $0.042 per million input tokens with free output and answers in hundreds of milliseconds changes what you can afford to check. Checking every tool call becomes cheap enough to be the default rather than a sampled audit.
What I would avoid is treating the integration’s defaults as a design. The threshold, the projection and the ledger are the design, and the API call is the easy part.
Do this
Wire Jev into an agent loop without losing the loop's invariants
The order matters. Each step removes a class of failure that the step after it would otherwise hide, and the first two are the ones teams skip because the API call works without them.
Keep a ledger of actions already executed
Record every tool call the loop has run this turn and remove those options from the next question's option set. Pydantic AI documents that Jev may propose a tool again once it has executed, and a loop that trusts the answer will call the same tool until something else stops it.
Project the state per question rather than sending the transcript
Retrieve and filter in code, then send only the fields the question needs. TypeSafe's jaggedness page states that accuracy falls as the state grows with content unrelated to the decision, so an append-only agent history is the worst possible payload shape.
Keep every question independent of every other
Do not rely on answers cohering across questions. The jaggedness page says structural invariants between separate questions are not guaranteed, so two questions that logically exclude each other can both come back affirmative. Enforce the exclusion in code or model it as one choice question.
Fit the escalation threshold on labelled data before shipping it
Bucket your own decisions by reported confidence and measure how often each bucket was right. The integrations expose decision_route_threshold and decision_boolean_threshold, and the default value is a convention rather than a measurement of your distribution.
Leave the language model in the loop for anything textual
Any free-form tool argument, plan or explanation has to come from a model that generates text. Pydantic AI escalates a route with a string field for this reason, and designing as though Jev can be the only model in the loop fails at the first free-form field.
Compute the numbers rather than asking for them
Arithmetic, counting and date ordering are documented as unreliable. Work them out in your own code and pass the result into the state as a named value or a bucket, which is exact, cheaper and removes a whole class of silent wrong answer.
Frequently asked questions
- How do you use Jev inside an agent loop?
- Through one of the two shipped integrations. Pydantic AI exposes Jev as a model via TypeSafeModel with a Pydantic output type whose field docstrings become the questions. LangChain wraps it as TypeSafeClassifier and uses it inside middleware, including a routing step and a tool-call gate that blocks risky calls before execution.
- Why does Jev repeat a tool call that already ran?
- Pydantic AI's documentation lists this as a known failure mode: once a tool executes, Jev may propose it again until the next turn. The model answers a question about the state you handed it and holds no notion of turn progress, so the loop has to record executed actions and filter them out of the option set.
- What confidence threshold should escalate a Jev decision to a language model?
- No source publishes a defensible default, and the Pydantic integration's 0.5 is a convention. Confidence describes how concentrated the probability distribution was rather than how likely the answer is to be correct, so the threshold has to be fitted on your own labelled data before it carries routing decisions.
- Can Jev replace the planning model in an agent?
- No. Jev is not trained to generate text, so anything requiring a sentence, a plan or a tool argument that is free-form has to come from a language model. The pattern in both integrations is a language model for open-ended work with Jev answering the bounded decisions taken along the way.
- Does an agent transcript fit inside a Jev request?
- Often not for long. TypeSafe documents 64k tokens per request and 32k for state plus the longest question, and a loop that appends tool results grows past that. Worse, accuracy is documented to fall as the state fills with content the question does not need, so sending the whole transcript is the wrong move even when it fits.