DeepThinking AI

Why does vLLM 0.30.0 cut some models' max context length?

AI Architect

Key takeaways

  • PR #56446 fixed vLLM's YaRN scaling for vendor-alias rope_scaling types that were not the exact string "yarn".
  • TeleChat3-36B-Thinking's derived max_model_len drops from 131072 to 32768 tokens after upgrading to v0.30.0.
  • Sarvam-105b's derived max_model_len drops from 5242880 to 131072 tokens, a 40x reduction.
  • The old, inflated max_model_len let requests index past the end of the precomputed rotary embedding cache.
  • The fix aligns vLLM with Transformers, which treats max_position_embeddings as the final context length and never multiplies it by the YaRN factor again.

vLLM 0.30.0 shipped on September 22, 2026 with a one-line release note about YaRN alignment. Behind it is a bug fix that had let some models advertise a context length they could not safely serve: requests near the old figure risked indexing past the end of the cache vLLM precomputes for rotary embeddings.

What did vLLM 0.30.0 change about YaRN context length?

PR #56446 stopped vendor YaRN aliases from double-applying the context scaling factor. vLLM’s existing safeguard against re-deriving an already-scaled length only matched the exact string "yarn", so a model whose config used a vendor-specific type such as deepseek_yarn or telechat3-yarn skipped that check and had its already-scaled max_position_embeddings multiplied by factor a second time. The fix aligns vLLM with Hugging Face Transformers’ implementation, which treats max_position_embeddings as the final context length and never re-multiplies it, and adopts Transformers’ own _compute_yarn_parameters logic for the attention math, dropping the now-unused attn_factor and extrapolation_factor fields. The change ships as an ordinary point release rather than a major version bump, so a deployment that pins vLLM loosely, or rebuilds a container image on a schedule, can pick it up without anyone deciding to. Nothing in the request or response shape changes: the same rope_scaling config that worked yesterday still parses, still starts the server, and still serves requests. Only the derived max_model_len a client can actually rely on moves, and it moves downward.

How the vendor alias bug inflated max_model_len

How the vendor alias bug inflated max_model_lenDiagram: 5 ordered layers. Vendor max_position_embeddings, then YaRN factor applied once, then Vendor alias bug, pre-0.30.0 (breakpoint), then Same factor applied again, then Inflated max_model_len shipped.1Vendor max_position_embeddingse.g. 32768 for TeleChat3-36B2YaRN factor applied oncematches Transformers outputVendor alias bug, pre-0.30.03Same factor applied againdeepseek_yarn, telechat3-yarn alias4Inflated max_model_len shipped131072 advertised, only 32768 safe
Show as text
How the vendor alias bug inflated max_model_len. Diagram: 5 ordered layers. Vendor max_position_embeddings, then YaRN factor applied once, then Vendor alias bug, pre-0.30.0 (breakpoint), then Same factor applied again, then Inflated max_model_len shipped.
#LayerNote
1Vendor max_position_embeddingse.g. 32768 for TeleChat3-36B
2YaRN factor applied oncematches Transformers output
·Vendor alias bug, pre-0.30.0 (breakpoint)
3Same factor applied againdeepseek_yarn, telechat3-yarn alias
4Inflated max_model_len shipped131072 advertised, only 32768 safe
The vendor alias types were not recognized by vLLM's safeguard against re-deriving an already-scaled length, so the same YaRN factor was multiplied in twice before the fix in v0.30.0.

Which models are confirmed affected, and by how much?

Two, named directly in the release notes and the PR. Tele-AI/TeleChat3-36B-Thinking’s derived max_model_len drops from 131,072 to 32,768 tokens, exactly a 4x reduction. sarvamai/sarvam-105b’s drops from 5,242,880 to 131,072 tokens, exactly 40x. Both ratios match the YaRN scaling factor each model’s config sets, which is the signature of a value being multiplied in twice: once correctly, once by the vendor alias the safeguard missed. Any model whose rope_scaling.type uses a YaRN variant other than the plain string "yarn" is worth checking against its Transformers config directly, since the two named models are the ones vLLM’s own maintainers happened to catch rather than an exhaustive list. The exact ratios are the tell: a derived length that is a clean multiple of the model’s own scaling factor almost always means that factor got applied more than once somewhere in the chain, which is a faster check than reading through the rope scaling code for every model you serve.

Derived max_model_len before and after v0.30.0

Derived max_model_len before and after v0.30.0Bar chart. TeleChat3-36B, before: 131,072 tokens. TeleChat3-36B, after: 32,768 tokens. sarvam-105b, before: 5,242,880 tokens. sarvam-105b, after: 131,072 tokens.TeleChat3-36B, before131,072 tokensTeleChat3-36B, after32,768 tokenssarvam-105b, before5,242,880 tokenssarvam-105b, after131,072 tokens
Show data
Derived max_model_len before and after v0.30.0. Bar chart. TeleChat3-36B, before: 131,072 tokens. TeleChat3-36B, after: 32,768 tokens. sarvam-105b, before: 5,242,880 tokens. sarvam-105b, after: 131,072 tokens.
ItemValue (tokens)
TeleChat3-36B, before131,072
TeleChat3-36B, after32,768
sarvam-105b, before5,242,880
sarvam-105b, after131,072
TeleChat3-36B-Thinking drops by exactly 4x and sarvam-105b by exactly 40x, because both ratios are the YaRN factor that vLLM had been applying a second time before the fix.

Was the old, larger context length actually safe to use?

No, and that is the detail easy to miss reading the release note as a downgrade. The PR states plainly that the double-scaled figure let requests index past the end of the precomputed cosine and sine cache vLLM builds for rotary position embeddings. A request that filled anywhere close to the old advertised length was not exercising real extra capacity. It was exercising a code path the safeguard was supposed to have prevented in the first place. Framing v0.30.0 as a version that “cuts context length” undersells what actually happened: the true, safe length was 32,768 and 131,072 tokens respectively the whole time, and the old numbers were reporting a capability that did not exist.

What should you actually do before upgrading past vLLM 0.30.0?

Check your served model’s rope_scaling.type before you upgrade. Waiting until a request fails means finding out in production instead of in staging. If it is one of the vendor aliases the PR names, or any YaRN variant that is not the literal string "yarn", expect max_model_len to drop and plan your token budgets and chunk sizes against the corrected figure rather than the one you deployed with yesterday. Do not roll back to recover the larger number. The larger number was the bug. If your capacity planning genuinely needs more than the corrected length, treat that as a model or hardware decision rather than a version pin. This is also a case, like Claude Opus 5.5’s silent progress-update change on the same day, where a version bump changes real behavior through a release note most people will skim past rather than through an error anyone has to acknowledge.

Do this

Check whether the vLLM 0.30.0 YaRN fix changes your deployment

The fix is a correctness change rather than an optional tuning knob, so the order below is about finding out where it applies before it surprises a production request.

  1. Find your served model's rope_scaling type in its config.json

    The bug applied to vendor-specific YaRN aliases such as deepseek_yarn and telechat3-yarn. It did not touch the exact string yarn. Grep your model repository's config for rope_scaling and read the type field before you upgrade.

  2. Compare your current advertised max_model_len against the corrected value

    For TeleChat3-36B-Thinking the corrected figure is 32768, for sarvam-105b it is 131072. For any other affected model, expect the old figure to be the YaRN factor times the corrected one.

  3. Audit request-time token budgets and RAG chunk sizes built against the old number

    A prompt template or context-window budget sized against the inflated figure will start failing with a context-length-exceeded error the moment a request needs more than the corrected max, immediately after the upgrade.

  4. Treat the old, larger figure as unsafe capacity rather than a regression to roll back

    PR #56446 states requests near the old figure could index past the end of the precomputed rotary embedding cache. Rolling back to keep the larger number keeps that correctness bug too.

  5. Set max_model_len explicitly at server startup

    An explicit value survives the next change to vLLM's derivation logic. Trusting the auto-derived figure means a future alignment fix can move your deployment's capacity again without any code on your side changing.

Frequently asked questions

Why did my vLLM-served model's context length shrink after upgrading to 0.30.0?
If it uses a vendor-specific YaRN rope_scaling type such as deepseek_yarn or telechat3-yarn, vLLM was multiplying an already-scaled max_position_embeddings by the YaRN factor a second time. Version 0.30.0 fixed that in PR #56446, so the derived max_model_len is now smaller and correct.
Which models are confirmed affected by the vLLM YaRN scaling fix?
The PR names Tele-AI/TeleChat3-36B, whose derived length drops from 131072 to 32768 tokens, and sarvamai/sarvam-105b, whose derived length drops from 5242880 to 131072 tokens. Any model using a YaRN alias other than the exact string "yarn" is worth checking.
Was the old, larger context length actually usable?
No. The PR states the double-scaled figure let requests index past the end of the precomputed cosine and sine cache vLLM builds for rotary embeddings. That made the larger figure a correctness bug rather than extra usable capacity.
Does this fix change anything for models using plain YaRN?
Only the attention math. attn_factor and extrapolation_factor are now ignored in favor of mscale and attention_factor, adopting Transformers' own _compute_yarn_parameters logic. The length-doubling bug was specific to the vendor aliases; the base yarn type was already correct.
How do I avoid being surprised by a future scaling change like this?
Set max_model_len explicitly at server startup instead of trusting the value vLLM derives from the model config, and confirm the number you set against Transformers' own documented behavior for your model's rope_scaling type.

Sources

  1. vLLM v0.30.0 release notesvLLM Project · 2026-09-22
  2. config: align vendor YaRN rope_scaling aliases with Transformers (#56446)vLLM Project · 2026-09-22
  3. modeling_rope_utils.py, _compute_yarn_parametersHugging Face · 2026-09-22

vllmcontext-windowrope-scalingself-hostedinference