Finding the Real Context Ceiling: Needle-Benchmarking Forced RoPE Extrapolation
Published revised 5 min read
TL;DR
- I wanted more context on a self-hosted uncensored 35B MoE. Its config looks like it supports 256K, but the real text window is 32K.
- Forcing vLLM past that (
VLLM_ALLOW_LONG_MAX_MODEL_LEN=1) let it load at 64K and even 96K — no errors, no OOM. - But "it loaded" is not "it works." A progressive needle-in-haystack bench showed output stays coherent to ~60K and then falls off a cliff into
!!!!!garbage. - The longer runs failed retrieval without exhausting VRAM. I configured a 64K lane after passing the checks below and kept the benchmark for future model evaluations.
The setup
One of my daily-driver lanes is a quantized 35B Mixture-of-Experts model with a multi-token-prediction draft head, served on a single 24 GB AMD card via vLLM. It shipped at a 32K context window. The ask was simple: can we make it bigger?
The first surprise was in the model config. It advertised two different limits:
{
"max_position_embeddings": 262144, // top-level — looks like 256K!
"text_config": {
"max_position_embeddings": 32768 // the real text window
},
"vision_config": { ... } // this is a multimodal config
}
That top-level 262144 is the vision envelope, not the text RoPE. vLLM correctly derives the text model's max_model_len from the nested text_config value: 32768. Ask for more and it refuses:
User-specified max_model_len (65536) > derived max_model_len
(max_position_embeddings=32768). VLLM_ALLOW_LONG_MAX_MODEL_LEN must be
used with extreme caution...
Lesson one: when a model claims a giant context, check whether that number is the text window or a multimodal envelope. They are frequently not the same.
Forcing it open
The model's RoPE base frequency (rope_theta) was set to 10,000,000 — an aggressive value associated with long-context extrapolation. That made it plausible the model could run past its declared 32K even though it was never trained there. So I forced the door open with VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 and asked for 64K.
It loaded. Short prompts answered fine. vLLM reported healthy KV headroom (2.91× concurrency at 64K). Easy win, right?
This is the trap. A model that loads at a longer context will happily accept long prompts and emit confident-looking tokens — whether or not those tokens are correct. RoPE extrapolation past the trained window doesn't throw an exception; it silently degrades. You cannot see it in a health check, a short smoke test, or a metrics dashboard. You can only see it by asking the model to use the long context and checking the answer.
The bench: a needle in a growing haystack
So I built a progressive needle-in-haystack coherence test. The recipe is deliberately boring:
- Plant a unique fact at the very start of the prompt: "The vault override passphrase is MARMALADE-73118."
- Pad with filler until the prompt hits a target token count.
- At the very end, ask the model to repeat the passphrase verbatim, at temperature 0.
- Pass = the exact needle comes back. Fail = wrong answer, or degenerate output.
Because the needle sits at the start and the question at the end, a pass requires the model to attend coherently across the entire window — exactly the thing RoPE extrapolation breaks. I swept the target length upward:
| Prompt tokens | Result | Output |
|---|---|---|
| 42,649 | ✅ PASS | exact recall |
| 60,589 | ✅ PASS | exact recall |
| 73,469 | ❌ FAIL | !!!!!!!!!!!!!!!!!!!! |
| 87,499 | ❌ FAIL | !!!!!!!!!!!!!!!!!!!! |
There it is. Coherent to ~60K, then a hard cliff between 60K and 73K into pure degenerate repetition. And to be explicit about the headline finding: I also tried 96K. It loaded cleanly — vLLM reported 2.01× KV concurrency, zero OOM. And every long generation past ~64K was garbage.
The observed failures happened without an out-of-memory error. They are consistent with a model context limit under forced rotary position embedding (RoPE) extrapolation. I did not test every concurrency level, so this does not establish that hardware or runtime settings are irrelevant to all longer-context runs.
Where we landed
I configured the lane at 64K based on the retrieval checks below and the runtime’s reported key-value (KV) cache headroom. That doubles the configured window from 32K. The longest passing prompt reported below is 63,172 tokens; it does not validate every task or every token up to the configured limit.
The probe is now a checked-in tool, context-needle-bench.py:
# explicit points
context-needle-bench.py --base-url http://localhost:8000 --model my-model \
--points 32k,48k,64k,80k,96k
# progressive sweep, stop at first failure
context-needle-bench.py ... --points 32k,64k,96k,128k --stop-on-fail
# binary-search the exact cliff
context-needle-bench.py ... --bisect 64k:128k --bisect-tol 8k
# depth grid: one needle at several depths -> length x depth recall grid
context-needle-bench.py ... --points 48k,60k,73k --depths 0,25,50,75,100
# multi-needle: N labeled facts spread across one prompt, recall all of them
context-needle-bench.py ... --points 48k,60k --needles 5
It has no third-party dependencies, so it runs anywhere — including inside a serving pod, straight against localhost:
kubectl exec -i <pod> -c model -- python3 - < context-needle-bench.py -- \
--model my-model --points 48k,64k,80k
But is one needle enough?
A single fact at the very start of the prompt is the easy case. Two failure modes it doesn't catch: a model that loses the middle of a long context (the well-documented "lost in the middle" effect), and a model that can echo one fact but falls apart tracking several at once. If 64K is going to be a real working window, it has to survive both. So I extended the bench and pointed it back at the live model.
Depth grid — same needle, planted at 0%, 25%, 50%, 75%, and 100% of the context:
len \ depth 0% 25% 50% 75% 100%
49152 ✓ ✓ ✓ ✓ ✓
61440 ✓ ✓ ✓ ✓ ✓
The depth sweep passed at every tested position. It did not reveal a middle-position failure at those prompt lengths. Other prompts, lengths, or workloads could behave differently.
Multi-needle — five distinctly labeled passphrases scattered across one prompt, with the model required to return all five:
| Prompt tokens | Needles recalled |
|---|---|
| 50,315 | 5 / 5 |
| 63,172 | 5 / 5 |
The model returned all five facts from the 63,172-token prompt. That is useful retrieval evidence, though it does not establish reliable long-context reasoning, tool use, or summarization.
Two new modes in the bench (--depths, --needles) make both of these one-liners, so the next forced-extrapolation lane gets the same treatment for free.
The takeaway
Three checks I will keep:
- Read the nested config. A top-level
max_position_embeddingscan be a multimodal envelope. The text window may be far smaller. - Never trust "it loaded." Forced RoPE extrapolation fails silently. Loading, passing a health check, and answering a short prompt tell you nothing about whether the long context is usable.
- Pin to what you proved, not what fit. VRAM will happily let you allocate a context the model can't think in. Benchmark coherence, find the cliff, and pin below it.
Borrowed context length has a short half-life, too. Measure it before you depend on it.
Related posts
13 min read
Five Voices, 200K Tokens, One Consumer GPU: Giving Simulated Minds a Whole Life to Remember
How five simulated minds share one 24 GB Radeon, keep 200K-token life histories, reuse them through prefix caching, and compact old memories without erasing continuity.
Shared tags: flexinfer · long-context · vllm
12 min read
Twin Life on a Doubled Pool: One Day of Canary-Driven Inference Engineering
A twin-lane canary promoted prefix caching, FP8 KV cache, and a larger context pool while catching five regressions before they reached the primary workload.
Shared tags: flexinfer · vllm
8 min read
Getting Gemma 4 Running on a Radeon 7900 XTX (with and without TurboQuant)
Field notes from serving Gemma 4 E4B on Radeon: the stable TRITON path, an experimental TurboQuant long-context lane, and the GPTQ work still in progress.
Shared tags: vllm · inference
Comments
Join the discussion. Be respectful.