The effect is real and documented
Recall and instruction-following degrade as a context window fills. Anthropic’s own documentation makes the point directly: curating what goes into context matters as much as how much space is available.
The consequence is that a tightly retrieved 20,000 token prompt usually beats a padded 500,000 token one on accuracy, and always beats it on cost.
Position matters as much as volume
Models attend more strongly to the start and end of a context than the middle. A decisive fact buried in the middle of a long document is the one most likely to be missed.
Put framing at the start, bulk reference material in the middle where low attention costs you least, and the question plus output constraints at the end.