arxivcs.AI2026-07-09
Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench
Siddhartha Jain, Ameya Velingker
Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summari…