Skip to content

benchmark

NoLiMa

Long-context benchmark that plants a fact whose wording barely overlaps the question, so a model must infer the association instead of string-matching it.

Known aliases

  • arXiv:2502.05167

Relationships

No evidence-backed relationships are recorded.

Current clusters