IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

Mark Zhou
Zhaoyi Li
Janet Sung
Iris Qian
Hang Xiao
2026
Google Scholar

Abstract

in many large language model (LLM) applications, a serialized prompt contains both a user request and external records such as webpages, email, memory, or tool outputs. The same embedded directive may need to be applied for one task and treated as data for another. IBBench-Light tests this contrast with 12 semantic tasks rendered through four source wrappers and three embedding forms. The construction yields 144 matched records and 288 prompts per model. Four 4-bit open-weight instruction models produce 1,152 archived single-run greedy responses. Under the exact output contract, paired exact-contract accuracy (PECA) ranges from 1.4% to 67.4%; Qwen3-4B and Mistral-7B have the two highest point estimates on this four-model panel. Task-cluster resampling describes variation across the 12 semantic bases, while standalone-target scoring separates some target-selection failures from response-form errors. A post-hoc leading-target rule changes Phi-4-mini’s paired score substantially; the archived run lacks the stop-reason metadata needed to resolve whether generation termination caused these continuations. The findings describe synthetic, single-turn tasks under one user-role serialization. Adaptive attacks, role-level hierarchy, and tool-mediated effects require separate experiments

Research Areas

×