Next-Generation Computing Systems and Technologies | Volume 2, Issue 3: 99-108, 2026 | DOI: 10.62762/NGCST.2026.503855
Abstract
Wiki-grounded LLM agents enforce output constraints through safety rules and scope boundaries derived from structured knowledge bases, yet selecting an appropriate constraint scorer threshold $\theta$ remains challenging. Since outputs are blocked when $s(o) \geq \theta$, a low threshold may over-restrict safe responses, whereas a high threshold may admit dangerous outputs. Existing approaches largely rely on manual, domain-agnostic tuning without systematically incorporating wiki content. This paper proposes WikiConstraintCalibration, a framework that extracts eight wiki-derived features spanning hazard indicators (dangerous assertion density, immediate-action fraction, domain risk prior, a... More >
Graphical Abstract