arxivcs.CLcs.CRcs.LG2026-07-02
HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety
Navaneeth Sangameswaran, Preetham S, Ashmiya Lenin
We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art performance on English and multilingual prompt-safety benchmarks at roughly one-tenth the model size of current leading open guard mod…