A framework for evaluating hallucinations in multi-turn conversations across challenging domains. Check out our website for the updates on newly released models ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results