Sharing on Mastodon:
An eval harness found what qualitative review couldn't: AI models are most confident when wrong
Save
Home
About