Source

Towards Understanding Sycophancy in Language Models

peer-reviewed conference paper Sharma et al. arXiv:2310.13548, ICLR 2024

Claim this source supports

“Five AI assistants showed sycophancy; responses matching a user's views were more likely to be preferred, and humans and preference models sometimes preferred convincing sycophantic responses over correct ones.”

View original source

Cited in