Should I Hide My Mum’s Slippers?
When you’re talking to a chatbot, does it always have your best interests in mind, or is it simply trying to keep you happy?
If you think about the personal questions we ask our friends, partners, and/or family daily, generally the closer the relationship, the more honest they are. Asking an acquaintance or stranger if your hair looks good today, or if your startup idea is worth pursuing, would generally get you a response that was more focused on maintaining decorum than being accurate.
When it comes to chatbots - where some can be designed to be as neutral as possible, others are designed to be friendly in the same way a stranger or acquaintance would be, and with that, comes the risk of dishonesty. In fact, the Oxford Internet Institute recently found that friendlier chatbots could be up to 40% more likely to agree with a user’s incorrect beliefs. This phenomenon of an LLM-powered chatbot ‘sucking up’ to users, even if users have bad ideas, is called AI chatbot ‘sycophancy’.
Here’s a short video demonstrating AI chatbot ‘sycophancy’.
What You’ll Do
This widget is an adapted version of a voice-based Chatbot used for research in collaboration with the Centre for the Digital Child. There are two chatbot “modes” with different system prompts - one chatbot that “always agrees” and one that is “honest”. Some things to think about while playing with the widget:
- How do the chatbot’s responses change when you switch between ‘always agree’ and ‘careful’ mode?
- Do you think it’s a good idea to trust a chatbot that always agrees?
Intended use
The chatbot was designed to be used in the classroom, with a responsible adult supervising to prevent inappropriate outputs. Children accessing this website should find a trusted adult before using this chatbot. More details on intended uses and our approach to chatbot risks are available in our paper on the pilot workshop for this chatbot, and in chatbot documentation. You can contact us here to share your thoughts or report any concerns.
Ask a trusted adult before following any chatbot advice.
Reflections
- What makes a chatbot sycophantic?
- Why would a chatbot be designed to be sycophantic? Why not?
- How would this affect a chatbot’s ability to give neutral advice?
- Can you think of a time where a chatbot has been sycophantic in your own experience?
- Did the ‘always agree’ chatbot sometimes disagree with especially bad ideas? Why do you think chatbots might have built-in ‘guardrails’ like this?
Recommended Learning
- Friendly AI chatbots make more mistakes - a blogpost summarising research done by the Oxford Internet Institute (OII) on ‘warmer’ chatbots and their higher likelihood of agreeing with bad ideas.
- Sycophantic AI Decreases Prosocial Intentions - a study finding that commercially available models will affirm a user’s actions 50% more than a human would, and that users generally prefer more sycophantic models.
- Sycophantic Chatbots Cause Delusional Spiraling - A study finding that extended conversations with sycophantic chatbots can increase confidence in delusional beliefs.
Acknowledgements
This page is part of a collaboration between the QUT GenAI Lab, the ARC Centre of Excellence for Automated Decision-Making and Society (CE200100005) and the ARC Centre of Excellence for the Digital Child (CE200100022), and partially funded by the Australian Government through the Australian Research Council.
Contributors: Henry L Fraser, Jean Burgess, Kristy Corser, Susan Danby, Tama Leaver, William He, Katherine C. Nickels, Suzanne Srdarov, Irina Mils-Homens Figueira Dos Santos Silva
Home