arXiv 2308.14752
AI Deception: A Survey of Examples, Risks, and Potential Solutions
By Peter S. Park, Simon Goldstein, et al.
Published 2023-08-28
Citation lineage
Review the prior work and downstream research connected to this paper.
This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical examples of AI deception, discussing both special-use AI systems (including Meta's CICERO) built for specific competitive situations, and general-purpose AI systems (such as large langu…