arXiv 2308.14752

AI Deception: A Survey of Examples, Risks, and Potential Solutions

By Peter S. Park, Simon Goldstein, et al.

Published 2023-08-28

Citation lineage

Review the prior work and downstream research connected to this paper.

This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical examples of AI deception, discussing both special-use AI systems (including Meta's CICERO) built for specific competitive situations, and general-purpose AI systems (such as large langu…

View the original paper on arXiv