Two-faced AI language models learn to hide deception

2 min read Original article ↗

Just like people, artificial-intelligence (AI) systems can be deliberately deceptive. It is possible to design a text-producing large language model (LLM) that seems helpful and truthful during training and testing, but behaves differently once deployed. And according to a study shared this month on arXiv1, attempts to detect and remove such two-faced behaviour are often useless — and can even make the models better at hiding their true nature.

doi: https://doi.org/10.1038/d41586-024-00189-3

References

Subjects

Latest on:

Nature Careers

Jobs

  • Faculty Positions in Westlake University

    Founded in 2018, Westlake University is a new type of non-profit research-oriented university in Hangzhou, China, supported by public a...

    Hangzhou, Zhejiang, China

    Westlake University