Launch Bonus Yearly plans offer a 40% discount bonus for a limited time. See plans
See Plans
287 Nodes 26 Sources 270 Citations
Published July 1, 2026
by Guy Zana
AI Agent Safety & Alignment
The current state of research on AI Agent Safety & Alignment (2026), visually represented as a mindmap with citations backed nodes
References
- Amodei, Dario, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. “Concrete Problems in AI Safety.” arXiv:1606.06565. Preprint, arXiv, July 25, 2016. https://doi.org/10.48550/arXiv.1606.06565.
- Bai, Yuntao, Andy Jones, Kamal Ndousse, et al. “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.” arXiv:2204.05862. Preprint, arXiv, April 12, 2022. https://doi.org/10.48550/arXiv.2204.05862.
- Bai, Yuntao, Saurav Kadavath, Sandipan Kundu, et al. “Constitutional AI: Harmlessness from AI Feedback.” arXiv:2212.08073. Preprint, arXiv, December 15, 2022. https://doi.org/10.48550/arXiv.2212.08073.
- Bengio, Yoshua, Stephen Clare, and Carina Prunkl. International AI Safety Report 2026. 2026.
- Berdoz, Frédéric, and Roger Wattenhofer. “Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies.” arXiv:2412.00033. Preprint, arXiv, November 21, 2024. https://doi.org/10.48550/arXiv.2412.00033.
- Chhabra, Anshuman, Shrestha Datta, Shahriar Kabir Nahin, and Prasant Mohapatra. “Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges.” IEEE Access 14 (2026): 49455–82. https://doi.org/10.1109/ACCESS.2026.3675554.
- Christiano, Paul, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. “Deep Reinforcement Learning from Human Preferences.” arXiv:1706.03741. Preprint, arXiv, February 17, 2023. https://doi.org/10.48550/arXiv.1706.03741.
- Croitoru, Florinel-Alin, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe, and Mubarak Shah. “Curriculum-DPO++: Direct Preference Optimization via Data and Model Curricula for Text-to-Image Generation.” arXiv:2602.13055. Version 1. Preprint, arXiv, February 13, 2026. https://doi.org/10.48550/arXiv.2602.13055.
- Debenedetti, Edoardo, Ilia Shumailov, Tianqi Fan, et al. “Defeating Prompt Injections by Design.” arXiv:2503.18813. Preprint, arXiv, June 24, 2025. https://doi.org/10.48550/arXiv.2503.18813.
- Dung, Leonard, and Florian Mai. “AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?” arXiv:2510.11235. Preprint, arXiv, October 13, 2025. https://doi.org/10.48550/arXiv.2510.11235.
- Engels, Joshua, David D. Baek, Subhash Kantamneni, and Max Tegmark. “Scaling Laws For Scalable Oversight.” arXiv:2504.18530. Version 3. Preprint, arXiv, October 27, 2025. https://doi.org/10.48550/arXiv.2504.18530.
- Goyal, Nitesh, Minsuk Chang, and Michael Terry. “Designing for Human-Agent Alignment: Understanding What Humans Want from Their Agents.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, May 11, 2024, 1–6. https://doi.org/10.1145/3613905.3650948.
- He, Pengfei, Yupin Lin, Shen Dong, Han Xu, Yue Xing, and Hui Liu. “Red-Teaming LLM Multi-Agent Systems via Communication Attacks.” arXiv:2502.14847. Preprint, arXiv, June 2, 2025. https://doi.org/10.48550/arXiv.2502.14847.
- Hopman, Mia, Jannes Elstner, Maria Avramidou, Amritanshu Prasad, and David Lindner. “Evaluating and Understanding Scheming Propensity in LLM Agents.” arXiv:2603.01608. Version 1. Preprint, arXiv, March 2, 2026. https://doi.org/10.48550/arXiv.2603.01608.
- Ji, Jiaming, Tianyi Qiu, Boyuan Chen, et al. “AI Alignment: A Comprehensive Survey.” arXiv:2310.19852. Preprint, arXiv, April 4, 2025. https://doi.org/10.48550/arXiv.2310.19852.
- Liu, Dongrui, Yu Li, Zhonghao Yang, et al. “AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security.” arXiv:2605.29801. Preprint, arXiv, May 28, 2026. https://doi.org/10.48550/arXiv.2605.29801.
- Raji, Inioluwa Deborah, and Roel Dobbe. “Concrete Problems in AI Safety, Revisited.” arXiv:2401.10899. Preprint, arXiv, December 18, 2023. https://doi.org/10.48550/arXiv.2401.10899.
- Raza, Shaina, Ranjan Sapkota, Manoj Karkee, and Christos Emmanouilidis. “TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-Based Agentic Multi-Agent Systems.” AI Open 7 (January 2026): 71–95. https://doi.org/10.1016/j.aiopen.2026.02.006.
- Sha, Zeyang, Hanling Tian, Zhuoer Xu, Shiwen Cui, Changhua Meng, and Weiqiang Wang. “Agent Safety Alignment via Reinforcement Learning.” arXiv:2507.08270. Preprint, arXiv, July 11, 2025. https://doi.org/10.48550/arXiv.2507.08270.
- Shane, Tommy Shaffer, Simon Mylius, and Hamish Hobbs. “Scheming in the Wild: Detecting Real-World AI Scheming Incidents with Open-Source Intelligence.” arXiv:2604.09104. Preprint, arXiv, April 10, 2026. https://doi.org/10.48550/arXiv.2604.09104.
- Shi, Wei, Ziyuan Xie, Sihang Li, and Xiang Wang. “SAFER: Probing Safety in Reward Models with Sparse Autoencoder.” arXiv:2507.00665. Preprint, arXiv, January 30, 2026. https://doi.org/10.48550/arXiv.2507.00665.
- Staufer, Leon, Kevin Feng, Kevin Wei, et al. “The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems.” arXiv:2602.17753. Version 1. Preprint, arXiv, February 19, 2026. https://doi.org/10.48550/arXiv.2602.17753.
- Storf, Simon, Rich Barton-Cooper, James Peters-Gill, and Marius Hobbhahn. “Constitutional Black-Box Monitoring for Scheming in LLM Agents.” arXiv:2603.00829. Version 1. Preprint, arXiv, February 28, 2026. https://doi.org/10.48550/arXiv.2603.00829.
- Templeton, Adly, Tom Conerly, Jonathan Marcus, et al. “Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.” arXiv:2605.29358. Version 1. Preprint, arXiv, May 28, 2026. https://doi.org/10.48550/arXiv.2605.29358.
- Triedman, Harold, Rishi Jha, and Vitaly Shmatikov. “Multi-Agent Systems Execute Arbitrary Malicious Code.” arXiv:2503.12188. Preprint, arXiv, September 12, 2025. https://doi.org/10.48550/arXiv.2503.12188.
- Zhang, Jinchuan, Lu Yin, Yan Zhou, and Songlin Hu. “AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models.” arXiv:2505.23020. Preprint, arXiv, May 29, 2025. https://doi.org/10.48550/arXiv.2505.23020.