jhalloran
- Karma
- 2
- Created
- 23 days ago
About
Researcher working on LLM alignment and preference optimization, focused on agentic/tool-use safety (MCP). Papers: arxiv.org/abs/2505.23634, arxiv.org/abs/2605.11217. Code: github.com/johnhalloran321Recent Submissions
- 1. ▲ Refusal training for LLM agents against disguised MCP attacks (github.com)