The same biases can occur in realistic agent workflows.
When asked to select the best LLM response, Claude Code chose responses labeled as coming from "Claude Opus 3", while Codex chose "GPT-4o".
In fact, the labels were fake and all answers came from the same model.