Wandr Benchmark: Evaluating Research Agents That Must Search Wide and Deep research.perplexity.ai 1 points by tagawa a month ago · 1 comment Reader PiP Save Collapse all Expand all 1 thread tagawaOP a month ago Repo with benchmark tasks, evaluation harness, tech report:https://github.com/perplexityai/wandr