Wandr Benchmark: Evaluating Research Agents That Must Search Wide and Deep research.perplexity.ai 1 points by tagawa 3 months ago · 1 comment Reader PiP Save Collapse all Expand all 1 thread tagawaOP 3 months ago Repo with benchmark tasks, evaluation harness, tech report:https://github.com/perplexityai/wandr