Perplexity has released the WANDR benchmark, an open evaluation designed to test whether AI research agents can find large sets of information and support every result with verifiable evidence.
This is a more demanding test than asking an AI tool to answer one question or produce a polished report. WANDR measures whether an agent can search broadly, investigate every qualifying result, and avoid quietly leaving large gaps in the final dataset.
That matters for marketers using AI for competitor mapping, account research, content planning, market analysis, sales prospecting, and customer intelligence. The benchmark results suggest that current systems can find useful information, but complete, consistently supported research remains difficult.
What was announced
Perplexity published the WANDR benchmark on July 14, 2026. WANDR stands for Wide ANd Deep Research and contains 500 realistic data-collection tasks based on de-identified patterns from production use.
Each task requires two types of work:
Wide research: Finding a large or open-ended set of qualifying entities.
Deep research: Investigating each entity far enough to support the requested claims with evidence.
A task might ask an agent to find dozens of companies, identify relevant executives at each company, and provide supporting sources for every person. A few accurate examples would not count as a complete answer.
Across the 500 tasks, the median assignment requires 50 qualifying members and 245 records in total. The full benchmark calls for 170,495 source-backed records. Perplexity says the tasks cover work such as competitive mapping, due diligence, literature review, market analysis, product comparison, and talent sourcing.
The grading process also differs from benchmarks built around a fixed answer key. WANDR re-fetches the pages cited by an agent and checks whether the evidence supports each submitted record. This helps the benchmark evaluate changing, open-ended questions without relying entirely on a static reference list.
Perplexity has released the tasks, evaluation harness, and technical report publicly. The results and comparisons remain Perplexity-reported rather than independently reproduced at scale.
Why the WANDR benchmark matters for marketers
The obvious marketing use case is competitive intelligence.
A marketer may ask an AI agent to find every competitor serving a certain industry, identify their pricing pages, record recent product launches, collect customer examples, and map their positioning. A system that finds 12 convincing companies can look useful, even when 30 other qualifying companies were missed.
That is the problem the WANDR benchmark exposes.
For B2B SaaS marketing, incomplete research can distort decisions. A missing competitor affects positioning. Missing accounts weaken an ABM list. Unsupported claims damage a comparison page. Incomplete customer research can lead a team to build campaigns around a market segment that only appears more important because the evidence was patchy.
The benchmark also changes how teams should think about AI marketing automation. Automating research is not valuable simply because an agent produces a spreadsheet faster. The output still needs coverage checks, evidence standards, duplicate handling, and a clear definition of what counts as complete.
The current results show why that control matters.
In Perplexity’s evaluation, its Search as Code system led with a soft F1 score of 0.363 and a hard F1 score of 0.133. Anthropic followed with 0.249 soft F1 and 0.072 hard F1. Perplexity reported that no evaluated system combined leading quality with leading efficiency.
These are vendor-reported benchmark results, so they should not be treated as an independent verdict on the systems involved. The larger signal is the low hard-completion score. Even the leading setup received complete credit for only a small share of the requested entities.
Evidence quality is the real bottleneck
The benchmark found that reaching a plausible webpage was usually easier than proving every part of a claim.
Across the evaluated systems, many submitted pages failed at least one substantive task requirement. An even larger share of selected excerpts failed to support everything the record claimed. For Perplexity’s system, 41.4 percent of submitted pages failed at least one substantive requirement, while 57.5 percent of excerpts did not fully support the associated claim.
This distinction matters for content and SEO teams.
An AI agent may find a company page that mentions analytics, for example, but that page may not prove the company serves enterprise SaaS businesses, operates in Europe, or launched the feature during the requested period. A relevant page is not automatically sufficient evidence.
Marketing teams working on AI search visibility face the same issue from the other direction. If agents increasingly collect, compare, and verify information before recommending a vendor, brands need pages that make important claims easy to locate and confirm.
Clear product descriptions, dated announcements, customer evidence, pricing context, structured comparisons, author information, and specific source links become more useful than vague category copy.
How WANDR compares with other research evaluations
WANDR is positioned as the wide counterpart to Perplexity’s DRACO benchmark.
DRACO evaluates whether an agent can create an accurate, complete, and objective long-form research report. WANDR evaluates whether it can build a large structured collection in which every member is backed by evidence.
Traditional question-answering benchmarks usually test whether a system reaches one correct answer. That is useful for measuring reasoning or factual accuracy, but it does not show whether the system can discover an unknown number of valid results and maintain evidence quality across all of them.
WANDR therefore tests a task shape closer to marketing operations:
| Evaluation type | Main output | Marketing relevance |
|---|---|---|
| Question-answering benchmark | One answer | Useful for factual lookup and simple analysis |
| Deep research benchmark | Long-form report | Useful for market reports and strategic briefs |
| WANDR benchmark | Large evidence-backed collection | Useful for competitor maps, account lists, category research and structured datasets |
The distinction also connects with Sakana Fugu and multi-agent AI workflows. Coordinating several agents can help divide a large research task, but orchestration alone does not guarantee complete discovery or reliable evidence. It can also distribute the same weak assumption across hundreds of records.
What marketing teams should watch next
Marketing teams should avoid judging research agents by how professional the final document looks.
A clean table can still be incomplete. Citations can still point to pages that only partially support a claim. A large result count can include duplicates, weak matches, or entities that fail the original criteria.
Before using an agent for production research, test it on a task where your team already knows most of the answer. Check four things:
- How many qualifying records did it miss?
- Does every source support the full claim?
- How does quality change as the requested list grows?
- Can another reviewer reproduce the result from the evidence provided?
Teams should also measure research cost at the workflow level. Higher effort settings can improve results, but Perplexity’s evaluation found costs ranging from a few cents to hundreds of dollars per task across systems and configurations. More compute did not improve every system consistently.
Finally, connect research output to actual marketing measurement. If AI-assisted content or research affects discovery, teams can use this guide to tracking AI traffic in GA4 as a starting point. Referral traffic will not measure every influence, but it can help identify whether AI platforms are sending visitors to the pages supplying the evidence.
OneMetrik Takeaway
The WANDR benchmark delivers an inconvenient but useful message: AI can produce research faster than most teams can review it, but speed is not the same as completeness.
At OneMetrik, we would use wide-and-deep research for first-pass market mapping, competitor discovery, account enrichment, and evidence collection. We would not let it make an unchecked positioning, targeting, or GTM decision.
The practical rule is simple. Use agents to expand the search. Use clear criteria and human review to decide what survives it.