Monday, August 31, 2026
Catatonic Times
No Result
View All Result
  • Home
  • Crypto Updates
  • Bitcoin
  • Ethereum
  • Altcoin
  • Blockchain
  • NFT
  • Regulations
  • Analysis
  • Web3
  • More
    • Metaverse
    • Crypto Exchanges
    • DeFi
    • Scam Alert
  • Home
  • Crypto Updates
  • Bitcoin
  • Ethereum
  • Altcoin
  • Blockchain
  • NFT
  • Regulations
  • Analysis
  • Web3
  • More
    • Metaverse
    • Crypto Exchanges
    • DeFi
    • Scam Alert
No Result
View All Result
Catatonic Times
No Result
View All Result

Perplexity Launches WANDR Benchmark For Measuring Large-Scale Research Capabilities Of AI Agents

by Catatonic Times
July 15, 2026
in Metaverse
Reading Time: 4 mins read
0 0
A A
0
Home Metaverse
Share on FacebookShare on Twitter

[ad_1]

Alisa Davidson

by
Alisa Davidson


Printed: July 15, 2026 at 6:57 am Up to date: July 15, 2026 at 6:57 am

by Anastasiia O


Edited and fact-checked:
July 15, 2026 at 6:57 am

To enhance your local-language expertise, generally we make use of an auto-translation plugin. Please word auto-translation might not be correct, so learn unique article for exact data.

Perplexity Launches WANDR Benchmark For Measuring Large-Scale Research Capabilities Of AI Agents

Perplexity AI has launched WANDR (Broad ANd Deep Analysis), an open benchmark designed to judge how successfully synthetic intelligence programs carry out large-scale analysis duties that require each broad data discovery and detailed proof assortment. The framework comprises 500 practical data-gathering duties modeled on skilled information work, together with market evaluation, due diligence, literature critiques, aggressive intelligence, product comparisons, and expertise sourcing.

In contrast to conventional AI benchmarks that concentrate on producing a single reply or a written report, WANDR measures an AI system’s capacity to establish giant numbers of related entities and confirm every consequence with supporting proof. The benchmark is meant to mirror real-world analysis workflows, the place success relies upon not solely on discovering correct data but in addition on attaining complete protection throughout lots of and even 1000’s of data.

Based on Perplexity, present AI programs proceed to face vital challenges on this space. Even the highest-performing mannequin within the firm’s analysis achieved a tender F1 rating of 0.363 and a tough F1 rating of 0.133, indicating that wide-scale, evidence-backed analysis stays removed from being absolutely automated. The benchmark contains greater than 170,000 source-backed data throughout its 500 duties, offering a large-scale testing atmosphere for research-oriented AI brokers.

We’re open sourcing WANDR.

WANDR is an inside benchmark we constructed and used for constructing deep and extensive analysis capabilities inside Perplexity Pc.https://t.co/gp2BWjFK4d

— Perplexity (@perplexity_ai) July 14, 2026

Benchmark Outcomes Spotlight Present AI Analysis Limitations

WANDR makes use of a reference-free analysis course of that verifies every submitted declare towards the proof cited by the AI system, reasonably than evaluating outcomes with a set reply key. Each declare is checked for supply high quality, factual accuracy, relevance, and whether or not the supporting excerpts genuinely substantiate the knowledge offered. This strategy is meant to higher mirror real-world analysis, the place data modifications over time and full reply units are troublesome to take care of.

The benchmark additionally supplies detailed diagnostics to establish the place AI programs fail throughout advanced analysis duties. Efficiency could be measured throughout a number of phases, together with data discovery, information enrichment, id matching, supply validation, and proof extraction, permitting builders to pinpoint weaknesses past total accuracy scores.

Perplexity evaluated six manufacturing AI analysis programs utilizing WANDR beneath similar testing circumstances. Its Search as Code (SaC) platform achieved the best total efficiency, recording a tender F1 rating of 0.363 and a tough F1 rating of 0.133. Anthropic ranked second with scores of 0.249 and 0.072, whereas different evaluated programs didn’t exceed a tender F1 rating of 0.121. The examine additionally discovered that rising computational effort typically improved efficiency for a number of fashions, though increased prices and longer processing occasions didn’t persistently translate into higher outcomes.

The corporate stated the benchmark is meant to function an open useful resource for researchers and builders engaged on AI-powered search and analysis programs. Past benchmarking, WANDR might also assist future reinforcement studying methods by offering structured suggestions at every stage of the analysis course of, enabling AI fashions to enhance not solely factual accuracy but in addition planning, protection, and proof assortment at scale.

Disclaimer

In step with the Belief Challenge pointers, please word that the knowledge offered on this web page just isn’t supposed to be and shouldn’t be interpreted as authorized, tax, funding, monetary, or another type of recommendation. It is very important solely make investments what you’ll be able to afford to lose and to hunt unbiased monetary recommendation in case you have any doubts. For additional data, we propose referring to the phrases and circumstances in addition to the assistance and assist pages offered by the issuer or advertiser. MetaversePost is dedicated to correct, unbiased reporting, however market circumstances are topic to alter with out discover.

About The Writer


Alisa, a devoted journalist on the MPost, makes a speciality of crypto, AI, investments, and the expansive realm of Web3. With a eager eye for rising developments and applied sciences, she delivers complete protection to tell and interact readers within the ever-evolving panorama of digital finance.

Extra articles


Alisa, a devoted journalist on the MPost, makes a speciality of crypto, AI, investments, and the expansive realm of Web3. With a eager eye for rising developments and applied sciences, she delivers complete protection to tell and interact readers within the ever-evolving panorama of digital finance.








Extra articles

[ad_2]

Source link

Tags: AgentsBenchmarkCapabilitiesLargeScaleLaunchesMeasuringPerplexityResearchWANDR
Previous Post

Hyperliquid Bitcoin Longs Just Topped Levels From Q2’s $83K Run

Next Post

The Math Behind ‘1,200 Digital Workers’

Related Posts

MEXC Introduces 31 New 0-Fee Stock and ETF Futures Across Key Market Sectors
Metaverse

MEXC Introduces 31 New 0-Fee Stock and ETF Futures Across Key Market Sectors

August 12, 2026
AI Diaries: Weekly AI News and Updates (August 12, 2026)
Metaverse

AI Diaries: Weekly AI News and Updates (August 12, 2026)

August 13, 2026
Syntetika Launches Tokenization Hub Bringing Regulated Investment Strategies Onchain
Metaverse

Syntetika Launches Tokenization Hub Bringing Regulated Investment Strategies Onchain

August 10, 2026
Quantum Teleportation: The Existential Philosophy of the Star Trek Transporter
Metaverse

Quantum Teleportation: The Existential Philosophy of the Star Trek Transporter

August 10, 2026
The Real-Life Mission Impossible Masks: Identity in the Age of Digital Face Hacking
Metaverse

The Real-Life Mission Impossible Masks: Identity in the Age of Digital Face Hacking

August 8, 2026
Gate Update: Exchange Tops Global Net Inflow Rankings And Opens Moonshot AI Pre-IPO As gStocks And Gold Drive Market Activity
Metaverse

Gate Update: Exchange Tops Global Net Inflow Rankings And Opens Moonshot AI Pre-IPO As gStocks And Gold Drive Market Activity

August 8, 2026
Next Post
The Math Behind ‘1,200 Digital Workers’

The Math Behind '1,200 Digital Workers'

BIL Suisse Renews Strategic Partnership with Avaloq

BIL Suisse Renews Strategic Partnership with Avaloq

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Catatonic Times

Stay ahead in the cryptocurrency world with Catatonic Times. Get real-time updates, expert analyses, and in-depth blockchain news tailored for investors, enthusiasts, and innovators.

Categories

  • Altcoin
  • Altcoins
  • Analysis
  • Base
  • Bitcoin
  • Blockchain
  • Blockchain Infrastructure
  • Crypto Exchanges
  • Crypto Updates
  • DeFi
  • Ethereum
  • Metaverse
  • NFT
  • Regulations
  • Scam Alert
  • Security / Scam Alerts
  • Solana
  • Uncategorized
  • Web3

Latest Updates

  • Solana Fees Surge as Inflation Cuts Transform Ecosystem
  • DeFi Security Vulnerabilities: $83M Loss Calls for Oversight
  • Tectonic Exploit: A Stark Reminder of DeFi Lending Security Vulnerabilities
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA
  • Cookie Privacy Policy
  • Terms and Conditions
  • Contact Us

Copyright © 2024 Catatonic Times.
Catatonic Times is not responsible for the content of external sites.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
âš¡ Get Real-time Live Crypto Updates to your email

No Result
View All Result
  • Home
  • Crypto Updates
  • Bitcoin
  • Ethereum
  • Altcoin
  • Blockchain
  • NFT
  • Regulations
  • Analysis
  • Web3
  • More
    • Metaverse
    • Crypto Exchanges
    • DeFi
    • Scam Alert

Copyright © 2024 Catatonic Times.
Catatonic Times is not responsible for the content of external sites.