Agent Benchmark Suite
5 standard evaluation tasks · measures factuality, narrative quality, risk calibration, actionability
Runs each task through the full swarm (auction → execution → eval). Cycles runs 1→4 to demonstrate the learning arc across mission types.
#1 · Run 1 — Baseline
Market Sizing
pending
Analyze the market size and competitive landscape for a B2B SaaS tool targeting mid-market legal firms that want to automate contract review.
lead: ResearchAgent (unverified)
Unverified claim detected. ResearchAgent penalized.
Quality
--
Factual
--
Useful
--
Specific
--
Action
--
Collab
--
#2 · Run 2 — Market Learned
Go-to-Market
pending
Design a go-to-market strategy for a developer productivity CLI tool that reduces boilerplate code setup. Target: senior engineers at Series A–C startups.
lead: ResearchAgent + SourceVerifierAgent
SourceVerifier paired with Research. Factuality +26 pts.
Quality
--
Factual
--
Useful
--
Specific
--
Action
--
Collab
--
#3 · Run 3 — Over-Rotation
Risk Analysis
pending
Conduct a comprehensive risk analysis for an AI-generated content startup planning to launch in the EU under the new AI Act regulatory environment.
lead: SkepticAgent (over-weighted)
SkepticAgent over-rotated. Risk-heavy pitch reduces usefulness.
Quality
--
Factual
--
Useful
--
Specific
--
Action
--
Collab
--
#4 · Run 4 — Calibrated Peak
Pitch Script
pending
Write a 90-second demo pitch for an IoT fleet management platform that reduces truck idle time by 34%. Target: Series A VCs with logistics portfolio.
lead: PitchAgent + BuilderAgent (synergy)
PitchAgent + BuilderAgent synergy. Best score across all dimensions.
Quality
--
Factual
--
Useful
--
Specific
--
Action
--
Collab
--
#5 · Run 2 — Verified
Positioning
pending
Define product positioning and landing page copy for a real-time collaborative whiteboard built for remote engineering teams doing system design interviews.
lead: BuilderAgent + SourceVerifierAgent
Collaborative swarm. All claims sourced. High actionability.
Quality
--
Factual
--
Useful
--
Specific
--
Action
--
Collab
--
Ask SwarmDAQ
Command the agent market. Run missions, inspect bids, rebalance swarms, replay traces, and expose weak agents.
Try: "Run a mission for RocketRide" or "Which agent is overvalued?"

Powered by CopilotKit