A quantitative agent exchange built on auction theory, bandit learning, and Bayesian reputation.
The user submits a free-form mission. PlannerAgent decomposes it into a structured task graph: market_research → positioning → landing_page_copy → pitch_script → risk_review → final_eval. Each task carries required skills that agents will bid against.
Each task triggers an open auction. All eligible agents submit bids with claimed confidence, cost, latency, and expected quality. MarketMaker runs a Vickrey-inspired mechanism: the highest utility bid wins, but pays the clearing price of the second-best bid. This approximates a truthful-style allocation mechanism for the demo.
MarketMaker combines seven signals into a final agent score: skill match, Bayesian trust, UCB1 exploration bonus, bid utility, normalized Elo, PageRank trust centrality, and collaboration history. The UCB1 term ensures under-tested agents get chances — preventing the system from collapsing to always the same agents.
Every LLM call is wrapped with weave.wrapGoogleGenAI() and traced to W&B Weave in real time. The live Weave traces panel in the demo fetches call counts, token totals, avg latency, and per-call op names via the Weave REST API. Per-run cost is computed from real usageMetadata (Gemini 2.5 Flash: $0.075/1M input, $0.30/1M output).
Agent state persists across runs via Redis Cloud (TCP). Redis Hashes store agent reputation snapshots, Sorted Sets power leaderboards, Streams preserve market events, and t-digest or rolling quantiles model price/latency anomalies. Mem0 stores agent memory across missions. Without credentials, an in-memory store provides full functionality with identical API.
EvaluatorAgent scores each output on six dimensions: quality, factuality, usefulness, specificity, actionability, and collaboration. Bayesian Beta reputation updates immediately. Elo ratings shift based on relative agent performance. The evaluator catches hallucinations — like ResearchAgent's unsourced market-size claim on run 1.
Run 1: ResearchAgent makes an unsourced claim — penalized −4 rep, −18 Elo. Run 2: market pairs ResearchAgent + SourceVerifierAgent, factuality +26. Run 3: market over-rotates on SkepticAgent — score regresses 91→85, but factuality holds. Run 4: PitchAgent + BuilderAgent synergy unlocked, all dimensions peak at 96. The regression in run 3 is intentional — real markets overshoot before calibrating.
SwarmDAQ is not just prompt chaining. Every routing decision is grounded in math.
Exploration/exploitation balance
Shrinking uncertainty over time
Truth-telling bid mechanism
Pairwise agent duel ratings
Swarm quality/risk/cost tradeoff
Leave-one-out credit attribution
Collaboration hub centrality
Run 1: ResearchAgent makes an unsourced claim — penalized. Run 2: SourceVerifier paired in — factuality jumps 26 pts. Run 3: Market over-rotates, SkepticAgent displaces PitchAgent — score regresses. Run 4: Balanced swarm. New peak. The regression in run 3 is the point — real markets overshoot.
Powered by CopilotKit