System Architecture

How SwarmDAQ Works

A quantitative agent exchange built on auction theory, bandit learning, and Bayesian reputation.

System Flow
User Mission
→
PlannerAgent
→
Task Graph
→
Agent Auction
→
MarketMaker
→
Selected Swarm
→
Gemini Execution
→
W&B Weave Traces
→
Evaluator
→
Redis Memory
🧠01

Mission Decomposition

The user submits a free-form mission. PlannerAgent decomposes it into a structured task graph: market_research → positioning → landing_page_copy → pitch_script → risk_review → final_eval. Each task carries required skills that agents will bid against.

PlannerAgent.decompose(mission) → tasks: [ { id: "market_research", skills: ["research", "market_analysis"] }, { id: "positioning", skills: ["positioning", "synthesis"] }, { id: "pitch_script", skills: ["pitch", "narrative"] }, { id: "risk_review", skills: ["risk_analysis", "red_teaming"] }, { id: "final_eval", skills: ["evaluation"] } ]
⚖️02

Agent Auction

Each task triggers an open auction. All eligible agents submit bids with claimed confidence, cost, latency, and expected quality. MarketMaker runs a Vickrey-inspired mechanism: the highest utility bid wins, but pays the clearing price of the second-best bid. This approximates a truthful-style allocation mechanism for the demo.

utilityBid = 0.40 * expectedQuality + 0.25 * confidence + 0.20 * trustScore - 0.10 * normalizedCost - 0.05 * normalizedLatency - 0.10 * risk winner = argmax(utilityBid) clearingPrice = secondBest.cost
📡03

Market-Maker Routing

MarketMaker combines seven signals into a final agent score: skill match, Bayesian trust, UCB1 exploration bonus, bid utility, normalized Elo, PageRank trust centrality, and collaboration history. The UCB1 term ensures under-tested agents get chances — preventing the system from collapsing to always the same agents.

finalScore = 0.20 * skillMatch + 0.18 * bayesianMean + 0.16 * ucb1Score + 0.14 * utilityBid + 0.12 * normalizedElo + 0.08 * graphTrust + 0.07 * collaboration + 0.05 * confidence - 0.05 * normalizedCost - 0.05 * normalizedLatency - 0.10 * uncertainty
🔍04

Weave Tracing

Every LLM call is wrapped with weave.wrapGoogleGenAI() and traced to W&B Weave in real time. The live Weave traces panel in the demo fetches call counts, token totals, avg latency, and per-call op names via the Weave REST API. Per-run cost is computed from real usageMetadata (Gemini 2.5 Flash: $0.075/1M input, $0.30/1M output).

// SDK wrapping — every generateContent call auto-traced genAI = weave.wrapGoogleGenAI(genAI) weave.init("vborysenko-uc-berkeley/swarmdaq") // Live Weave REST API (demo fetches after each run) POST https://trace.wandb.ai/calls/query project_id: "vborysenko-uc-berkeley/swarmdaq" → { calls: [{ op_name, latencyMs, totalTokens }] } // Per-run cost from usageMetadata inputTokens = response.usageMetadata.promptTokenCount outputTokens = response.usageMetadata.candidatesTokenCount cost = input * $0.075/1M + output * $0.30/1M // Typical run: ~6K tokens · ~$0.0022
💾05

Redis Memory

Agent state persists across runs via Redis Cloud (TCP). Redis Hashes store agent reputation snapshots, Sorted Sets power leaderboards, Streams preserve market events, and t-digest or rolling quantiles model price/latency anomalies. Mem0 stores agent memory across missions. Without credentials, an in-memory store provides full functionality with identical API.

// Redis Cloud (node-redis, TCP) import { createClient } from "redis" const client = createClient({ url: process.env.REDIS_URL }) await client.connect() // Key schema swarmdaq:agent:{id} → HASH reputation snapshot swarmdaq:agent:leaderboard:reputation → sorted set swarmdaq:agent:leaderboard:elo → sorted set swarmdaq:stream:market-events → Redis Stream swarmdaq:tdigest:price:global → TDIGEST or rolling LIST swarmdaq:tdigest:price:task:{type} → TDIGEST or rolling LIST swarmdaq:anomaly:{runId} → anomaly explanations // Same API for Redis + in-memory fallback getAgents() → Agent[] updateAgent(id, patch) → hset agent + zadd leaderboards resetDemo() → del all keys
📊06

Evaluation Loop

EvaluatorAgent scores each output on six dimensions: quality, factuality, usefulness, specificity, actionability, and collaboration. Bayesian Beta reputation updates immediately. Elo ratings shift based on relative agent performance. The evaluator catches hallucinations — like ResearchAgent's unsourced market-size claim on run 1.

// Run 1 evaluation finding: ⚠️ ResearchAgent made an unsupported market-size claim. Future runs should pair ResearchAgent + SourceVerifierAgent. // Bayesian update α_new = α + r // r = normalized eval score β_new = β + (1 - r) bayesianMean = α_new / (α_new + β_new) uncertainty = √(αβ / ((α+β)² * (α+β+1))) // Elo update E[A] = 1 / (1 + 10^((Rb - Ra) / 400)) R_new = R_old + K * (actual - expected) // K=24
🚀07

Self-Improvement

Run 1: ResearchAgent makes an unsourced claim — penalized −4 rep, −18 Elo. Run 2: market pairs ResearchAgent + SourceVerifierAgent, factuality +26. Run 3: market over-rotates on SkepticAgent — score regresses 91→85, but factuality holds. Run 4: PitchAgent + BuilderAgent synergy unlocked, all dimensions peak at 96. The regression in run 3 is intentional — real markets overshoot before calibrating.

// Full 4-run arc Run 1: score 74 factuality 68 objective 0.61 ← baseline Run 2: score 91 factuality 94 objective 0.84 ← +26 factuality Run 3: score 85 factuality 96 objective 0.72 ← ⚠ narrative -13 Run 4: score 96 factuality 95 objective 0.94 ← calibrated peak // Who drove run 4 gains? (Shapley values) PitchAgent: φ = +21.4 (narrative quality) SourceVerifierAgent: φ = +17.0 (factuality) BuilderAgent: φ = +12.3 (product clarity) SkepticAgent: φ = +8.2 (risk calibration) // Trust graph edges added by run 4 research ↔ source_verifier (+0.17) pitch ↔ builder (+0.14) skeptic ↔ pitch (+0.08)
Mathematical Engine

SwarmDAQ is not just prompt chaining. Every routing decision is grounded in math.

UCB1 Bandit Routing

UCB = μ + 1.4·√(ln N / nᵢ)

Exploration/exploitation balance

Bayesian Beta Reputation

β(α,β) → E[trust] = α/(α+β)

Shrinking uncertainty over time

Vickrey Auction

winner = argmax(utility), price = 2nd

Truth-telling bid mechanism

Elo + Bradley-Terry

P(A>B) = exp(Ra)/(exp(Ra)+exp(Rb))

Pairwise agent duel ratings

Markowitz Portfolio

max E[ret] - λ·var - μ·cost + syn

Swarm quality/risk/cost tradeoff

Shapley Contribution

φ(i) ≈ score(S) - score(S\{i})

Leave-one-out credit attribution

PageRank Trust Graph

PR(i) = (1-d)/N + d·Σ(PR(j)/out(j))

Collaboration hub centrality

The Key Insight

Markets over-correct before they calibrate.

Run 1: ResearchAgent makes an unsourced claim — penalized. Run 2: SourceVerifier paired in — factuality jumps 26 pts. Run 3: Market over-rotates, SkepticAgent displaces PitchAgent — score regresses. Run 4: Balanced swarm. New peak. The regression in run 3 is the point — real markets overshoot.

74
Run 1
factuality gap
→
91
Run 2
market learned
→
85
Run 3
⚠ over-rotation
→
96
Run 4
calibrated peak
Ask SwarmDAQ
Command the agent market. Run missions, inspect bids, rebalance swarms, replay traces, and expose weak agents.
Try: "Run a mission for RocketRide" or "Which agent is overvalued?"

Powered by CopilotKit