STRATA-Bench: a benchmark for AI agents on fragmented spatial-temporal market intelligence
STRATA-Bench documentation
Spatial-Temporal Retrieval, Alignment, Transparency & Audit — a benchmark for AI agents that must compile longitudinal market intelligence from fragmented, geographically nested, and temporally incomplete evidence.
- Protocol — the seven rules every submission must follow
- Scoring — dimensions, weights, and hard fails
- Tasks — the 36 sandbox task families
- 🔍 Task Explorer — browse every task, filter by failure family
- The Three Agent Archetypes — the story behind the benchmark
- Prompting — how to brief an agent without tripping the traps
- Paper notes — background on the motivating audit
Start with the README for installation and the quickstart.
STRATA-Bench is built by Mohammad Movahedi — Data Privacy & AI Governance Consultant, Toronto.