
You find out it's wrong from a customer
Without a score or a test suite, the first person to catch a bad answer is whoever acted on it.
Agents
Each agent carries a confidence score, an evaluation suite it has to keep passing, and 7 tuning surfaces — identity, knowledge, memory, rules, tables, verified queries and skills — under your control.
Revenue Analyst
Created by John Doe · 1,284 questions answered























Without a score or a test suite, the first person to catch a bad answer is whoever acted on it.

One text box, no measure of whether the edit made the answers better or quietly broke three of them.

You fix the same misread column every week, because the fix lives in a chat thread instead of the agent.
An agent is more than its instructions. Who it is, what it knows, what it remembers, the rules it obeys, the context on each table, the queries you have signed off and the skills it can load are each their own surface, so a fix lands where the problem is.

An agent reads the sources you gave it and writes back what it learned. Every step it takes is named as it happens, so the run is something you read rather than something you trust.

A suite of your real questions runs against the agent, each one scored and marked passed or failed, and the results roll up into one accuracy score. A regression shows up as a failed eval, not as a wrong answer in front of your team.

By the businesses running Supaboard in production.
Measured on a customer agent in production, after tuning.
Identity, knowledge, memory, rules, tables, verified queries, skills.
An AI analyst you configure and own: it has an identity, scoped access to your sources, and its own knowledge, rules and memory. Every analyst your team asks is one.
A per-agent score out of a hundred, computed from its evaluation results and open issues. It moves when you tune the agent, so you can see whether a change helped before anyone relies on it.
You collect real questions into a suite and run it against the agent. Each prompt gets a score and a pass or fail, and reruns track the trend, so a regression is caught by a failed eval instead of a user.
7 surfaces: who the agent is, its knowledge base, what it remembers, the rules it must follow, per-table context, the verified queries it reuses, and the skills it can load.
It keeps a memory of what it learned from past conversations, and with self-improve on it proposes its own fixes — but learnings stay reviewable, and verified queries only enter the pool when you sign them off.