Skip to content
Supaboard

Agents

An analyst you can hold accountable

Each agent carries a confidence score, an evaluation suite it has to keep passing, and 7 tuning surfaces — identity, knowledge, memory, rules, tables, verified queries and skills — under your control.

Revenue Analyst

Created by John Doe · 1,284 questions answered

Confidence Score

87 /100

GoodLast evaluated 2h ago · 3 issues
Evaluation E3
71/100
Evaluation E4
74/100
Evaluation E5
72/100
Evaluation E6
78/100
Evaluation E7
81/100
Evaluation E8
79/100
Evaluation E9
83/100
Evaluation E10
85/100
Evaluation E11
84/100
Evaluation E12
87/100

+16 points over 10 evals

Issues

2 blocking
  • 2 tables in orders_db are missing descriptions (blocking)
  • Join path orders → refunds is unverified (blocking)
  • Memory has 3 unreviewed learnings

Quick Actions

Self Improve
AI Insights

Fine Tuning Setup

  • Who am I

    Setup your agent's identity, access and Integrations.

  • Knowledge Base

    The brain of your agent—your business tribal knowledge.

  • Memory

    What your Agent has learned from past conversation.

  • Rules

    Guardrails the agent follows for every single query.

  • Needs attention:

    Table Configuration

    Add context to your data sources so query lands right.

  • Needs attention:

    Verified Queries

    Answers you have signed off on accuracy compounds here.

  • Skills

    Reusable procedures your agent loads on demand.

Evaluation

E12 (Latest)
PromptsStatusScoreEval TrendCreated by
What was net revenue last quarter, by brand?Passed94/100+3John Doe
Which return reason grew fastest this month?Passed92/100+4Amara Chen
Refund rate for repeat customers vs first ordersFailed58/100-12John Doe
Top 5 states by average order value, trailing 90 daysPassed90/100+1Priya Nair
How did the March promo change basket size?Passed88/100+2John Doe

Trusted by 1,000+ businesses across the world

  • Jindal Healthcare
  • Shopee
  • NetEase Games
  • Holcim
  • Publicis Sapient
  • Nilfisk
  • ClickBus
  • AAA
  • Scopely markScopely
  • FunPlus
  • Bilt
  • GlobalSign
  • Katalon
  • StockGro
  • Yami
  • Immutable
  • Plane
  • Gabriella
  • Visily
  • Poka
  • Trafilea

An AI you can't test is an AI you can't trust

A briefcase labeled data insights locked away from a business stakeholder

You find out it's wrong from a customer

Without a score or a test suite, the first person to catch a bad answer is whoever acted on it.

Mismatched abstract shapes that refuse to fit one visual style

Tuning is a prompt someone edits blind

One text box, no measure of whether the edit made the answers better or quietly broke three of them.

Stacked database discs frozen mid-refresh

It never learns what you corrected

You fix the same misread column every week, because the fix lives in a chat thread instead of the agent.

Tune it, watch it work, score it

7 dials, not one prompt

An agent is more than its instructions. Who it is, what it knows, what it remembers, the rules it obeys, the context on each table, the queries you have signed off and the skills it can load are each their own surface, so a fix lands where the problem is.

  • Identity, knowledge base and memory
  • Rules and per-table configuration
  • Verified queries and reusable skills
The agent tuning surfaces as separate cards: agent's role, memory, table configuration, verified queries, knowledge base and rules

Watch it do the work

An agent reads the sources you gave it and writes back what it learned. Every step it takes is named as it happens, so the run is something you read rather than something you trust.

  • Each step named while it runs
  • The documents it opened, listed
  • What it learned written into its knowledge base
An agent working through a task step by step: searching a document store, reading three quarterly OKR files, then updating its knowledge base

Evaluations that keep it honest

A suite of your real questions runs against the agent, each one scored and marked passed or failed, and the results roll up into one accuracy score. A regression shows up as a failed eval, not as a wrong answer in front of your team.

  • Your own prompts as the test suite
  • Pass or fail, with a score per prompt
  • One accuracy score across the suite
An agent evaluation listing churn rate, customer acquisition cost and MRR prompts each passed with its own accuracy, above an overall agent accuracy score of 96.7%

What a tuned agent looks like

  • 2K+ Agents trained

    By the businesses running Supaboard in production.

  • 97.8% Accuracy a tuned agent holds

    Measured on a customer agent in production, after tuning.

  • 7 Surfaces you can tune

    Identity, knowledge, memory, rules, tables, verified queries, skills.

Frequently asked questions

What is an agent in Supaboard?

An AI analyst you configure and own: it has an identity, scoped access to your sources, and its own knowledge, rules and memory. Every analyst your team asks is one.

What is the confidence score?

A per-agent score out of a hundred, computed from its evaluation results and open issues. It moves when you tune the agent, so you can see whether a change helped before anyone relies on it.

How do evaluations work?

You collect real questions into a suite and run it against the agent. Each prompt gets a score and a pass or fail, and reruns track the trend, so a regression is caught by a failed eval instead of a user.

What can I actually tune?

7 surfaces: who the agent is, its knowledge base, what it remembers, the rules it must follow, per-table context, the verified queries it reuses, and the skills it can load.

Does an agent learn on its own?

It keeps a memory of what it learned from past conversations, and with self-improve on it proposes its own fixes — but learnings stay reviewable, and verified queries only enter the pool when you sign them off.

Your data has the answers

Start free trialBook a demo