Skip to main content
An outcome is a session’s definition of done: an objective and a rubric. Each time the agent finishes its work, an evaluator agent checks it against every criterion in the rubric. It either accepts the work or sends the agent back with what fell short. Use an outcome when “the agent stopped” isn’t proof enough that the work is right. How well this works depends mostly on the rubric, so this page starts with how to write one.
Launching a graded session and saving an agent’s default rubric need the Developer or Admin role in the organization. The User role can read outcomes. See Organizations and roles. You need an agent to do the work, an environment, and an agent to act as the evaluator, usually a separate one set up for grading.

Write a rubric

A rubric is Markdown. Each list item is one criterion. The evaluator owes a verdict on every list item and on nothing else.
This rubric has six criteria in three sections. Every criterion counts the same: you can’t set weights, and a strong result on one criterion never makes up for a failure on another.

Make each criterion stand on its own

The evaluator receives the objective and the criteria. It doesn’t receive the rubric’s paragraphs or headings, so each criterion has to carry everything needed to check it.
  • Make it checkable. “The CSV has a numeric price column” can be checked. “The data looks good” can’t, and a criterion like that sends the work back until the passes run out.
  • Give each list item one requirement. When one item holds two requirements, a failure doesn’t say which one the work missed.
  • Keep each criterion on one line. Text on a second line isn’t part of the criterion.
  • Say what counts as evidence in the criterion. Write “The test suite passes when run with make test”, not “Tests pass” followed by a paragraph about how to run them.
  • Ask for work that leaves evidence. A file, a passing test, or a record in another system can be verified. A claim in the agent’s last message can’t.
  • Put shared background in the objective. The opening message is the objective, and both the agent and the evaluator read it. Facts both need, such as which repository to use, belong there.
  • Start from a good example. Take a result you would accept, list what makes it good, and turn each point into a criterion.

Name the deliverable

Only files the agent saves as deliverables are kept with the session. Agents already know where to save them, so you don’t need to give a path. Name the file in the opening message and in the rubric, for example “Save the notes as release-notes.md” and ”- The notes are saved as a deliverable named release-notes.md.” Deliverables are saved before every grading pass, so the evaluator checks the version the agent just finished. See Deliverables and artifacts.

Name your criteria

A list item written as - key: text gives the criterion a name:
A key starts with a lowercase letter and uses lowercase letters and digits, with single hyphens or underscores between them. The evaluator reports its verdicts against these names. The agent and the result show the text after the colon without the name, so that text has to make sense by itself. Name every criterion or none: a rubric that mixes named and unnamed items is refused.
A list item that starts with a lowercase word, a colon, and a space is read as a named criterion. - output: saved as report.csv gets the name output, and the agent is shown “saved as report.csv” without it. If your other items have no name, the rubric is refused. To keep an item unnamed, start it with a capital letter or reword it.
A rubric is refused when:
  • It has no list item, or more than 200.
  • A list item is empty, such as a lone - or - [ ], or a named item has no text, such as - todo:.
  • A line starts with a dash and no space, such as -The file exists.
  • It mixes named and unnamed criteria, uses the same key twice, or has a key longer than 64 characters.
  • It is larger than 256 KiB, or one criterion is larger than 4 KiB.

Choose an evaluator

Any agent in your organization can be the evaluator. It grades with its own model and its own instructions, so you decide how strict it is.
  • Use a capable model. The evaluator has to read the work, check it against each criterion, and explain every failure.
  • Put standing guidance in the evaluator’s instructions. How strict to be, what counts as evidence, and what to ignore (such as formatting or naming style) apply to every rubric. They belong on the evaluator agent, because rubric paragraphs don’t reach it.
  • Keep task details out of it. One evaluator can grade many agents. What a particular session must produce belongs in that session’s objective and rubric.
To create one, see Agents.

Grade a session against a rubric

1

Open the launch form

In the sidebar, click Sessions, then click Launch session. In Launch a session, choose the Agent and Environment.
2

Turn on grading

Under Task, select Grade this session against a rubric. In Grade with, choose the evaluator.
3

Write the rubric

Write your criteria in Rubric. The field starts with the agent’s default rubric, or with a short template. Replace the template’s lines with your own. The paragraphs under its ## Evidence and ## Ignore headings are notes, not criteria.
4

Write the objective and launch

Write the task in Opening message. It’s the objective the evaluator grades against. Click Launch session.
Grading is on from the first turn. Each grading pass is a session of the evaluator agent, so it’s charged like any other session. See Pricing.

Grade every session of an agent

Store a rubric and an evaluator on the agent to grade the sessions launched from it.
  1. Open the agent and stay on the Configuration tab.
  2. In Outcome, click Add if the section is closed.
  3. Write the rubric in the text box and choose the Default evaluator.
  4. Click Save new version.
In Launch a session, an agent with both has Grade this session against a rubric already selected and shows the stored rubric. Clear the box to run that session ungraded. An edit to the rubric in the form applies to that session only. A stored rubric with no default evaluator doesn’t grade anything by itself. The launch form offers the rubric, and you choose the evaluator for that run.

Read the result

  1. In the sidebar, click Sessions and open the session.
  2. Read the Outcome line above the transcript: the result so far, the number of unmet criteria if any, and the objective.
  3. Click Details to see the objective, the pass limit, the Rubric, and one block per pass, newest first. Each pass shows how many criteria were met and every criterion with its verdict. A failed criterion also shows the evaluator’s reason.
  4. Click Open evaluation and evidence on a pass to read the evaluator’s own session, including what it checked.
After a final result, the session still accepts messages, but those turns aren’t graded. To grade more work, launch a new session.

How grading works

  1. The agent works until it ends a turn with nothing left to do. A turn you interrupt or cancel isn’t graded.
  2. The agent’s deliverables are saved, and a grading pass starts. The evaluator runs in a separate session on a copy of the session’s work, with the same environment, tools, and credentials. It can open files, run the project’s tests, and check systems the agent changed. It can read the transcript, but not the agent’s reasoning.
  3. The evaluator gives every criterion one verdict: Pass, Fail, or Not applicable. One Fail is enough to send the work back.
  4. The agent receives the whole scorecard, with the evaluator’s reason for each failed criterion, and revises. Then the next pass starts.
For a session launched from the console, grading stops after five passes, the first one included. The outcome is Satisfied when no criterion failed and at least one passed.

What can go wrong

The most common problems:
  • The work keeps coming back, then the outcome reads Max iterations reached. A criterion is vague or can’t be met. Read the evaluator’s reason on the failed criterion, rewrite it so it can be checked, and launch again.
  • A deliverable criterion fails. The agent saved the file under another name, or didn’t save it as a deliverable. Name the file in both the opening message and the rubric.
  • A session wasn’t graded. The box was cleared in the launch form, or the agent has a stored rubric but no default evaluator.
  • The transcript says the outcome check returned no result. The evaluator didn’t return a usable verdict, so the work wasn’t judged. Send a message to continue. Grading runs again when the agent next finishes.

Limits

Other limits are on Limits.

Next steps

Deliverables and artifacts

Ask for named deliverables the evaluator can check.

Sessions

Launch, steer, and follow a session while it’s graded.

Agents

Create an evaluator, or store a default rubric on an agent.

Pricing

See what a session hour covers.