> ## Documentation Index
> Fetch the complete documentation index at: https://docs.labelbox.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Write outcome rubrics

> Write a rubric an evaluator can grade, launch a session against it, and read the verdict for each criterion.

An outcome is a session's definition of done: an objective and a rubric. Each time the agent finishes its work, an evaluator agent checks it against every criterion in the rubric. It either accepts the work or sends the agent back with what fell short.

Use an outcome when "the agent stopped" isn't proof enough that the work is right. How well this works depends mostly on the rubric, so this page starts with how to write one.

<Note>
  Launching a graded session and saving an agent's default rubric need the Developer or Admin role in the organization. The User role can read outcomes. See [Organizations and roles](/recursion/organizations-and-roles). You need an [agent](/recursion/agents) to do the work, an [environment](/recursion/environments), and an agent to act as the evaluator, usually a separate one set up for grading.
</Note>

## Write a rubric

A rubric is Markdown. Each list item is one criterion. The evaluator owes a verdict on every list item and on nothing else.

| You write | How it's read |
| - | - |
| A list item (`-`, `*`, `+`, `1.`, or `1)`) | One criterion. A task box in front of the text, as in `- [ ] The file exists`, is dropped and the text is kept. Indented items are criteria too. |
| A heading (`#`, `##`, and so on) | The section for the criteria below it. The agent sees its criteria grouped by section, and each criterion in the result shows its section. |
| A paragraph | A note for people who read the rubric. It isn't scored, and neither the agent nor the evaluator receives it. |
| A fenced code block | Ignored, including list items inside it. |
| A horizontal rule (`---`) | Ignored. |

```markdown theme={"theme":"css-variables"}
# Release notes for v4.12.0

## Deliverable
- The notes are saved as a deliverable named `release-notes.md`.

## Coverage
- Every pull request merged between the tags v4.11.0 and v4.12.0 appears exactly once.
- No entry describes a pull request outside that range.

## Format
- Breaking changes are listed first, under a heading named "Breaking changes".
- Each entry is one sentence and ends with the pull request number in parentheses.
- No entry contains an internal tracker id.
```

This rubric has six criteria in three sections. Every criterion counts the same: you can't set weights, and a strong result on one criterion never makes up for a failure on another.

### Make each criterion stand on its own

The evaluator receives the objective and the criteria. It doesn't receive the rubric's paragraphs or headings, so each criterion has to carry everything needed to check it.

* **Make it checkable.** "The CSV has a numeric `price` column" can be checked. "The data looks good" can't, and a criterion like that sends the work back until the passes run out.
* **Give each list item one requirement.** When one item holds two requirements, a failure doesn't say which one the work missed.
* **Keep each criterion on one line.** Text on a second line isn't part of the criterion.
* **Say what counts as evidence in the criterion.** Write "The test suite passes when run with `make test`", not "Tests pass" followed by a paragraph about how to run them.
* **Ask for work that leaves evidence.** A file, a passing test, or a record in another system can be verified. A claim in the agent's last message can't.
* **Put shared background in the objective.** The opening message is the objective, and both the agent and the evaluator read it. Facts both need, such as which repository to use, belong there.
* **Start from a good example.** Take a result you would accept, list what makes it good, and turn each point into a criterion.

### Name the deliverable

Only files the agent saves as deliverables are kept with the session. Agents already know where to save them, so you don't need to give a path. Name the file in the opening message and in the rubric, for example "Save the notes as `release-notes.md`" and "- The notes are saved as a deliverable named `release-notes.md`." Deliverables are saved before every grading pass, so the evaluator checks the version the agent just finished. See [Deliverables and artifacts](/recursion/artifacts).

### Name your criteria

A list item written as `- key: text` gives the criterion a name:

```markdown theme={"theme":"css-variables"}
- breaking-first: Breaking changes are listed first, under a heading named "Breaking changes".
- one-sentence: Each entry is one sentence and ends with the pull request number in parentheses.
```

A key starts with a lowercase letter and uses lowercase letters and digits, with single hyphens or underscores between them. The evaluator reports its verdicts against these names. The agent and the result show the text after the colon without the name, so that text has to make sense by itself. Name every criterion or none: a rubric that mixes named and unnamed items is refused.

<Warning>
  A list item that starts with a lowercase word, a colon, and a space is read as a named criterion. `- output: saved as report.csv` gets the name `output`, and the agent is shown "saved as report.csv" without it. If your other items have no name, the rubric is refused. To keep an item unnamed, start it with a capital letter or reword it.
</Warning>

<Accordion title="Rules a rubric must pass">
  A rubric is refused when:

  * It has no list item, or more than 200.
  * A list item is empty, such as a lone `-` or `- [ ]`, or a named item has no text, such as `- todo:`.
  * A line starts with a dash and no space, such as `-The file exists`.
  * It mixes named and unnamed criteria, uses the same key twice, or has a key longer than 64 characters.
  * It is larger than 256 KiB, or one criterion is larger than 4 KiB.
</Accordion>

## Choose an evaluator

Any agent in your organization can be the evaluator. It grades with its own model and its own instructions, so you decide how strict it is.

* **Use a capable model.** The evaluator has to read the work, check it against each criterion, and explain every failure.
* **Put standing guidance in the evaluator's instructions.** How strict to be, what counts as evidence, and what to ignore (such as formatting or naming style) apply to every rubric. They belong on the evaluator agent, because rubric paragraphs don't reach it.
* **Keep task details out of it.** One evaluator can grade many agents. What a particular session must produce belongs in that session's objective and rubric.

To create one, see [Agents](/recursion/agents).

## Grade a session against a rubric

<Steps>
  <Step title="Open the launch form">
    In the sidebar, click **Sessions**, then click **Launch session**. In **Launch a session**, choose the **Agent** and **Environment**.
  </Step>

  <Step title="Turn on grading">
    Under **Task**, select **Grade this session against a rubric**. In **Grade with**, choose the evaluator.
  </Step>

  <Step title="Write the rubric">
    Write your criteria in **Rubric**. The field starts with the agent's default rubric, or with a short template. Replace the template's lines with your own. The paragraphs under its `## Evidence` and `## Ignore` headings are notes, not criteria.
  </Step>

  <Step title="Write the objective and launch">
    Write the task in **Opening message**. It's the objective the evaluator grades against. Click **Launch session**.
  </Step>
</Steps>

Grading is on from the first turn. Each grading pass is a session of the evaluator agent, so it's charged like any other session. See [Pricing](/recursion/pricing).

## Grade every session of an agent

Store a rubric and an evaluator on the agent to grade the sessions launched from it.

1. Open the agent and stay on the **Configuration** tab.
2. In **Outcome**, click **Add** if the section is closed.
3. Write the rubric in the text box and choose the **Default evaluator**.
4. Click **Save new version**.

In **Launch a session**, an agent with both has **Grade this session against a rubric** already selected and shows the stored rubric. Clear the box to run that session ungraded. An edit to the rubric in the form applies to that session only.

A stored rubric with no default evaluator doesn't grade anything by itself. The launch form offers the rubric, and you choose the evaluator for that run.

## Read the result

1. In the sidebar, click **Sessions** and open the session.
2. Read the **Outcome** line above the transcript: the result so far, the number of unmet criteria if any, and the objective.
3. Click **Details** to see the objective, the pass limit, the **Rubric**, and one block per pass, newest first. Each pass shows how many criteria were met and every criterion with its verdict. A failed criterion also shows the evaluator's reason.
4. Click **Open evaluation and evidence** on a pass to read the evaluator's own session, including what it checked.

| The outcome shows | When | What happens next |
| - | - | - |
| **Pending** | No grading pass has run yet. | The first pass starts when the agent finishes its work. |
| **Evaluating** | A grading pass is running. | Wait for the verdict. The session's status also reads **Evaluating**. |
| **Running** | The agent is revising after a failed pass, or a pass returned no result. | The next pass starts when the agent finishes. |
| **Satisfied** | No criterion failed and at least one passed. | Grading ends with the work accepted. |
| **Max iterations reached** | A criterion still failed on the last pass. | Grading ends. The last pass keeps its verdicts. |
| **Not applicable** | The evaluator judged every criterion not applicable. | Grading ends. Check that the rubric fits the task. |
| **Failed** | The work couldn't be graded. | Grading ends. Open **Details** to read the reason. |

After a final result, the session still accepts messages, but those turns aren't graded. To grade more work, launch a new session.

## How grading works

```mermaid theme={"theme":"css-variables"}
flowchart TD
  work["Agent works"] --> done["Turn ends"]
  done --> grade["Evaluator grades"]
  grade -->|"a criterion failed"| revise["Agent revises"]
  revise --> done
  grade -->|"none failed"| accepted["Outcome satisfied"]
```

1. The agent works until it ends a turn with nothing left to do. A turn you interrupt or cancel isn't graded.
2. The agent's deliverables are saved, and a **grading pass** starts. The evaluator runs in a separate session on a copy of the session's work, with the same environment, tools, and credentials. It can open files, run the project's tests, and check systems the agent changed. It can read the transcript, but not the agent's reasoning.
3. The evaluator gives every criterion one verdict: **Pass**, **Fail**, or **Not applicable**. One **Fail** is enough to send the work back.
4. The agent receives the whole scorecard, with the evaluator's reason for each failed criterion, and revises. Then the next pass starts.

For a session launched from the console, grading stops after five passes, the first one included. The outcome is **Satisfied** when no criterion failed and at least one passed.

## What can go wrong

The most common problems:

* **The work keeps coming back, then the outcome reads Max iterations reached.** A criterion is vague or can't be met. Read the evaluator's reason on the failed criterion, rewrite it so it can be checked, and launch again.
* **A deliverable criterion fails.** The agent saved the file under another name, or didn't save it as a deliverable. Name the file in both the opening message and the rubric.
* **A session wasn't graded.** The box was cleared in the launch form, or the agent has a stored rubric but no default evaluator.
* **The transcript says the outcome check returned no result.** The evaluator didn't return a usable verdict, so the work wasn't judged. Send a message to continue. Grading runs again when the agent next finishes.

## Limits

| Limit | Value |
| - | - |
| Criteria in one rubric | 1 to 200 |
| Rubric size | 256 KiB |
| One criterion's text | 4 KiB |
| Grading passes for one outcome | 5 for a session launched from the console, the first one included |
| Criterion weights | None. Every criterion counts the same. |

Other limits are on [Limits](/recursion/limits).

## Next steps

<CardGroup cols={2}>
  <Card title="Deliverables and artifacts" href="/recursion/artifacts">
    Ask for named deliverables the evaluator can check.
  </Card>

  <Card title="Sessions" href="/recursion/sessions">
    Launch, steer, and follow a session while it's graded.
  </Card>

  <Card title="Agents" href="/recursion/agents">
    Create an evaluator, or store a default rubric on an agent.
  </Card>

  <Card title="Pricing" href="/recursion/pricing">
    See what a session hour covers.
  </Card>
</CardGroup>
