Launching a graded session and saving an agent’s default rubric need the Developer or Admin role in the organization. The User role can read outcomes. See Organizations and roles. You need an agent to do the work, an environment, and an agent to act as the evaluator, usually a separate one set up for grading.
Write a rubric
A rubric is Markdown. Each list item is one criterion. The evaluator owes a verdict on every list item and on nothing else.Make each criterion stand on its own
The evaluator receives the objective and the criteria. It doesn’t receive the rubric’s paragraphs or headings, so each criterion has to carry everything needed to check it.- Make it checkable. “The CSV has a numeric
pricecolumn” can be checked. “The data looks good” can’t, and a criterion like that sends the work back until the passes run out. - Give each list item one requirement. When one item holds two requirements, a failure doesn’t say which one the work missed.
- Keep each criterion on one line. Text on a second line isn’t part of the criterion.
- Say what counts as evidence in the criterion. Write “The test suite passes when run with
make test”, not “Tests pass” followed by a paragraph about how to run them. - Ask for work that leaves evidence. A file, a passing test, or a record in another system can be verified. A claim in the agent’s last message can’t.
- Put shared background in the objective. The opening message is the objective, and both the agent and the evaluator read it. Facts both need, such as which repository to use, belong there.
- Start from a good example. Take a result you would accept, list what makes it good, and turn each point into a criterion.
Name the deliverable
Only files the agent saves as deliverables are kept with the session. Agents already know where to save them, so you don’t need to give a path. Name the file in the opening message and in the rubric, for example “Save the notes asrelease-notes.md” and ”- The notes are saved as a deliverable named release-notes.md.” Deliverables are saved before every grading pass, so the evaluator checks the version the agent just finished. See Deliverables and artifacts.
Name your criteria
A list item written as- key: text gives the criterion a name:
Rules a rubric must pass
Rules a rubric must pass
A rubric is refused when:
- It has no list item, or more than 200.
- A list item is empty, such as a lone
-or- [ ], or a named item has no text, such as- todo:. - A line starts with a dash and no space, such as
-The file exists. - It mixes named and unnamed criteria, uses the same key twice, or has a key longer than 64 characters.
- It is larger than 256 KiB, or one criterion is larger than 4 KiB.
Choose an evaluator
Any agent in your organization can be the evaluator. It grades with its own model and its own instructions, so you decide how strict it is.- Use a capable model. The evaluator has to read the work, check it against each criterion, and explain every failure.
- Put standing guidance in the evaluator’s instructions. How strict to be, what counts as evidence, and what to ignore (such as formatting or naming style) apply to every rubric. They belong on the evaluator agent, because rubric paragraphs don’t reach it.
- Keep task details out of it. One evaluator can grade many agents. What a particular session must produce belongs in that session’s objective and rubric.
Grade a session against a rubric
1
Open the launch form
In the sidebar, click Sessions, then click Launch session. In Launch a session, choose the Agent and Environment.
2
Turn on grading
Under Task, select Grade this session against a rubric. In Grade with, choose the evaluator.
3
Write the rubric
Write your criteria in Rubric. The field starts with the agent’s default rubric, or with a short template. Replace the template’s lines with your own. The paragraphs under its
## Evidence and ## Ignore headings are notes, not criteria.4
Write the objective and launch
Write the task in Opening message. It’s the objective the evaluator grades against. Click Launch session.
Grade every session of an agent
Store a rubric and an evaluator on the agent to grade the sessions launched from it.- Open the agent and stay on the Configuration tab.
- In Outcome, click Add if the section is closed.
- Write the rubric in the text box and choose the Default evaluator.
- Click Save new version.
Read the result
- In the sidebar, click Sessions and open the session.
- Read the Outcome line above the transcript: the result so far, the number of unmet criteria if any, and the objective.
- Click Details to see the objective, the pass limit, the Rubric, and one block per pass, newest first. Each pass shows how many criteria were met and every criterion with its verdict. A failed criterion also shows the evaluator’s reason.
- Click Open evaluation and evidence on a pass to read the evaluator’s own session, including what it checked.
After a final result, the session still accepts messages, but those turns aren’t graded. To grade more work, launch a new session.
How grading works
- The agent works until it ends a turn with nothing left to do. A turn you interrupt or cancel isn’t graded.
- The agent’s deliverables are saved, and a grading pass starts. The evaluator runs in a separate session on a copy of the session’s work, with the same environment, tools, and credentials. It can open files, run the project’s tests, and check systems the agent changed. It can read the transcript, but not the agent’s reasoning.
- The evaluator gives every criterion one verdict: Pass, Fail, or Not applicable. One Fail is enough to send the work back.
- The agent receives the whole scorecard, with the evaluator’s reason for each failed criterion, and revises. Then the next pass starts.
What can go wrong
The most common problems:- The work keeps coming back, then the outcome reads Max iterations reached. A criterion is vague or can’t be met. Read the evaluator’s reason on the failed criterion, rewrite it so it can be checked, and launch again.
- A deliverable criterion fails. The agent saved the file under another name, or didn’t save it as a deliverable. Name the file in both the opening message and the rubric.
- A session wasn’t graded. The box was cleared in the launch form, or the agent has a stored rubric but no default evaluator.
- The transcript says the outcome check returned no result. The evaluator didn’t return a usable verdict, so the work wasn’t judged. Send a message to continue. Grading runs again when the agent next finishes.
Limits
Other limits are on Limits.
Next steps
Deliverables and artifacts
Ask for named deliverables the evaluator can check.
Sessions
Launch, steer, and follow a session while it’s graded.
Agents
Create an evaluator, or store a default rubric on an agent.
Pricing
See what a session hour covers.