> ## Documentation Index
> Fetch the complete documentation index at: https://docs.genai.scale.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Rubrics

> Contributor defined rating criteria

Rubrics tasks are a specific type of Evaluation task where models are assessed based on a dynamic rubric, which is defined by the contributor during the task. These tasks provide quantitative and qualitative insight into how well a model can follow arbitrary criteria.

## Overview

Each rubrics task is designed to test specific aspects of model performance, from basic comprehension to complex reasoning abilities, across natural human defined criteria.

## Task Structure

A Rubrics task consists of:

* A prompt (question or instruction)
* A list of model evaluation criteria (contributor defined based on the prompt)
* Multiple model generated responses
* Human evaluation of the model responses against the criteria

<Note>
  Additional annotations can be included when necessary for standard evaluation
  of the responses.
</Note>

<img className="block dark:hidden" src="https://mintcdn.com/data-engine-gen-ai/7rAP6-o6Oxbv3RYk/images/archetype-rubrics-light.svg?fit=max&auto=format&n=7rAP6-o6Oxbv3RYk&q=85&s=c4b38b89c0d06a97ca47c44571cfd3a0" width="720" height="1008" data-path="images/archetype-rubrics-light.svg" />

<img className="hidden dark:block" src="https://mintcdn.com/data-engine-gen-ai/7rAP6-o6Oxbv3RYk/images/archetype-rubrics-dark.svg?fit=max&auto=format&n=7rAP6-o6Oxbv3RYk&q=85&s=9594bbb3eb9db2be87716789693cf821" width="720" height="1008" data-path="images/archetype-rubrics-dark.svg" />

### Messages

Each turn consists of sequential [messages](../core-resources/message) that represent a user prompt, multiple model responses, and corresponding human evaluation of the model responses against the criteria.

#### User message

Contains the initial prompt or instruction (role: `user`).

<CodeGroup>
  ```json Sample Instruction theme={null}
  {
    "content": {
      "text": "Put the following data into a table and sort by name:\n\nSue,$20\nJared,$40\nMike,$60"
    },
    "role": "user",
    "source_id": "user",
    "annotations": []
  }
  ```

  ```python get_rubric_user_prompts.py theme={null}
  # task: output of `/v2/task`
  # returns a map of turn_id to user prompt for every turn
  def get_rubric_user_prompts(task):
  	thread = task['threads'][0]
  	user_prompts = {}
  	for turn in thread['turns']:
  		responses = [msg for msg in turn['messages'] if msg['source_id'].lower() == "user"]
  		assert len(responses) == 1, "Turn must contain a user prompt. Turn ID: {}".format(turn['id'])
  		user_prompts[turn['id']] = responses[0]
  	return user_prompts
  ```
</CodeGroup>

#### Model Responses

Contains responses from different models (role: `assistant`).

<CodeGroup>
  ```json Sample Model Response theme={null}
  {
    "content": {
      "text": "This is model 1 response"
    },
    "role": "assistant",
    "source_id": "model_1",
    "annotations": [
      {
        "id": "rubric_0_criteria_0_rating",
        "title": "The model must respond with a formatted table.",
        "value": "no_issues",
        "metadata": { "criteria": "rubric_0_criteria_0" }
      },
      {
        "id": "rubric_0_criteria_1_rating",
        "title": "The response must sort the names in alphabetical order.",
        "value": "major_issues",
        "metadata": { "criteria": "rubric_0_criteria_1" }
      },
      {
        "id": "rubric_0_criteria_2_rating",
        "title": "The response must include bold headers.",
        "value": "minor_issues",
        "metadata": { "criteria": "rubric_0_criteria_2" }
      }
    ]
  }
  ```

  ```python get_rubrics_model_responses.py theme={null}
  # task: output of `/v2/task`
  # returns a map of turn_id to model responses for every turn
  def get_rubrics_model_responses(task):
  	thread = task['threads'][0]
  	model_responses = {}
  	for turn in thread['turns']:
  		responses = [msg for msg in turn['messages'] if msg['role'].lower() == "assistant"]
  		assert len(responses) > 0, "Turn must contain at least one model response. Turn ID: {}".format(turn['id'])
  		model_responses[turn['id']] = responses
  	return model_responses
  ```
</CodeGroup>

<Note>
  Each model response includes a `source_id` that uniquely identifies the model
  that generated the response.
</Note>

#### Message Annotations

Each model response is evaluated across the contributor-defined criteria. For example:

* "The response must display information in a table."
* "The response must sort the names in alphabetical order."
* "The response must use metric units."

Message's [`annotations`](../core-resources/annotation) include the evaluation results for each dimension.

<CodeGroup>
  ```json Sample Message Annotations theme={null}
  [
    {
      "id": "rubric_0_criteria_0_rating",
      "title": "The model must respond with a formatted table.",
      "value": "no_issues",
      "metadata": {
        "criteria": "rubric_0_criteria_0" // rubric criteria defined at turn-level annotation
      }
    }
  ]
  ```
</CodeGroup>

<Note>
  The rubric criteria is unique to the prompt and is different per task.
  Additional evaluation dimensions can be included based on project requirements
  and evaluation objectives.
</Note>

### Turn-Level Annotations

The `annotations` at the **turn** level, specifies contributor-defined rubrics, preference related or aggregated information. Some common examples are:

* Selected model identifier: `selected_model_id`
* Likert scale rating: `likert_value`
* Detailed justification for the selection: `justification`
* Any other comparative analysis between model responses

<CodeGroup>
  ```json Sample Turn-Level Annotations theme={null}
  [
    {
      "key": "selected_model_id",
      "value": "model_2"
    },
    {
      "key": "rubric_0_criteria_0",
      "title": "The model must respond with a formatted table.",
      "value": "objective",
    },
    {
      "key": "rubric_0_criteria_1",
      "title": "The response must sort the names in alphabetical order.",
      "value": "objective",
    },
    {
      "key": "rubric_0_criteria_2",
      "title": "The response must include bold headers.",
      "value": "implicit",
    }
  ]
  ```

  ```python get_turn_annotations.py theme={null}
  # task: output of `/v2/task`
  # returns a map of turn_id to turn-level annotations for every turn
  def get_turn_annotations(task):
  	thread = task['threads'][0]
  	turn_annotations = {}
  	for turn in thread['turns']:
  		turn_annotations[turn['id']] = turn['annotations']
  	return turn_annotations
  ```
</CodeGroup>

### Expanded Rubrics Task Output

This is a sample expanded sample Rubrics Task output returned by [`/v2/task`](../v2/task).

```json Sample Rubrics Task Output theme={null}
{
  "task_id": "task_123",
  "project": "project_123",
  "batch": "batch_123",
  "status": "completed",
  "created_at": "2025-01-01T08:31:03.169Z",
  "completed_at": "2025-01-02T04:00:39.923Z",
  "threads": [
    {
      "id": "thread_0",
      "turns": [
        {
          "id": "turn_0",
          "messages": [
            {
              "content": {
                "text": "Put the following data into a table and sort by name:\n\nSue,$20\nJared,$40\nMike,$60"
              },
              "role": "user",
              "source_id": "user",
              "annotations": []
            },
            {
              "content": {
                "text": "This is model 1 response"
              },
              "role": "assistant",
              "source_id": "model_1",
              "annotations": [
                {
                  "id": "rubric_0_criteria_0_rating",
                  "title": "The model must respond with a formatted table.",
                  "value": "no_issues",
                  "metadata": { "criteria": "rubric_0_criteria_0" }
                },
                {
                  "id": "rubric_0_criteria_1_rating",
                  "title": "The response must sort the names in alphabetical order.",
                  "value": "major_issues",
                  "metadata": { "criteria": "rubric_0_criteria_1" }
                },
                {
                  "id": "rubric_0_criteria_2_rating",
                  "title": "The response must include bold headers.",
                  "value": "minor_issues",
                  "metadata": { "criteria": "rubric_0_criteria_2" }
                }
              ]
            },
            {
              "content": {
                "text": "This is model 2 response"
              },
              "role": "assistant",
              "source_id": "model_2",
              "annotations": [
                {
                  "id": "rubric_0_criteria_0_rating",
                  "title": "The model must respond with a formatted table.",
                  "value": "no_issues",
                  "metadata": { "criteria": "rubric_0_criteria_0" }
                },
                {
                  "id": "rubric_0_criteria_1_rating",
                  "title": "The response must sort the names in alphabetical order.",
                  "value": "no_issues",
                  "metadata": { "criteria": "rubric_0_criteria_1" }
                },
                {
                  "id": "rubric_0_criteria_2_rating",
                  "title": "The response must include bold headers.",
                  "value": "major_issues",
                  "metadata": { "criteria": "rubric_0_criteria_2" }
                }
              ]
            }
          ],
          "annotations": [
            {
              "key": "selected_model_id",
              "value": "model_2"
            },
            {
              "key": "rubric_0_criteria_0",
              "title": "The model must respond with a formatted table.",
              "value": "objective",
            },
            {
              "key": "rubric_0_criteria_1",
              "title": "The response must sort the names in alphabetical order.",
              "value": "objective",
            },
            {
              "key": "rubric_0_criteria_2",
              "title": "The response must include bold headers.",
              "value": "implicit",
            }
          ]
        }
      ],
      "annotations": []
    }
  ]
}
```
