Skip to main content
You only need a local directory containing SKILL.md to evaluate a Skill using comet eval. Even if you haven’t pre-written the evaluation use cases, Eval will automatically generate 2 to 4 evaluation use cases based on the content of the Skill and run them. This article introduces in the order of actual usage: First, run it through once, then select the Agent and evaluate the backend, and then configure the task, generate the report and locate the cause of failure.

What you need before running

Required content

  • @rpamis/comet has been installed. The npm package comes with an eval harness, so there is no need to clone the Comet repository.
  • A local Skill directory containing SKILL.md.
  • When running the real evaluation, uv, Python 3.11+, Docker, the selected Agent CLI and the corresponding model credentials are required.
--collect only performs static discovery and configuration checks, and does not start agents, Docker, plugins, credentials, or network requests. Therefore, it can be used to check the Skill first, without having to prepare a complete model environment from the very beginning. If you want to select an Agent other than Claude Code, first read the Eval Agent startup configuration to confirm the corresponding CLI, authentication method and model protocol.

User-level .env configuration

Ordinary users do not need to clone the Comet source code, nor do they need to enter eval/ in the source code directory. Run arbitrary for the first time When the comet eval command is used, the CLI will automatically create a complete user-level configuration file:
  • Windows:%USERPROFILE%\.comet\eval\.env
  • macOS/Linux:~/.comet/eval/.env
The CLI only creates missing files and does not overwrite existing ones. The first run will display the actual path in the output. Open this file Just fill in the parameters as needed and run it again. The environment variables already set in the current shell take precedence over .env. The automatically generated template contains all user-configurable parameters and is organized in the following groups: The parameters in the template are all comments by default. Items that are not filled in will not change their default behavior. The real API key should only be placed at the user level .env or the current shell cannot write skills, manifest, reports, or public repositories.

Which evaluation agents are supported

By default, comet eval uses claude-code. The main Agent, user simulator and optional Judge can all use the following agents:
The custom Agent adapter is a user-level advanced capability and does not require cloning the Comet source code. Only the built-in Eval harness needs to be modified. When it comes to tasks or Docker environments, it is only necessary to refer to Advanced Configuration ].
For registration, credentials, and CLI conventions for custom Agents, see Eval Agent setup · Extend Eval with a custom Agent. Placing the executable file on PATH alone does not automatically enable it. If evaluation ends instantly and does not actually run, Docker, Agent CLI, or model credentials are often not ready. In these cases, harness marks the sample as skipped to avoid false-positive success.

The first run

First, perform a pre-check without calling the model, then run the low-cost smoke once, and finally run the complete task set as needed
When there is no comet/eval.yaml, or when there are no evaluation.tasks and recommendedTasks in the manifest, the normal run will automatically generate 2 to 4 evaluation use cases based on the Skill content, freeze and cache them before execution. You don’t need to write down the task first to start the assessment. --quick does not use these automatically generated use cases but runs generic-skill-smoke consistently to verify that the Skill can be injected, invoked, and produce result.md.

How to pass “target”

Directories, direct SKILL.md and manifest can all be used as targets:
When there is no manifest, Comet will synthesize the basic configuration in memory and will not rewrite the Skill source file. When the Skill is outside the repository, use --project to specify the project directory for saving the running status and reports:

Configure the evaluation Agent and Judge

The simplest way is to select the main Agent in the command line:
It is also possible to create comet/eval.yaml in the root directory of the evaluated Skill and configure the default Agent in this file. For example, the directory structure is as follows:
When you run comet eval ./my-skill, Comet will automatically detect ./my-skill/comet/eval.yaml. You can also directly pass this file as the target
Configure the default Agent in eval.yaml. CLI options take precedence over manifest:
The main Agent and LLM-as-Judge can use different agents, models, API addresses and credentials. Judge does not silently inherit the model or credentials of the main Agent. When enabling Judge, a separate Judge configuration needs to be provided.

Select the evaluation backend

--suite selects the evaluation backend. The three backends use the same set of target, task and Agent configurations. The difference lies in whether the results are synchronized to the external evaluation service: Langfuse example
If it is only for checking tasks and configurations, any backend can be used together with --collect:
--collect --suite langfuse does not initialize the SDK, does not connect to the Internet, nor does it download plugins.

What types of reports are there

The assessment status and reports are written by default to the following directory in the project the target belongs to:
When --html is not included, the assessment will still generate summary.md. Run when HTML is needed:
You can also customize the report output using --report-config <path> or COMET_EVAL_REPORT_CONFIG. The CLI will print Experiment and Report path. When searching for reports, the output of this time shall prevail. When reading the report, pay attention to three things first:
  1. Evaluate whether it is passed.
  2. Is the failure attribution harness, workflow, task or model?
  3. Whether the expected artifact is missing, and whether the token, cost, and duration are abnormal.
Failure attribution means: harness points to environment or dependency issues, workflow means the Skill process did not meet expectations, task indicates task-definition or validation-condition problems, and model indicates unstable model behavior or calls.

How to customize tasks

Tasks are selected according to the following priorities:
  1. The CLI specifies --task.
  2. --quick uses generic-skill-smoke.
  3. The evaluation.tasks in the manifest.
  4. The recommendedTasks in the manifest or the task package source.
  5. When there are no available tasks, automatically generate and cache 2 to 4 evaluation use cases based on the content of the Skill.
If the automatically generated tasks do not fit your Skill, please create or edit comet/eval.yaml in the root directory of the evaluated Skill and declare inline task under evaluation.tasks within it. For example:

Reference the task package within the Skill package

If the task requires an independent task.toml, instruction.md, Docker environment or validation script, the task package can be placed in the Skill directory and then referenced through source. The recommended directory structure is as follows:
In the comet/eval.yaml of the Skill root directory, point to the task package with evaluation.tasks[].source. Here, source is relative to the root directory of the Skill package (the directory containing SKILL.md), not to the comet/ directory:
The following are the minimum sample contents of several files in this task package. For the complete task.toml field, validation script, profile, and Docker configuration instructions, please refer to Configuration Evaluation and Customizing Task task.toml declares task metadata, Docker environment, products to be checked, and verification scripts:
instruction.md is the task instruction sent to the Agent:
environment/Dockerfile provides the basic environment required for task operation and verification:
validation/test_summary.py checks within the task container whether the Agent has produced files that meet the requirements
eval-tasks/writes-summary/task.toml describes the environment and verification method of the task, and instruction.md is the task instruction sent to the Agent. If Docker or deterministic verification scripts are needed, place them respectively under environment/ and validation/ in the same task package. The source task cannot simultaneously write the prompt or expect fields of the inline task, and Comet will check that task.toml and instruction.md do indeed exist in the task package. After configuration, you can directly transfer the Skill directory, and Comet will automatically discover comet/eval.yaml. You can also directly pass manifest:
The task can check files, text, JSON or command results; The task workspace and the expected product must remain within the permitted Skill package or evaluation workspace. Only when it is necessary to customize the Docker environment, validation script, profile or treatment is it necessary to continue reading Configuration Evaluation and Customizing Task

Common evaluation process

Read further

Last modified on September 4, 2026