SKILL.md to evaluate a Skill using comet eval. Even if you haven’t pre-written the evaluation use cases, Eval will automatically generate 2 to 4 evaluation use cases based on the content of the Skill and run them. This article introduces in the order of actual usage: First, run it through once, then select the Agent and evaluate the backend, and then configure the task, generate the report and locate the cause of failure.
What you need before running
Required content
@rpamis/comethas been installed. The npm package comes with an eval harness, so there is no need to clone the Comet repository.- A local Skill directory containing
SKILL.md. - When running the real evaluation,
uv, Python 3.11+, Docker, the selected Agent CLI and the corresponding model credentials are required.
--collect only performs static discovery and configuration checks, and does not start agents, Docker, plugins, credentials, or network requests. Therefore, it can be used to check the Skill first, without having to prepare a complete model environment from the very beginning.
If you want to select an Agent other than Claude Code, first read the Eval Agent startup configuration to confirm the corresponding CLI, authentication method and model protocol.
User-level .env configuration
Ordinary users do not need to clone the Comet source code, nor do they need to enter eval/ in the source code directory. Run arbitrary for the first time
When the comet eval command is used, the CLI will automatically create a complete user-level configuration file:
- Windows:
%USERPROFILE%\.comet\eval\.env - macOS/Linux:
~/.comet/eval/.env
.env.
The automatically generated template contains all user-configurable parameters and is organized in the following groups:
The parameters in the template are all comments by default. Items that are not filled in will not change their default behavior. The real API key should only be placed at the user level
.env or the current shell cannot write skills, manifest, reports, or public repositories.
Which evaluation agents are supported
By default,comet eval uses claude-code. The main Agent, user simulator and optional Judge can all use the following agents:
The custom Agent adapter is a user-level advanced capability and does not require cloning the Comet source code. Only the built-in Eval harness needs to be modified.
When it comes to tasks or Docker environments, it is only necessary to refer to Advanced Configuration ].
The first run
First, perform a pre-check without calling the model, then run the low-cost smoke once, and finally run the complete task set as neededcomet/eval.yaml, or when there are no evaluation.tasks and recommendedTasks in the manifest, the normal run will automatically generate 2 to 4 evaluation use cases based on the Skill content, freeze and cache them before execution. You don’t need to write down the task first to start the assessment. --quick does not use these automatically generated use cases but runs generic-skill-smoke consistently to verify that the Skill can be injected, invoked, and produce result.md.
How to pass “target”
Directories, directSKILL.md and manifest can all be used as targets:
--project to specify the project directory for saving the running status and reports:
Configure the evaluation Agent and Judge
The simplest way is to select the main Agent in the command line:comet/eval.yaml in the root directory of the evaluated Skill and configure the default Agent in this file. For example, the directory structure is as follows:
comet eval ./my-skill, Comet will automatically detect ./my-skill/comet/eval.yaml. You can also directly pass this file as the target
eval.yaml. CLI options take precedence over manifest:
Select the evaluation backend
--suite selects the evaluation backend. The three backends use the same set of target, task and Agent configurations. The difference lies in whether the results are synchronized to the external evaluation service:
Langfuse example
--collect:
--collect --suite langfuse does not initialize the SDK, does not connect to the Internet, nor does it download plugins.
What types of reports are there
The assessment status and reports are written by default to the following directory in the project the target belongs to:--html is not included, the assessment will still generate summary.md. Run when HTML is needed:
--report-config <path> or COMET_EVAL_REPORT_CONFIG. The CLI will print Experiment and Report path. When searching for reports, the output of this time shall prevail.
When reading the report, pay attention to three things first:
- Evaluate whether it is passed.
- Is the failure attribution
harness,workflow,taskormodel? - Whether the expected artifact is missing, and whether the token, cost, and duration are abnormal.
harness points to environment or dependency issues, workflow means the Skill process did not meet expectations, task indicates task-definition or validation-condition problems, and model indicates unstable model behavior or calls.
How to customize tasks
Tasks are selected according to the following priorities:- The CLI specifies
--task. --quickusesgeneric-skill-smoke.- The
evaluation.tasksin the manifest. - The
recommendedTasksin the manifest or the task packagesource. - When there are no available tasks, automatically generate and cache 2 to 4 evaluation use cases based on the content of the Skill.
comet/eval.yaml in the root directory of the evaluated Skill and declare inline task under evaluation.tasks within it. For example:
Reference the task package within the Skill package
If the task requires an independenttask.toml, instruction.md, Docker environment or validation script, the task package can be placed in the Skill directory and then referenced through source. The recommended directory structure is as follows:
comet/eval.yaml of the Skill root directory, point to the task package with evaluation.tasks[].source. Here, source is relative to the root directory of the Skill package (the directory containing SKILL.md), not to the comet/ directory:
task.toml field, validation script, profile, and Docker configuration instructions, please refer to Configuration Evaluation and Customizing Task
task.toml declares task metadata, Docker environment, products to be checked, and verification scripts:
instruction.md is the task instruction sent to the Agent:
environment/Dockerfile provides the basic environment required for task operation and verification:
validation/test_summary.py checks within the task container whether the Agent has produced files that meet the requirements
eval-tasks/writes-summary/task.toml describes the environment and verification method of the task, and instruction.md is the task instruction sent to the Agent. If Docker or deterministic verification scripts are needed, place them respectively under environment/ and validation/ in the same task package. The source task cannot simultaneously write the prompt or expect fields of the inline task, and Comet will check that task.toml and instruction.md do indeed exist in the task package.
After configuration, you can directly transfer the Skill directory, and Comet will automatically discover comet/eval.yaml. You can also directly pass manifest:

