Skip to main content
Frequently asked questions regarding the comet eval assessment system, environmental preparation, report interpretation, and release of evidence.

Basic concepts

comet eval performs real model tasks through a shared eval harness to verify whether a Skill as a product capability can pass the assessment and produce pre-release evidence. It encapsulates pytest, task registry, profile, and report generation. You don’t need to manually spell parameters. For details, please refer to Overview of the Evaluation System
comet eval assesses the product capabilities of Skill and produces reviewable evaluation reports. It can read comet/eval.yaml (the complete package generated by /comet-any), or you can directly consume any local Skill Table of Contents When there is no manifest, normal operation will generate and cache 2 to 4 tasks, with --quick being fixed generic-skill-smoke is smoking. comet skill check checks whether a certain Skill run is missing artifact or state, does not perform model tasks and does not produce release evidence (read) (comet/checks.yaml). Release readiness requires the complete package of comet eval Evidence. For details, see Runtime check.
No need. comet eval encapsulates the underlying details. You only need to know whether to use --manifest or --skill-path. The underlying details are handled by the harness. For more details, see Eval harness。
Creation-time evaluation (comet eval, reading a local Skill directory or comet/eval.yaml) and runtime checks (comet skill check, reading comet/checks.yaml) are separate. The former produces release evidence; the latter checks Run completion. See Evaluation overview · Two evaluation systems.

Environmental preparation

You need uv, Python 3.11+, Docker, the selected Agent CLI, and its model credentials. Core Comet Runtime does not require them; only comet eval does. See Evaluation quickstart · What you need before running.
In most cases the environment is not ready, so samples are skipped. Common causes are Docker not running, missing model credentials, or an unavailable Agent CLI. These runs are marked as “skip”, not as failure. Verify the environment first, then rerun.
macOS/Linux:curl -LsSf https://astral.sh/uv/install.sh | sh。Windows PowerShell:powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1: iex"。uv 'will automatically manage Python versions and eval/.venv`。
Eval harness is missing at ... indicates that the accompanying package eval/ in the npm package is incomplete, or --project points to the wrong repository. Reinstall @rpamis/comet or correct the project path. uv is not installed or not in PATHIt indicates that the harness has been found and only needs to be installed or repaired`uv ’. npm Most users do not need to clone the Comet repository separately.
export ANTHROPIC_API_KEY=sk-ant-..., or configure it in the user-level %USERPROFILE%\.comet\eval\.env/~/.comet/eval/.env (comet eval will load automatically). When using proxy credentials (BigModel/OpenRouter), use ANTHROPIC_AUTH_TOKEN.

Two entrances

To evaluate any local Skill directory, use --skill-path (just pass the directory; when there is no manifest, a task will be generated during normal operation). This is the default entry. When a fixed smoke is needed, explicitly add --quick. Evaluate the complete package generated by /comet-any (with comet/eval.yaml) and run the entire profile task set with --manifest. The result is the evidence for release. To enter the release readiness, the full package manifest must be used. The two are mutually exclusive.
No need. The <current-bundle-hash> generated by /comet-any will be parsed as the current Bundle draft hash before collect/run and written into the temporary manifest; The source Bundle and source comet/eval.yaml will not be modified. An error will only be reported when the manifest has left the original Bundle, bundle.yaml cannot be found, or the draft cannot be loaded.
When --quick is used in conjunction with --skill-path, clearly select generic-skill-smoke task. When a regular directory is running, it generates and caches 2 to 4 tasks based on Skill snapshots. --quick It is low-cost smoke, but it does not mean complete evidence before release. When preparing for release, the complete package manifest path still needs to be followed.
collect only conducts discovery pre-checks (verification) manifest, task, path), without consuming model calls, with the lowest cost, suitable for those who have just generated a Skill The back row is wrong. Only run will carry out the real assessment. collect first It can quickly identify configuration issues and avoid wasting model calls. For more details, please refer to Quick Start
authoring-skill (an 11-dimension rubric with auto_user and maxTurns=8 outer round trips) evaluates Skills generated by /comet-any. It checks package integrity, resolved-Skill evidence, the Engine contract, routing consistency, authoring lanes, and review gates. generic (a 7-dimension rubric with one round by default) handles general Skill smoke tests. See Evaluation overview · Profile.

Reports and Failures

The CLI output prints Report path, commonly .comet/eval/runs/<experiment-id>/summary.html. Always trust the current run’s Experiment and Report path values. For details, see Read the Evaluation Report.
Look at “failure attribution” in the report. harness Explain the environment/dependencies/path issues. workflow indicates that the Skill process did not meet expectations. task Explain the task definition /fixture issue, model It indicates that the model behavior is unstable. Attribution determines what you should change. For details, please refer to Read the Evaluation Report ]
Check if the path is correct: The local Skill directory should contain SKILL.md, and comet/eval.yaml should point to a real existing file. If it is not in the root directory of the Comet repository, adding --project <dir> will point to the correct root directory.
model attribution indicates unstable model behavior or tool usage, and rerunning often helps. If it fails repeatedly, consider lowering The dependency of Skill on non-deterministic behaviors.
It’s considered passed. Rubric is an informative score ([RUBRIC] line and RubricAvg), which does not directly determine whether one passes or not. The true pass/fail is determined by the invocation of the validator and the required skill. A low Rubric score is a diagnostic signal and can be used to optimize skills without affecting the release of access control.

Publish evidence

Eval passing is just one of the conditions. /comet-any or the backend will include the eval evidence in readiness: no evidence, failure, or the corresponding old hash cannot be published. Only after passing and hash matching can you enter review/publish. For details, please refer to Overview of the Evaluation System
It indicates that the Skill was modified after the last evaluation (draft hash The old assessment results are no longer valid. It is necessary to run comet eval ... --html again to generate the current binding Evidence of hash.
No. /comet-any will record structured evidence through the Bundle backend. Manually edit the Bundle The state or internal JSON will break the hash binding and readiness check.
No. --skill-path --quick is only emitting smoke in the early stage, with a limited coverage area. Before release, comet/eval.yaml must be generated through /comet-any, and then a complete evaluation should be run with --manifest.
Last modified on September 4, 2026