comet eval assessment system, environmental preparation, report interpretation, and release of evidence.
Basic concepts
What exactly does comet eval evaluate
What exactly does comet eval evaluate
comet eval performs real model tasks through a shared eval harness to verify whether a Skill as a product capability can pass the assessment and produce pre-release evidence. It encapsulates pytest, task registry, profile, and report generation. You don’t need to manually spell parameters. For details, please refer to Overview of the Evaluation System What are the differences between comet eval and comet skill check
What are the differences between comet eval and comet skill check
comet eval assesses the product capabilities of Skill and produces reviewable evaluation reports. It can read
comet/eval.yaml (the complete package generated by /comet-any), or you can directly consume any local Skill
Table of Contents When there is no manifest, normal operation will generate and cache 2 to 4 tasks, with --quick being fixed
generic-skill-smoke is smoking. comet skill check checks whether a certain Skill run is missing
artifact or state, does not perform model tasks and does not produce release evidence (read)
(comet/checks.yaml). Release readiness requires the complete package of comet eval
Evidence. For details, see Runtime check.Do I need to understand pytest or Docker
Do I need to understand pytest or Docker
comet eval encapsulates the underlying details. You only need to know whether to use --manifest or
--skill-path. The underlying details are handled by the harness. For more details, see Eval
harness。What do the two sets of assessment systems mean
What do the two sets of assessment systems mean
comet eval, reading a local Skill directory or comet/eval.yaml) and runtime checks (comet skill check, reading comet/checks.yaml) are separate. The former produces release evidence; the latter checks Run completion. See Evaluation overview · Two evaluation systems.Environmental preparation
What environment is needed to run eval
What environment is needed to run eval
uv, Python 3.11+, Docker, the selected Agent CLI, and its model credentials. Core Comet Runtime does not require them; only comet eval does. See Evaluation quickstart · What you need before running.The evaluation passed instantly, but it looks like it never ran
The evaluation passed instantly, but it looks like it never ran
How to install uv
How to install uv
curl -LsSf https://astral.sh/uv/install.sh | sh。Windows
PowerShell:powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1: iex"。uv 'will automatically manage Python versions and eval/.venv`。What's the difference between Eval harness is missing and uv missing
What's the difference between Eval harness is missing and uv missing
Eval harness is missing at ... indicates that the accompanying package eval/ in the npm package is incomplete, or
--project points to the wrong repository. Reinstall @rpamis/comet or correct the project path. uv is not installed or not in PATHIt indicates that the harness has been found and only needs to be installed or repaired`uv ’. npm
Most users do not need to clone the Comet repository separately.How to configure an API key
How to configure an API key
export ANTHROPIC_API_KEY=sk-ant-..., or configure it in the user-level %USERPROFILE%\.comet\eval\.env/~/.comet/eval/.env (comet eval will load automatically). When using proxy credentials (BigModel/OpenRouter), use ANTHROPIC_AUTH_TOKEN.Two entrances
Which one should be used, --manifest or --skill-path
Which one should be used, --manifest or --skill-path
--skill-path (just pass the directory; when there is no manifest, a task will be generated during normal operation). This is the default entry. When a fixed smoke is needed, explicitly add --quick. Evaluate the complete package generated by /comet-any (with comet/eval.yaml) and run the entire profile task set with --manifest. The result is the evidence for release. To enter the release readiness, the full package manifest must be used. The two are mutually exclusive.Does the current-bundle-hash in the manifest need to be replaced manually
Does the current-bundle-hash in the manifest need to be replaced manually
<current-bundle-hash> generated by /comet-any will be parsed as the current Bundle draft hash before collect/run and written into the temporary manifest; The source Bundle and source comet/eval.yaml will not be modified. An error will only be reported when the manifest has left the original Bundle, bundle.yaml cannot be found, or the draft cannot be loaded.What is quick and what is the default task
What is quick and what is the default task
--quick is used in conjunction with --skill-path, clearly select generic-skill-smoke
task. When a regular directory is running, it generates and caches 2 to 4 tasks based on Skill snapshots. --quick
It is low-cost smoke, but it does not mean complete evidence before release. When preparing for release, the complete package manifest path still needs to be followed.Why collect first and then run
Why collect first and then run
collect only conducts discovery pre-checks (verification)
manifest, task, path), without consuming model calls, with the lowest cost, suitable for those who have just generated a Skill
The back row is wrong. Only run will carry out the real assessment. collect first
It can quickly identify configuration issues and avoid wasting model calls. For more details, please refer to Quick Start Reports and Failures
Where can I find the report
Where can I find the report
Report path, commonly .comet/eval/runs/<experiment-id>/summary.html. Always trust the current run’s Experiment and Report path values. For details, see Read the Evaluation Report.The assessment failed. How can I determine where the problem lies
The assessment failed. How can I determine where the problem lies
harness
Explain the environment/dependencies/path issues. workflow indicates that the Skill process did not meet expectations. task
Explain the task definition /fixture issue, model
It indicates that the model behavior is unstable. Attribution determines what you should change. For details, please refer to Read the Evaluation Report ]collect reports that it cannot find the target
collect reports that it cannot find the target
SKILL.md, and comet/eval.yaml should point to a real existing file. If it is not in the root directory of the Comet repository, adding --project <dir> will point to the correct root directory.Do I need to run again if the model attribution fails
Do I need to run again if the model attribution fails
model attribution indicates unstable model behavior or tool usage, and rerunning often helps. If it fails repeatedly, consider lowering
The dependency of Skill on non-deterministic behaviors.If the Rubric score is very low but all the checks are passed, does that count as passing
If the Rubric score is very low but all the checks are passed, does that count as passing
[RUBRIC] line and RubricAvg), which does not directly determine whether one passes or not. The true pass/fail is determined by the invocation of the validator and the required skill. A low Rubric score is a diagnostic signal and can be used to optimize skills without affecting the release of access control.Publish evidence
Can it be published once Eval is approved
Can it be published once Eval is approved
/comet-any or the backend will include the eval evidence in readiness: no evidence, failure, or the corresponding old hash cannot be published. Only after passing and hash matching can you enter review/publish. For details, please refer to Overview of the Evaluation System What does it mean that the eval evidence corresponds to the old hash
What does it mean that the eval evidence corresponds to the old hash
comet eval ... --html again to generate the current binding
Evidence of hash.Can I manually write the report path into the publication status
Can I manually write the report path into the publication status
/comet-any will record structured evidence through the Bundle backend. Manually edit the Bundle
The state or internal JSON will break the hash binding and readiness check.Can quick smoke be used as release evidence
Can quick smoke be used as release evidence
--skill-path --quick is only emitting smoke in the early stage, with a limited coverage area. Before release, comet/eval.yaml must be generated through /comet-any, and then a complete evaluation should be run with --manifest.
