comet eval is Comet’s standalone Skill-evaluation entry. You can evaluate an existing local Skill directly, without running /comet-any first or preparing comet/eval.yaml ahead of time.
Shortest path
target can be a Skill directory, a direct SKILL.md, or comet/eval.yaml. When you pass a directory, Comet discovers its manifest automatically. Without a manifest, a normal run generates 2–4 bounded tasks from a Skill snapshot without rewriting the Skill source.
Three ways to point at a target cover different situations:
--collectperforms static discovery and configuration checks without starting an Agent, Docker, plugins, credentials, or network requests.--quickuses the fixedgeneric-skill-smoketask for early smoke testing.--htmlproduces both Markdown and HTML reports.- Reports and run state are written under
.comet/eval/runs/in the Skill or the project passed with--project.
User-level .env configuration
Published Comet CLI users do not need to pull the source repository. Place the configuration file beside the existing user-level Eval adapter directory:
.env; CLI options and the manifest take precedence over environment defaults. Claude Code, Codex, and CodeBuddy map the common values to their native variables; Codex uses its isolated runtime config.toml, and Qoder uses only its officially supported authentication and service configuration. Explicit Agent-native variables override the common fallbacks.
Security boundary: Eval never writes API keys into the published package, manifest, report, or Skill workspace. During Docker Agent runs, Codex, Qoder, and CodeBuddy use isolated container-local temporary config roots. Codex’s config.toml contains only an environment-variable reference, CodeBuddy’s settings.json uses only apiKeyHelper, and the actual key exists only in the current container process and is destroyed with the temporary config root after the run.
Task selection
Tasks are selected in this order: explicit--task, --quick, evaluation.tasks in the manifest, recommendedTasks, and finally generated tasks. Generated tasks are cached by the Skill snapshot, Agent, and evaluation configuration so later runs can reuse them.
For stable acceptance conditions, define an inline task in comet/eval.yaml:
source to reference a task package containing task.toml and instruction.md inside the Skill package. Inline tasks support file, text, JSON, and command checks. The task workspace and expected artifacts must remain inside the allowed Skill package or evaluation workspace.
Choose an evaluation Agent
The default Agent isclaude-code. You can also select codex, qoder, or codebuddy:
Choose an evaluation suite
local is the default and is suitable for local development and HTML reports. For team tracing, use the existing LangSmith suite:
--collect --suite langfuse does not initialize the SDK, make network requests, or download plugins.
Common options
Reading the result
Before running, the CLI prints the target, suite, task, experiment, and report path. The report attributes failures toharness, workflow, task, or model:
harness: dependency, Docker, path, or environment problem;workflow: the Skill flow did not meet its intended behavior;task: the task definition or acceptance condition is incorrect;model: model behavior or invocation instability.
comet eval evaluates a Skill as a product capability. comet skill check checks whether one Skill Run satisfies its runtime checks; they serve different purposes.
Next steps
- Eval Agent startup configuration — prepare the CLI, authentication, and model protocol for Claude Code, Codex, or Qoder
- Evaluation system overview — evaluation configuration and scoring
- Reading evaluation reports — pass rates and failure attribution
- comet publish — use evaluation evidence for publishing decisions

