Skip to main content
The design of Comet did not come out of thin air. There are already relevant practices and discussions in the industry regarding the evaluation methodology, Skill creation and distribution, and runtime and workflow. When Comet began its design, some articles had not yet been published. Near completion, we found that many of Comet’s practices were in line with these practical directions **, and some of them were extensions of Comet for workflow Skill scenarios. This article lists each practice of Comet by dimension, indicates which industry article it corresponds to, and provides references at the end of the article. The article does not elaborate on the internal details of these reference articles - if you wish to understand the specific mechanism of Comet evaluation, please refer to Evaluation The compass for Skill evolution and How Comet uses evaluation to drive Skill evolution ; the purpose of this article is to provide an index of “Which industry references does Comet’s approach correspond to?”

I. Evaluation Methodology

The object evaluated by Comet eval is “whether the Skill bound to the Agent is effective”, rather than the capabilities of the model itself. This positioning determines its approach in terms of evaluation objects, comparison methods, indicator design and interaction forms.

The evaluation object: The incremental value of Skill, rather than the model’s capabilities

Comet used the “treatment system” (CONTROL baseline without Skill vs. COMET_FULL injection of the complete Skill stack) for control with and without to measure the marginal effect of Skill under the same task and the same isolated environment. This approach - taking Skill as the first-class evaluation object and using paired experiments to measure “whether adding Skill improves” - is consistent with the positioning of SkillsBench. [1]. The evaluation experiment of Comet was designed with reference to the practical article [2] of LangChain. The latter also emphasizes that evaluating skill is actually assessing whether the “agent + skill “system can effectively utilize skill information.

Multi-round interaction: Dual agents automatically simulate users

The workflow Skill will pause and wait for the user at the decision point. The evaluation needs to be able to automatically run the entire interaction process. Comet solves this problem by using automatic interaction between two agents (the Agent under test runs Skill, and the user simulates the Agent’s response at the decision point). The practice of using agents to simulate users and enabling the evaluation to run the entire workflow unattended shares the same concept in Alibaba’s Harness engineering solution.

Indicator design: rubric multi-dimensional scoring + separation of capability and reliability

Comet’s rubric breaks down quality into multiple dimensions, each of which consists of binary checkitems and is weighted and aggregated. This approach of “replacing subjective scoring with objectively verifiable binary signals” is consistent with the evaluation methodologies of SWE-bench, SWE-bench Pro, tau-bench, HumanEval, etc. [5][6][7]. The rubric scoring methodology section in the Comet documentation clearly aligns with practices such as Galileo, Hebbia, and tau-bench. Comet uses both pass@k (upper limit of capability) and pass^k (lower limit of reliability) to distinguish between “can do” and “can do every time”. pass@k using HumanEval ‘s unbiased estimator [7]. It is worth noting that SkillsBench, for the sake of certainty, does not use pass@k but instead adopts a fixed average pass rate of three repetitions - this is a difference point between Comet and it in the selection of metrics.

Harness engineering: Environmental isolation and observability

Comet uses Docker to isolate each task, pytest drivers, and automatically generates reports. The approach of treating the evaluation itself as a system that needs to be engineered (environmental isolation, task scheduling, result collection, report generation) It is reflected in both Tencent’s large-scale Agent evaluation practice [4] and Alibaba’s Harness engineering solution [3]. Comet’s LangSmith integration reuses the same local task suite and connects the evaluation to the enterprise-level observability platform. The idea of this observability platform integration also references the practice of LangChain. [2].

Skill trigger guarantee: Call evidence as hard access control

If the model does not invoke the Skill, what is actually being evaluated is the bare model. The practical article of LangChain directly points out this problem - even with the addition of prompt words, the call rate is only about 70%[2]. Comet’s approach is to use the CLAUDE.md contract, PreToolUse hook, and stream-json invocation for evidence parsing, treating Skill invocation as hard access control. This is the engineering solution proposed by Comet for this issue. The reference article points out the existence of the problem.

Summary of comparison

Ii. Runtime and Workflow

Cross-platform Node runtime

0.4.0 rewrites the 7 Bash scripts as the Node .mjs launcher, backed by a shared TypeScript runtime, so that Windows users no longer need Git Bash/WSL. Cross-platform runtime is a common engineering issue in the agent framework. This article does not list the reference articles on this dimension one by one. For the specific approach of Comet, you can refer to Runtime Refactoring from 0.3.9 to 0.4.0 ]

Recoverable workflow and state layering

Comet breaks down the status of each change into three layers: user-readable .comet.yaml, machine-managed run-state.json, and appended audit logs state-events.jsonl. After an interruption, it can be precisely restored from the checkpoint instead of having the Agent re-guess the context. Recoverability is the core design of Comet workflow. For more details, please refer to State Management ]. State persistence and recovery are common topics in workflow engines. This article does not list reference articles on this dimension one by one.

Intent recognition and slot extraction

Comet structures “what users want to do” into intent frames (CometIntentFrame), which include intent classification (intent with confidence), entity extraction (entities), and slot filling (slots) Such as risk signals like requested_action, workflow_candidate, scope, public_api_change, etc., and then the route is re-scored and determined by the runtime according to a fixed priority. This is the ** mainstream technology in the industry ** of natural language understanding (NLU) and task-based dialogue systems - intent classification and slot filling are the standard methods for industrial-grade dialogue systems (intelligent customer service, voice assistants, task robots). Comet applies this mature technology to the routing of the AI coding workflow, upgrading the routing from “natural language rule guessing” to “structured evidence + fixed priority scoring”. For details, please refer to Intent Recognition and Routing

External access control: eval detection + hook interception

Comet’s phase protection is supported by two mechanisms: comet-guard and comet-state transition share the same TypeScript transfer semantic table, and the process is directly blocked when any phase condition is not met. The PreToolUse hook (comet-hook-guard.mjs) intercepts before the tool call is executed, imposing runtime hard constraints on stage boundaries and dangerous operations. This stability fulcrum of “eval detection + hook interception “is consistent with the harness layering practice of the Alibaba Cloud Developer community. This paper regards the verification mechanism as the core of harness stability and describes the same idea with G1-G8 access control walls (deterministic checksum forced blocking) + hook interception (real-time fencing). This “external access control” judgment is supported by research: The explanatory power of Effective Feedback Compute for the variance of agent success rate reaches R²=0.94-0.99, which is much higher than that of the original token consumption and tool invocation (R²=0.33-0.42) [8]. This indicates that the impact of strict quality inspection on reliability exceeds the budget scale. This is consistent with Comet’s design motivation of using eval evidence as a release access and hook as a hard constraint at runtime. The Harness Engineering practice of Alibaba technology further summarizes this system as a complete closed loop of constraint mechanisms, feedback loops, and structured Rubric evaluations. [9] corresponds to Comet’s “create → evaluate → release” closed loop.

Iii. Skill Creation and Distribution

Combine any Skill into a distributable Bundle

/comet-any combines any Skill into a stable Skill Bundle with workflow protocol, resolved skill evidence, Engine four-piece set and required control plane. The complete chain of /comet-any is “Create → Evaluate → review → publish → distribute”, and preview is forced before distribution. This is Comet’s solution for the requirement of “enabling users to create reusable, reviewable, and distributable Skill packages”. For more details, please refer to Complete process of Skill Creator . The skill ecosystem in the industry (such as Anthropic’s Agent Skills and the common skills add CLI) provides the definition format and installation mechanism of skill. Comet has added composition, deterministic Bundle, release evidence chain and distribution preview on this basis. The industry reference for this dimension is mainly based on the official documentation of each skill ecosystem, which is not listed one by one in this article.

Release access control: eval evidence binding draft hash

Comet makes the eval evidence a mandatory access control for publishing - publishing cannot be done without the eval evidence of the current draft hash, if the eval fails, or if the evidence corresponds to an old hash. The practice of integrating evaluation results into the release process and using evidence to constrain release decisions originated from LangChain’s idea of “iterating skills with evaluations”. The extension of Comet lies in implementing this constraint as a hard access control that binds hash rather than a soft suggestion.

Iv. Extensions of Comet and Unresolved Issues

The purpose of the comparison is not to prove that Comet is fully aligned with the industry - this article does not list reference articles for all dimensions one by one, and Comet has not completely solved some problems either. ** An extension of Comet ** (There are relevant discussions in the industry, and Comet has been further engineered) :
  • Dual-agent interaction + invocation evidence hard access control to ensure that Skill is truly invoked;
  • The state is stratified into three layers, translating recoverability from a concept to specific documents and contracts.
  • Intent recognition and slot extraction, applying the mainstream NLU technology to the AI coding workflow routing;
  • Release the access control binding draft hash to make eval evidence a hard constraint.
** Unresolved issues ** (SkillsBench and others also admit that the field is still in its early stages) :
  • There is no gold standard for the evaluation of Skill binding to date. [1];
  • Skill injection increases the context length, but the attribution of gain remains impure [1];
  • The authenticity of user simulation, the reliability of LLM-as-judge, and the evaluation coverage of long-term workflows are all open questions [1].
Reference articles for these issues have been listed at the end of the text. It is recommended to read them as an extension.

References

  1. Li, X., et al. SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.arxiv preprint arXiv:2602.12670, 2026. [Paper link ]
  2. Xu, R. Evaluating Skills.langchain Blog, 2026-03-05. [Original link ]
  3. Alibaba Technology. * Evaluation Scheme of Harness Engineering Construction Business Agent Based on Top-level Agent (Claude Code) *. Wechat Public Platform, 2026. [Original link: ]
  4. Tencent Technology Engineering. *AI Agent & Skill Evaluation Plan and Implementation Practice * Wechat Public Platform, 2026. [Original link: ]
  5. Jimenez, C. E., et al. *SWE-bench: Can Language Models Resolve Real-World GitHub Issues? * ICLR, 2024. [Paper link ]
  6. Deng, X., et al. *SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? * arXiv preprint arXiv:2509.16941, 2025. [Paper link ]
  7. Chen, M., et al. Evaluating Large Language Models Trained on Code (HumanEval). arXiv preprint arXiv:2107.03374 2021. [Paper link ]
  8. Alibaba Cloud Developer Community. *AI is not short of intelligence but discipline: harness Claude Code to ensure the stable implementation of long-term tasks *. Wechat Public Platform, 2026. [Original link: ]
  9. Alibaba Technology. *Harness Engineering: Long-Term Automation AI Coding/Skills Development Practice *. Wechat Public Platform, 2026. [Original link: ]

Extended Reading

  • From 0.3.9 to 0.4.0: How Comet Evolved from a Workflow Layer to a Skill Platform ](/en/tech-blog/comet-0.3.9-to-0.4.0) - Comet’s Own Architectural Evolution
  • Evaluation: The Compass of Skill Evolution - Detailed Explanation of the Evaluation Methodology
  • How does Comet use evaluation to drive Skill evolution? ](/en/eval/eval-driven-evolution) - Open Challenges in Evaluation and Comet’s Engineering Solutions
  • Scoring criteria and dual-agent evaluation - rubric dimension details and pass@k/pass^k formula
最后修改于 2026年8月31日