mimo-2.5 pro as the main Agent and glm-5.2 as Judge. The three treatments ran 240 times in total (pass@5). The sections cover scope, task results, repeated-run metrics, rubric, cost, and failure attribution.
This article is based on a snapshot of the report from this experiment. The
CONTROL in the report are business completion baselines without Comet skills and do not require the production of Comet workflow artifacts. Therefore, it is suitable for business completion comparison, but not for determining whether Comet workflow has been correctly executed.Experimental scope
This experiment compares three treatments:
The experiment consists of 16 Comet workflow tasks. Each treatment incorporates 80 analysis runs - the runs counted in the statistics, with explicit environment failures excluded - equivalent to approximately 5 repeated runs for each task. The report simultaneously counts task results, rubric dimensions,
pass@k/pass^k, costs, runtime overhead, and run-level failed checks.
Core data
The conclusion given in the report is that the comprehensive workflow score ofCOMET_FULL_040_BETA is 0.89, which is higher than 0.82 of COMET_FULL_039, and there is no dimensional regression exceeding the tolerance line of 0.05.
Meanwhile, strict overall pass shows another layer of signal: COMET_FULL_040_BETA is 71/80, lower than COMET_FULL_039 is 76/80. This difference mainly stems from run-level Workflow Contract failure; Business failure in the task matrix is not the main cause.
Therefore, this report needs to be read by perspective: 0.4.0 beta is higher than 0.3.9 in terms of weighted workflow quality, status recovery, and partial process evidence; Meanwhile, the run-level Skill invocation contract still needs to be checked.
pass@k and pass^k
pass@k represents the probability of success at least once out of k attempts and is used to observe the upper limit of ability. pass^k represents the probability of all k attempts being successful and is used to observe the lower limit of reliability.
The main signal here is gap:
COMET_FULL_040_BETA’s pass@5 is 1.00, but pass^5 is 0. This indicates that success can be observed in multiple attempts, but continuous and stable success has not yet been achieved.
Task result
The task matrix reflects whether business tasks have passed
The only task-level failure occurred in
comet-api-cache-ttl of COMET_FULL_039. Under the task result standard, 0.4.0 beta covers all 16 tasks.
However, the task matrix only answers “Whether the task has been completed”. It does not answer whether the Skill was triggered as expected, whether sufficient workflow evidence was retained, or whether the decision point was stably adhered to. Therefore, we still need to look at rubric and failed checks.
Rubric dimension
The weighted score of 0.4.0 beta is higher, mainly frommain_flow, gate_guard and recovery_resilience:
recovery_resilience is the dimension with the highest difference this time. It indicates that 0.4.0 beta scores higher in interrupt recovery, state preservation, and recovery evidence.
There was no improvement in decision_point_compliance and skill_invocation. They correspond to two subsequent inspection directions: whether the decision point is stably exposed to the user, and whether the call evidence dependent on the Skill is stably included in the report.
Cost and runtime
The total token, total cost and average cost of 0.4.0 beta are lower than those of 0.3.9:
The runtime overhead is similar, but the number of tool calls varies:
This indicates that the average number of rounds in 0.4.0 beta is higher, but the number of tool calls is smaller, and the overall time consumption is basically the same as that in 0.3.9.
Failed checks
run-level failed checks mainly focus on Skill invocation contracts:
Such failures do not necessarily mean the business task failed. They indicate that the report did not consistently observe expected Skill-call evidence. For workflow Skills this still matters, because evaluation checks both final outcomes and whether the process is traceable and recoverable.
LLM judge overlay
After re-reading the artifact, the LLM judge gave independent scores to the three qualitative dimensions:
This set of readings indicates that the regularized rubric observed better process and recovery evidence at 0.4.0 beta. However, in terms of the content quality of artifact, 0.4.0 beta is close to 0.3.9, with some dimensions slightly lower. Subsequent optimization can simultaneously focus on two types of signals: structured process evidence, as well as the content density of Proposals, designs, tasks, and verify Artifacts.
How to use this result
This report can be used as a baseline reading that includes multiple indicators, avoiding simplifying the results into a single pass/fail conclusion.
If the 0.4.0 beta is to be further advanced, two types of issues should be prioritized for inspection: whether the evidence of dependency Skill calls stably enters the stream-json, and whether the decision points stably require user confirmation. After the repair, the same set of tasks and the same number of repetitions should be used for re-running and comparison to avoid mistaking task differences or sample differences for version differences.
Integration with LangSmith and Langfuse
Comet Eval can sync results to LangSmith or Langfuse for experiment records, run traces, and rubric metrics.
manage your Skill baseline in LangSmith, view detailed evaluation metrics, latency and Token consumption

track your Claude Code full link in LangSmith

tracks the custom Rubric metric
in LangSmith via PytestMimo real Token consumption

Before the experiment begins

experiment,
This experiment consumed more than 13 billion Credits, equivalent to over 1.5 billion tokens.Original report
A snapshot of the original HTML report is embedded below. It contains complete charts, task matrices, source evidence, raw vs analysis sensitivity, failed checks and LLM judge overlay.Next step
- Scoring Criteria and Dual-Agent Evaluation - Understanding
pass@k,pass^k, rubric and LLM judge - Read the evaluation report - Learn to locate problems from summary, report JSON and failed checks.
- Eval harness - Understand how the evaluation runs and how to collect stream-json evidence.

