.claude/skills/ab-test-agent-workflowRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
多智能体双盲 A/B 测试工作流。对多个 AI 模型/Agent 进行多轮次、双盲对照测试。 核心角色:协调者(Coordinator)、受测者 A/B(Contestant)、评测者(Judge)。 触发场景:"A/B 测试"、"双盲测试"、"比较 AI 模型"、"模型评测"、"测试工作流"、 "compare...
These states come from the source or distribution context. None of the entries below are SkillVetAI compatibility test results.
These checks parse the fixed package against dated platform rules. They do not execute the Skill or verify task behavior.
.claude/skills/ab-test-agent-workflowRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
.agents/skills/ab-test-agent-workflowRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
skills/ab-test-agent-workflowRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
This command is recorded from the source ecosystem and resolves the registry's latest release. The fixed release shown on this page should be inspected before adoption.
clawhub install @johnsmithfan/ab-test-agent-workflow-1-1-0clawhub inspect @johnsmithfan/ab-test-agent-workflow-1-1-0 --version 1.0.0This automated, non-executing scan is bound to this release hash. It is not a safety certification and may contain false positives or false negatives.
This is registry-supplied evidence for the recorded release, not an independent SkillVetAI scan. Check the canonical source for the full report, scanner versions, scope, and current moderation state.
The catalog stores hashes and an inventory summary for change detection. It does not republish the package contents.
sha256:d329094a2eb6dfbb10aa236342606a77828fc9f44734b0b6f9d6f8d5c2cdffe5_meta.jsonreferences/rubric_templates.mdreferences/workflow_guide.mdscripts/anonymizer.pyscripts/judge_prompts.pyscripts/runner.pyskill-card.mdSKILL.mdab-test-agent-workflow v1.1.0 introduces a structured, multi-agent double-blind A/B testing workflow for model comparison. - Adds support for multi-round, double-blind evaluation of two models/agents via coordinator, contestants, and judge roles. - Presents complete workflow architecture, including role definitions and communication flow. - Provides detailed prompt templates for each role and various task types (general, code generation), ensuring standardized outputs. - Includes guidance for both fully automated (skill-based) and script-driven execution modes. - Supplies example report formats, rubric quick-reference, and troubleshooting guidance on anonymization, timeouts, and parser fallback. - Lists new scripts for runner, prompt construction/parsing, and output anonymization.