Real Skill packageSource verifiedClawHub registry

Ab Test Agent Workflow 1.1.0

多智能体双盲 A/B 测试工作流。对多个 AI 模型/Agent 进行多轮次、双盲对照测试。 核心角色:协调者(Coordinator)、受测者 A/B(Contestant)、评测者(Judge)。 触发场景:"A/B 测试"、"双盲测试"、"比较 AI 模型"、"模型评测"、"测试工作流"、 "compare...

Identity and source

Publisher attributionJohnSmithfanregistry owner unverified by skillvetai
Functional categoryAgent Engineering, Security & Governanceautomatically inferred · 55% rule confidence
Package forminstruction with code8 recorded files
Canonical sourceClawHub registryclawhub:johnsmithfan:ab-test-agent-workflow-1-1-0
Open canonical source ↗

Platform declarations

These states come from the source or distribution context. None of the entries below are SkillVetAI compatibility test results.

OpenClawnative officialProvenance: registry distribution

Independent structural checks

These checks parse the fixed package against dated platform rules. They do not execute the Skill or verify task behavior.

Claude Codepasses structure
Checker 0.1.0 · agent-skills-2026-08-13+claude-code-docs-2026-08-13 · 9/16/2026.claude/skills/ab-test-agent-workflow

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

OpenAI Codexpasses structure
Checker 0.1.0 · agent-skills-2026-08-13+codex-docs-2026-08-13 · 9/16/2026.agents/skills/ab-test-agent-workflow

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

OpenClawpasses structure
Checker 0.1.0 · agent-skills-2026-08-13+openclaw-docs-2026-08-13 · 9/16/2026skills/ab-test-agent-workflow

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

Installation and inspection

This command is recorded from the source ecosystem and resolves the registry's latest release. The fixed release shown on this page should be inspected before adoption.

clawhub install @johnsmithfan/ab-test-agent-workflow-1-1-0
clawhub inspect @johnsmithfan/ab-test-agent-workflow-1-1-0 --version 1.0.0

Security evidence

SkillVetAI static result: no findings detected

This automated, non-executing scan is bound to this release hash. It is not a safety certification and may contain false positives or false negatives.

Status
completed
Coverage
full text content
Files
8 / 8 inspected as text
Checked
9/16/2026, 4:58:45 PM
Scanner
0.1.3
Policy
1.0.3
2 inferred permission indicators
  • network access — automatically inferred
  • filesystem read — automatically inferred
1 dependency and API indicators
  • api: clawhub.ai
External clawhub result: suspicious

This is registry-supplied evidence for the recorded release, not an independent SkillVetAI scan. Check the canonical source for the full report, scanner versions, scope, and current moderation state.

Evidence checked
9/16/2026, 6:43:44 AM
Release binding
Matches this record
  • vt: clean
  • skillspector: suspicious
  • llm: suspicious

Recorded files

The catalog stores hashes and an inventory summary for change detection. It does not republish the package contents.

Package content hashsha256:d329094a2eb6dfbb10aa236342606a77828fc9f44734b0b6f9d6f8d5c2cdffe5
Show up to 8 recorded paths
  • _meta.json
  • references/rubric_templates.md
  • references/workflow_guide.md
  • scripts/anonymizer.py
  • scripts/judge_prompts.py
  • scripts/runner.py
  • skill-card.md
  • SKILL.md

Source changelog

ab-test-agent-workflow v1.1.0 introduces a structured, multi-agent double-blind A/B testing workflow for model comparison. - Adds support for multi-round, double-blind evaluation of two models/agents via coordinator, contestants, and judge roles. - Presents complete workflow architecture, including role definitions and communication flow. - Provides detailed prompt templates for each role and various task types (general, code generation), ensuring standardized outputs. - Includes guidance for both fully automated (skill-based) and script-driven execution modes. - Supplies example report formats, rubric quick-reference, and troubleshooting guidance on anonymization, timeouts, and parser fallback. - Lists new scripts for runner, prompt construction/parsing, and output anonymization.