Real Skill packageSource verifiedClawHub registry

Agent Evaluation Benchmark Engine

Objectively evaluates OpenClaw agent improvements using baseline benchmarks, regression checks, golden tests, scoring, and upgrade gating across skills and w...

Identity and source

Publisher attributionSuga Nickregistry owner unverified by skillvetai
Functional categoryAgent Engineering, Security & Governanceautomatically inferred · 64% rule confidence
Package forminstruction bundle5 recorded files
Canonical sourceClawHub registryclawhub:pmuhammadagus-byte:agent-evaluation-benchmark-engine
Open canonical source ↗

Platform declarations

These states come from the source or distribution context. None of the entries below are SkillVetAI compatibility test results.

OpenClawnative officialProvenance: registry distribution

Independent structural checks

These checks parse the fixed package against dated platform rules. They do not execute the Skill or verify task behavior.

Claude Codeissues found
Checker 0.1.0 · agent-skills-2026-08-13+claude-code-docs-2026-08-13 · 8/23/2026.claude/skills/<skill-name>
1 structural issue
  • error: YAML frontmatter could not be parsed: Nested mappings are not allowed in compact mappings at line 6, column 12: changelog: ClawHub professional standard: Overview, When to Use, How to Use, Co… ^ SKILL.md

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

OpenAI Codexissues found
Checker 0.1.0 · agent-skills-2026-08-13+codex-docs-2026-08-13 · 8/23/2026.agents/skills/<skill-name>
1 structural issue
  • error: YAML frontmatter could not be parsed: Nested mappings are not allowed in compact mappings at line 6, column 12: changelog: ClawHub professional standard: Overview, When to Use, How to Use, Co… ^ SKILL.md

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

OpenClawissues found
Checker 0.1.0 · agent-skills-2026-08-13+openclaw-docs-2026-08-13 · 8/23/2026skills/<skill-name>
1 structural issue
  • error: YAML frontmatter could not be parsed: Nested mappings are not allowed in compact mappings at line 6, column 12: changelog: ClawHub professional standard: Overview, When to Use, How to Use, Co… ^ SKILL.md

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

Installation and inspection

This command is recorded from the source ecosystem and resolves the registry's latest release. The fixed release shown on this page should be inspected before adoption.

clawhub install @pmuhammadagus-byte/agent-evaluation-benchmark-engine
clawhub inspect @pmuhammadagus-byte/agent-evaluation-benchmark-engine --version 1.0.0

Security evidence

SkillVetAI static result: medium signal

This automated, non-executing scan is bound to this release hash. It is not a safety certification and may contain false positives or false negatives.

Status
completed
Coverage
full text content
Files
5 / 5 inspected as text
Checked
8/23/2026, 2:01:35 AM
Scanner
0.1.3
Policy
1.0.3
1 automated finding
mediumSKILL.md YAML frontmatter is invalidSKILL.md:1 · confidence 100%Nested mappings are not allowed in compact mappings at line 6, column 12: changelog: ClawHub professional standard: Overview, When to Use, How to Use, Co… ^
2 inferred permission indicators
  • network access — automatically inferred
  • filesystem write — automatically inferred
2 dependency and API indicators
  • api: clawhub.ai
  • api: github.com
External clawhub result: clean

This is registry-supplied evidence for the recorded release, not an independent SkillVetAI scan. Check the canonical source for the full report, scanner versions, scope, and current moderation state.

Evidence checked
8/22/2026, 6:47:26 PM
Release binding
Matches this record
  • vt: clean
  • skillspector: clean
  • llm: clean

Recorded files

The catalog stores hashes and an inventory summary for change detection. It does not republish the package contents.

Package content hashsha256:4303974af7f425a0b193f660e2b430d32eff1aaa31968fc105b5d24a1fa6cb11
Show up to 5 recorded paths
  • _meta.json
  • LICENSE
  • README.md
  • skill-card.md
  • SKILL.md

Source changelog

- Added comprehensive documentation outlining evaluation philosophy, methodology, and metrics for benchmarking OpenClaw agents. - Defines objective skill, workflow, agent, and system-level testing using baselines, golden tests, and structured evaluation loops. - Details test case structure, scoring rubrics, and explicit criteria for assessing task success, reasoning quality, tool/skill use, regression, and robustness. - Provides clear red flags and rationalizations to avoid; emphasizes evidence-based acceptance or rejection of changes. - Establishes pro-level procedures for regression detection, security evaluation, and context/resource stress testing.