Real Skill packageSource verifiedClawHub registry

benchmark-robustness-auditor

Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detection, option-letter selection bias (chi2), few-shot curve noise, LLM-judge position/verbosity/rubric-echo bias and hidden-instruction payload detection, paired McNemar + Wilson + deterministic bootstrap for score comparisons, WORKED mitigations (permutation majority ensemble, blind content normalization), documented 0-100 severity formula, h…

Identity and source

Publisher attributionorionshaowswmwregistry owner unverified by skillvetai
Functional categoryDocuments, PDF & Presentationsautomatically inferred · 62% rule confidence
Package forminstruction with code10 recorded files
Canonical sourceClawHub registryclawhub:orionshaowswmw:benchmark-robustness-auditor
Open canonical source ↗

Platform declarations

These states come from the source or distribution context. None of the entries below are SkillVetAI compatibility test results.

OpenClawnative officialProvenance: registry distribution

Independent structural checks

These checks parse the fixed package against dated platform rules. They do not execute the Skill or verify task behavior.

Claude Codepasses structure
Checker 0.1.0 · agent-skills-2026-08-13+claude-code-docs-2026-08-13 · 9/6/2026.claude/skills/benchmark-robustness-auditor

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

OpenAI Codexpasses structure
Checker 0.1.0 · agent-skills-2026-08-13+codex-docs-2026-08-13 · 9/6/2026.agents/skills/benchmark-robustness-auditor

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

OpenClawpasses structure
Checker 0.1.0 · agent-skills-2026-08-13+openclaw-docs-2026-08-13 · 9/6/2026skills/benchmark-robustness-auditor

Runtime, accounts, dependencies, permissions, network behavior and task quality remain untested.

Installation and inspection

This command is recorded from the source ecosystem and resolves the registry's latest release. The fixed release shown on this page should be inspected before adoption.

clawhub install @orionshaowswmw/benchmark-robustness-auditor
clawhub inspect @orionshaowswmw/benchmark-robustness-auditor --version 2.0.0

Security evidence

SkillVetAI static result: high signal

This automated, non-executing scan is bound to this release hash. It is not a safety certification and may contain false positives or false negatives.

Status
completed
Coverage
full text content
Files
10 / 10 inspected as text
Checked
9/6/2026, 4:36:52 AM
Scanner
0.1.3
Policy
1.0.3
2 automated findings
highDynamic code or shell execution is presentscripts/selftest.sh:92 · confidence 78%subprocess.run(\"python3 scripts/benchscan.py report --name selftest-bench --benchmark $SBX/bench.jsonl --corpus $SBX/corp.jsonl --cutoff 2024-06-01 --results $SBX/repro.jsonl --ru
mediumPrompt-override language requires reviewscripts/selftest.sh:45 · confidence 62%ignore previous instructions
1 High/Critical review queue entry
STATIC_DYNAMIC_CODE_EXECUTIONpending
Open human review queue →
5 inferred permission indicators
  • shell execution — automatically inferred
  • network access — automatically inferred
  • filesystem read — automatically inferred
  • filesystem write — automatically inferred
  • credential access — automatically inferred
1 dependency and API indicators
  • api: clawhub.ai
External clawhub result: clean

This is registry-supplied evidence for the recorded release, not an independent SkillVetAI scan. Check the canonical source for the full report, scanner versions, scope, and current moderation state.

Evidence checked
9/6/2026, 4:06:31 AM
Release binding
Matches this record
  • vt: clean
  • skillspector: suspicious
  • llm: clean

Recorded files

The catalog stores hashes and an inventory summary for change detection. It does not republish the package contents.

Package content hashsha256:2e44264b02b5bf07e1cd50f00203b3361743a65af9d9a786cf63e559d32b6425
Show up to 10 recorded paths
  • CHANGELOG.md
  • docs/evidence.md
  • docs/integration.md
  • docs/operations.md
  • manifest.json
  • README.md
  • scripts/benchscan.py
  • scripts/selftest.sh
  • skill-card.md
  • SKILL.md

Source changelog

Full functional rewrite: offline benchscan.py engine (12 subcommands: contam w/ C-1/C-2/C-3, selection, fewshot, judge E-1/E-2/E-3/T-3, compare McNemar+Wilson+bootstrap, tsguess G-1, ensemble + blind worked mitigations, severity, report w/ trends+ledger, audit, doctor); static 17-id catalogue anti-hallucination discipline; hash-chained per-target ledger; 33-check offline selftest (33/33); hardened via 4-lens distributed cross-model review + R1 consolidation pass.