.claude/skills/benchmark-robustness-auditorRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detection, option-letter selection bias (chi2), few-shot curve noise, LLM-judge position/verbosity/rubric-echo bias and hidden-instruction payload detection, paired McNemar + Wilson + deterministic bootstrap for score comparisons, WORKED mitigations (permutation majority ensemble, blind content normalization), documented 0-100 severity formula, h…
These states come from the source or distribution context. None of the entries below are SkillVetAI compatibility test results.
These checks parse the fixed package against dated platform rules. They do not execute the Skill or verify task behavior.
.claude/skills/benchmark-robustness-auditorRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
.agents/skills/benchmark-robustness-auditorRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
skills/benchmark-robustness-auditorRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
This command is recorded from the source ecosystem and resolves the registry's latest release. The fixed release shown on this page should be inspected before adoption.
clawhub install @orionshaowswmw/benchmark-robustness-auditorclawhub inspect @orionshaowswmw/benchmark-robustness-auditor --version 2.0.0This automated, non-executing scan is bound to this release hash. It is not a safety certification and may contain false positives or false negatives.
subprocess.run(\"python3 scripts/benchscan.py report --name selftest-bench --benchmark $SBX/bench.jsonl --corpus $SBX/corp.jsonl --cutoff 2024-06-01 --results $SBX/repro.jsonl --ruignore previous instructionsThis is registry-supplied evidence for the recorded release, not an independent SkillVetAI scan. Check the canonical source for the full report, scanner versions, scope, and current moderation state.
The catalog stores hashes and an inventory summary for change detection. It does not republish the package contents.
sha256:2e44264b02b5bf07e1cd50f00203b3361743a65af9d9a786cf63e559d32b6425CHANGELOG.mddocs/evidence.mddocs/integration.mddocs/operations.mdmanifest.jsonREADME.mdscripts/benchscan.pyscripts/selftest.shskill-card.mdSKILL.mdFull functional rewrite: offline benchscan.py engine (12 subcommands: contam w/ C-1/C-2/C-3, selection, fewshot, judge E-1/E-2/E-3/T-3, compare McNemar+Wilson+bootstrap, tsguess G-1, ensemble + blind worked mitigations, severity, report w/ trends+ledger, audit, doctor); static 17-id catalogue anti-hallucination discipline; hash-chained per-target ledger; 33-check offline selftest (33/33); hardened via 4-lens distributed cross-model review + R1 consolidation pass.