01-eval-design
What this skill does
Design high-quality evaluation datasets for AI systems, including creating test cases, stratifying samples, generating adversarial examples, and outputting datasets in OpenJudge-compatible format.
github/agentscope-ai - general - 733 stars
Threat analysis
Skill info
pkg:github/agentscope-ai/OpenJudge@2151def?skill=01-eval-designAssessments
No risk patterns detected in this scan. A clean automated scan is a good signal, not a guarantee.
Badge
Add the Anomity scan badge for 01-eval-design to your README.
How Anomity governs this at runtime
Scan-time vetting tells you what a skill says it will do. Anomity's Endpoint Sensor sees what agents actually do: it discovers skills alongside every other AI artifact on the endpoint, and runtime governance can allow, deny, or log the tool calls a skill triggers. Policy violations route to your SIEM, Slack, email, or Jira, backed by a queryable 90-day audit trail.
Book a 30-minute demo to see your own skill inventory.
Methodology and disputes
Every skill is assessed by the Anomity Skill Intelligence engine against its public source; findings indicate risk patterns, not confirmed exploitation. Maintainer of 01-eval-design? Report an issue or request a rescan.




