Evals
What this skill does
AI evaluation framework for testing and benchmarking AI agents
github/danielmiessler - Excessive Agency - 17.9k stars
Threat analysis
Skill info
pkg:github/danielmiessler/LifeOS@58381b3?skill=evalsAssessments (2)
Excessive Agency
Excessive Agency via local-llm-review
Workflows/CreateJudge.md
The workflow allows creation of custom LLM-as-Judge configurations, which could be misused to create biased or malicious grading systems if not properly controlled.Excessive Agency via local-llm-review
Workflows/CreateUseCase.md
The workflow enables creation of new evaluation use cases with test cases and scoring criteria, which could be used to create misleading or harmful evaluation scenarios if not properly validated.Badge
Add the Anomity scan badge for Evals to your README.
How Anomity governs this at runtime
Scan-time vetting tells you what a skill says it will do. Anomity's Endpoint Sensor sees what agents actually do: it discovers skills alongside every other AI artifact on the endpoint, and runtime governance can allow, deny, or log the tool calls a skill triggers. Policy violations route to your SIEM, Slack, email, or Jira, backed by a queryable 90-day audit trail.
Book a 30-minute demo to see your own skill inventory.
Methodology and disputes
Every skill is assessed by the Anomity Skill Intelligence engine against its public source; findings indicate risk patterns, not confirmed exploitation. Maintainer of Evals? Report an issue or request a rescan.




