Book a 30-minute demo →
Skill scan report

Evals

View on GitHub
20 Low Automated analysis flagged 2 potential risk patterns.

What this skill does

AI evaluation framework for testing and benchmarking AI agents

github/danielmiessler - Excessive Agency - 17.9k stars

ThreatsExcessive Agency

Threat analysis

Excessive Agency2 findings

Skill info

Namedanielmiessler/evals
Registrygithub
Version58381b3
PURLpkg:github/danielmiessler/LifeOS@58381b3?skill=evals
Stars17.9k

Assessments (2)

Excessive Agency2 findings MEDIUM
MEDIUM

Excessive Agency via local-llm-review

Workflows/CreateJudge.md

The workflow allows creation of custom LLM-as-Judge configurations, which could be misused to create biased or malicious grading systems if not properly controlled.
MEDIUM

Excessive Agency via local-llm-review

Workflows/CreateUseCase.md

The workflow enables creation of new evaluation use cases with test cases and scoring criteria, which could be used to create misleading or harmful evaluation scenarios if not properly validated.

Badge

Add the Anomity scan badge for Evals to your README.

Anomity Skill Check badge

Markdown
[![Anomity Skill Check](https://anomity.ai/skills/badge.svg)](https://anomity.ai/skills/github/danielmiessler/evals/)
HTML
<a href="https://anomity.ai/skills/github/danielmiessler/evals/"><img src="https://anomity.ai/skills/badge.svg" alt="Anomity Skill Check"></a>
Image URL
https://anomity.ai/skills/badge.svg

How Anomity governs this at runtime

Scan-time vetting tells you what a skill says it will do. Anomity's Endpoint Sensor sees what agents actually do: it discovers skills alongside every other AI artifact on the endpoint, and runtime governance can allow, deny, or log the tool calls a skill triggers. Policy violations route to your SIEM, Slack, email, or Jira, backed by a queryable 90-day audit trail.

Book a 30-minute demo to see your own skill inventory.

Methodology and disputes

Every skill is assessed by the Anomity Skill Intelligence engine against its public source; findings indicate risk patterns, not confirmed exploitation. Maintainer of Evals? Report an issue or request a rescan.

Ask AI about Anomity
ChatGPT Claude Perplexity Google AI Grok