Skip to main content
Back to all repositories

Evaluation discipline for agentic apps — a vendored, local-first LLM-evaluation framework operated through your coding agent, with typed authority and a single Verdict Engine that gates only what you approve. The pure core; you decide what to measure and it derives no metrics on its own.

0stars0forks0watchers/subscribers2issues
agenticai-safetybaselinecalibrationci-gatingclaude-codecodexevalsevaluation-frameworkllmllm-evaluationlocal-firstpythonscorecardverdict
Language
Python
License
Apache License 2.0
Size
2.4 MB
Created
Aug 20, 2026
Last Updated
Aug 31, 2026
Last Pushed
Aug 31, 2026

Available Plugins

Loading plugins...

Evaluate before installing

  1. Review the source repository, recent maintenance, and license on GitHub.
  2. Read the marketplace manifest and plugin source files before running commands.
  3. Start with the smallest required permission set and validate behavior in a safe environment.