Skip to main content
Back to all repositories
MI

deepeval-agent-reliability

Tool-calling reliability evaluation pipeline for AI agents — extends DeepEval with deterministic metrics, 8-category failure attribution, and baseline-vs-enhanced benchmarking.

0stars0forks0watchers/subscribers0issues
llm-evaluation-agent-tool-calling-python
Language
Python
License
Apache License 2.0
Size
19.6 MB
Created
Jul 2, 2026
Last Updated
Jul 10, 2026
Last Pushed
Jul 10, 2026

Available Plugins

Loading plugins...

Evaluate before installing

  1. Review the source repository, recent maintenance, and license on GitHub.
  2. Read the marketplace manifest and plugin source files before running commands.
  3. Start with the smallest required permission set and validate behavior in a safe environment.