Skip to main content
Back to all repositories
EL

craftbench

Benchmarks for LLMs on real software, data science, ML and AI engineering work. 20 eval categories, deterministic grading, ~$1.60 per full sweep. Runs inside Claude Code.

0stars0forks0watchers/subscribers0issues
Language
Python
Size
495 KB
Created
Aug 5, 2026
Last Updated
Aug 5, 2026
Last Pushed
Aug 5, 2026

Available Plugins

Loading plugins...

Evaluate before installing

  1. Review the source repository, recent maintenance, and license on GitHub.
  2. Read the marketplace manifest and plugin source files before running commands.
  3. Start with the smallest required permission set and validate behavior in a safe environment.