Skip to main content
Back to all repositories

Context optimization layer for LLMs. 65-90% token savings with zero quality loss. Drop-in proxy for Claude, GPT, Gemini, and local LLMs (Ollama/VLLM/llama.cpp). Features KV cache-aware compression, session dedup, error cards, and Pichay-proven context paging.

19stars1forks0watchers/subscribers6issues
agentaianthropicclaude-codecontext-engineeringcontext-windowlangchainllmmcpollamaopencodeprompt-engineeringproxypythontoken-optimization
Language
Python
License
Apache License 2.0
Size
49.1 MB
Created
Jun 20, 2026
Last Updated
Sep 3, 2026
Last Pushed
Aug 22, 2026

Available Plugins

Loading plugins...

Evaluate before installing

  1. Review the source repository, recent maintenance, and license on GitHub.
  2. Read the marketplace manifest and plugin source files before running commands.
  3. Start with the smallest required permission set and validate behavior in a safe environment.