API that compresses LLM context by up to 90% fewer tokens while maintaining answer accuracy.
Compresr is a drop-in API and SDK for reducing LLM token usage before sending prompts to any model. It performs question-aware compression that keeps only tokens relevant to the user query, achieving up to 90% token reduction with maintained or improved accuracy at light compression. Two deployment options are offered: a cloud SDK with per-million-token pricing and a private deployment inside the customer's own VPC with custom SLAs. Supported clients include Python 3.8+ and Node.js 18+. Compresr is a product of Compresr.
For people
For agents
Nothing listed yet.