Cache-Preserving Zero-Loss Slimming,
The New Standard for Multi-Agent AI
Safely maintain 100% of Frontier LLM Prompt Cache Hits (50~90% off) across OpenAI, Anthropic, DeepSeek, and vLLM, while trimming dynamic tool payloads and conversational turns with sub-0.08ms lossless pruning for maximum efficiency.
Zero-Rewrite Transparent In-Memory Proxy
CDX A2A Slimmer sits transparently between your multi-agent applications and upstream LLM provider endpoints. Without modifying existing application code, all agent traffic is losslessly pruned in <0.08ms.
100% Native Compatibility with Latest Frontier LLMs, Agent Platforms & Inference Engines
2-Tier Prefix-Tail Split: Eliminating Cache Invalidation for 93% Net Savings
Traditional token reduction methods lacking prompt cache awareness alter prefix structures across turns, triggering cache misses that invalidate provider discounts and can increase net bills by up to 2.5x. CDX A2A Slimmer resolves this through an architectural 2-Tier split, preserving 100% KV-Cache Hits (50~90% baseline discount) while delivering 30% lossless slimming on dynamic turns simultaneously.
- Scope: System Prompts, Tool / MCP Definition Schemas.
- Mechanism: Deep canonical lexicographical parameter sorting, locking serialized SHA256 hashes byte-for-byte.
- Impact: 100% Prompt Cache Hit Rate on Anthropic, OpenAI, DeepSeek, and vLLM (50~90% baseline discount preserved).
- Scope: A2A Intermediate Turns, CoT reasoning, verbose tool returns.
- Mechanism: Boilerplate greeting stripping, whitespace pruning, AST metadata cleanup (code & numbers 100% locked).
- Impact: 30% lossless payload trimming beyond the cache boundary, maintaining maximum cost efficiency alongside full cache discounts.
• Cache Threshold Safeguard (Over-Slimming Prevention): Compressing static schemas without cache awareness below 1,024 tokens causes an immediate forfeiture of the provider cache discount.
• CDX Safe Guardrails: Static prefixes are preserved above 1,024 tokens (sorting keys losslessly), while 30% slimming is focused strictly on the dynamic conversation tail to eliminate cache dropouts.
1. CDX A2A Enterprise Sidecar (Production Servers & K8s Pods)
An ultra-lightweight 50MB in-memory proxy designed to accelerate autonomous multi-agent communications within private enterprise VPCs, cloud Kubernetes pods, and Docker containers.
Prunes full end-to-end multi-turn dialogue traffic across LangGraph, CrewAI, and AutoGen swarms losslessly by 20%~31.55% (average ~30%) in < 0.08ms.
HTTP REST / SSE proxy (Port :8080). Simply set your agent application's base_url to the CDX sidecar to begin instant acceleration with zero code rewrites.
Generated Python/SQL code, UUIDs, and financial figures are strictly 100% preserved under RFC 8259 invariance while superfluous CoT boilerplate is pruned, converging naturally to 30% savings.
Streams per-agent token reduction and P99 latency metrics to Prometheus, Datadog, and Grafana via /metrics, while maintaining strictly isolated enterprise audit logs.
2. CDX MCP Stdio Streaming Proxy (Developer Desktop & Local IDEs)
An ultra-low latency stdio JSON-RPC 2.0 streaming proxy tailored for Claude Desktop, Cursor, Windsurf, and Zed IDE developers.
Prunes repeated, token-heavy tool definition schema metadata by up to 40%~50% in < 0.05ms, dramatically lowering local AI development API costs.
Prefix existing tool commands in claude_desktop_config.json or .cursor/mcp.json with python -m cdx_a2a_slimmer mcp-proxy --. Setup is instant with zero configuration.
Strips redundant $schema, title, additionalProperties, and UI formatting tags from PostgreSQL and GitHub tools via P vs NP AST Pruning (required properties 100% intact).
Bypasses corrupted payloads without alteration to eliminate runtime crashes, while emitting real-time token reduction telemetry to sys.stderr for clean observability.
Measured Telemetry Benchmarks & Interactive Verification
Examine verified reduction telemetry across realistic production workloads, and directly input raw JSON payloads into the interactive playground to test live trimming.
① Complex Agents (Tool Calling + CoT)
CrewAI-based research and report synthesis multi-agent crew workload
② Code Review Crew (AutoGen)
Python source code analysis, lint inspection, and patch proposal workload
③ Enterprise SQL & Big Data
SQL queries and enterprise table schemas with strict literal byte lock
10-Second Drop-in Integration
No need to refactor existing model architectures or agent loops. Run the lightweight 50MB Docker sidecar container or wrap requests with the 1-line Python SDK.
Terminal Deployment & Package Installation (10 Seconds)
1-Line Integration in Application Code
from cdx_a2a_slimmer import slim
# Raw multi-agent payload (OpenAI / Claude / DeepSeek schema)
raw_payload = {
"model": "gpt-4o",
"tools": [{"type": "function", "function": {"name": "query_db", "parameters": {...}}}],
"messages": [{"role": "user", "content": "Hello! As an AI agent, please execute SQL."}]
}
# 1-line lossless slimming (latency < 0.08ms, 31.55% token reduction)
slimmed_payload = slim(raw_payload)Official Global Registries & Package Downloads
pip install cdx-a2a-slimmer
docker pull cdxno1/cdx-a2a-sidecar:latest
git clone https://github.com/oddkyu/cdx-a2a-slimmer.git
Enterprise Lossless Governance & Reliability
Unlike stochastic LLM summarizers, our deterministic AST engine prevents syntax corruption and parameter distortion, guaranteeing complete enterprise data integrity.
Source Code & Literal Data Lock
Financial digits, IDs, SQL statements, and Python/C++ identifiers are locked byte-for-byte without alteration.
JSON Schema & Type Conformity
Strict RFC 8259 compliance ensuring `required`, `properties`, and type specifications are permanently preserved.
Context Noise Reduction & Attention Optimization
Prunes AST noise and boilerplate tokens to optimize Transformer attention efficiency and improve multi-agent response quality.
Workload-based Estimated ROI Simulation & Pricing
Simulate your monthly net savings range (20% to up to 31.55%) based on API traffic volume, then select the optimal tier engineered for your infrastructure scale.
• Simulated figures represent an estimated range based on empirical multi-agent benchmarks (20.0% to maximum 31.55% token reduction across Tool Calling, CoT traces, and Multi-turn dialogues).
• Actual enterprise savings vary depending on prompt density, tool schema complexity, dialogue turn depth, and target model pricing tiers (GPT-5, Claude 3.7, DeepSeek-R1, etc.).
• To evaluate your specific workload telemetry, run test payloads in the Live Playground or request an Enterprise PoC.
Developer
Individual developers, prototypes, and open-source research (PoC)
- ✓ 1 Local Machine (Dev Environment)
- ✓ Up to 5 Concurrent Agent Sessions
- ✓ 100k Monthly Slimming Calls
- ✓ MCP Stdio 1-Click Proxy (Cursor / Claude)
- ✓ 20% ~ Up to 31.5% Lossless Token Optimization & Full KV-Cache Alignment Included
- ✓ Community Discord Support
Team
Fast-growing startups and dedicated multi-agent services
- ✓ Up to 5 Nodes (Pod) Cluster
- ✓ Up to 50 Concurrent Agent Sessions
- ✓ 1M Monthly Slimming Calls (~$2.2k Savings)
- ✓ 10 Team Seats for MCP (Local & Remote SSE)
- ✓ ~20%~30% Lossless Token Pruning
- ✓ Prometheus Telemetry Endpoint
Business
Mid-market, high-traffic B2B SaaS, and agent platforms
- ✓ Up to 20 Nodes (Pod) Cluster
- ✓ Up to 200 Concurrent Agent Sessions
- ✓ 5M Monthly Slimming Calls
- ✓ Central Enterprise MCP Gateway & Cache
- ✓ K8s Autoscaling & Multi-VPC
- ✓ Dedicated Priority Support (4h SLA)
Enterprise
Fortune 500, Banks, Defense, and Air-Gapped Closed Networks
- ✓ Custom Node Cluster (50~100+ Pods)
- ✓ Unlimited Concurrent Agent Swarms
- ✓ 15M+ Monthly Slimming Volume
- ✓ Air-Gapped On-Premises MCP Hub & Audit Logs
- ✓ PII Zero-Trust Anonymization Vault
- ✓ Enterprise SSO (SAML/LDAP) & 24/7 Hotline
Engineering & Architecture FAQ
Clear technical answers regarding ultra-low latency benchmarks, lossless parsing fidelity, security compliance, and proxy integration.
required fields, properties, enums, and data types are strictly preserved. Verified across 500+ multi-tool benchmark testbeds with 100.0% tool calling accuracy and schema compliance.
base_url="http://localhost:8080/v1", or wrap payloads in Python with slim(payload). It works drop-in with all standard OpenAI/Anthropic SDKs.
2-Tier Prefix-Tail Split to canonicalize tool schemas and system instructions with strict lexicographical determinism, maintaining byte-level invariance. Developers retain full upstream prompt caching discounts while enjoying 30% lossless slimming on dynamic conversation payloads.