Python SDK & API Reference
Official enterprise reference for the cdx-a2a-slimmer package, including installation, synchronous & async pipelines, constructor parameters, real-time telemetry, Pydantic integration, and fail-open exception handling.
1. Installation & Environment Requirements
Compatible with Python 3.9+ environments, powered by a zero-dependency, ultra-lightweight C/Python AST engine.
# Install official PyPI package
pip install cdx-a2a-slimmer
2. API Key Authentication & Security Governance
CDX Sidecar Proxy and SDK enforce Zero-Trust client authentication via cryptographically generated Secret API Keys (cdx_live_sk_...).
•
Key Issuance: Generated from
Developer Console > 5. Proxy & API Key Governance with zero-downtime rolling.
•
HTTP Auth Header: Standard RFC 6750 Bearer token format:
Authorization: Bearer cdx_live_sk_....
•
Security Best Practices: Never commit API keys into front-end code or public repositories. Store them in
.env or an
Enterprise Secret Vault (AWS Secrets Manager, HashiCorp Vault, Azure Key Vault).
# [1] Set Environment Variables (.env)
export CDX_API_KEY="cdx_live_sk_your_actual_key_here" # 👈 Replace with actual Secret Key from Console
export CDX_PROXY_URL="http://127.0.0.1:8090/v1" # 👈 Local Sidecar or VPC Private Endpoint
# [2] Initialize Python SDK with Environment Variable
import os
from cdx_a2a_slimmer import CDXA2ATokenSlimmer
slimmer = CDXA2ATokenSlimmer(
license_key=os.getenv("CDX_API_KEY")
)
3. Synchronous Base Invocation: slim(payload: dict) -> dict
Wrap existing agent request dictionaries with slim() for sub-0.08ms lossless pruning before dispatching to LLMs.
from cdx_a2a_slimmer import slim
from openai import OpenAI
client = OpenAI()
# [1] Standard multi-agent request dictionary
payload = {
"model": "gpt-5-preview",
"messages": [
{"role": "system", "content": "You are an enterprise code reviewer agent."},
{"role": "user", "content": "Analyze the pull request changes."}
],
"tools": my_enterprise_tools # List of enterprise tool/MCP schemas
}
# [2] Wrap with slim() ➔ 0.08ms lossless pruning (20~30% token savings)
slimmed_payload = slim(payload)
response = client.chat.completions.create(**slimmed_payload)
print(response.choices[0].message.content)
4. Asynchronous High-Throughput Pipeline: async_slim(payload: dict) -> dict
Non-blocking AST execution optimized for FastAPI, Asyncio, and LangGraph multi-agent clusters.
import asyncio
from cdx_a2a_slimmer import async_slim
from openai import AsyncOpenAI
client = AsyncOpenAI()
async def run_agent_workflow(payload: dict):
# [1] Non-blocking asynchronous AST slimming
slimmed = await async_slim(payload)
# [2] Real-time streaming LLM dispatch
stream = await client.chat.completions.create(**slimmed, stream=True)
async for chunk in stream:
content = chunk.choices[0].delta.content or ""
print(content, end="", flush=True)
# Run event loop
asyncio.run(run_agent_workflow(payload))
5. Core Engine Class: CDXA2ATokenSlimmer
Initializes a high-concurrency in-memory parser instance with license verification and proprietary schema whitelisting.
from cdx_a2a_slimmer import CDXA2ATokenSlimmer
# [1] Initialize dedicated in-memory slimmer instance
slimmer = CDXA2ATokenSlimmer(
license_key="cdx_live_enterprise_sk_...", # None for Open Beta / Developer Tier
auto_canonicalize=True, # RFC 8259 deterministic schema sorting
preserve_keys=["company_auth_id"], # Whitelisted keys exempt from pruning
strip_docstring_redundancy=True # Lossless AST pruning of repetitive descriptions
)
# [2] Execute slimming via instance method
slimmed_payload = slimmer.slim(payload)
| Constructor Parameter | Type | Default | Description |
| license_key | Optional[str] | None | Commercial tier Ed25519 authentication key (Defaults to Free Tier if None) |
| auto_canonicalize | bool | True | Sorts tool definitions deterministically in lexicographical order for RFC 8259 integrity |
| preserve_keys | Optional[List[str]] | [] | List of internal company parameter keys 100% exempt from docstring reduction |
| strip_docstring_redundancy | bool | True | Prunes conversational chatter in tool docstrings while preserving exact types & constraints |
6. Real-Time Telemetry Retrieval: slimmer.slim_payload()
Returns both the slimmed payload and detailed telemetry (latency, raw/slimmed bytes, reduction ratio).
# [1] Extract slimmed payload and telemetry dictionary
slimmed_payload, telemetry = slimmer.slim_payload(raw_payload)
# [2] Print empirical telemetry logs
print(f"• In-Memory Latency : {telemetry['latency_ms']} ms")
print(f"• Payload Size : {telemetry['raw_bytes']}B ➔ {telemetry['slimmed_bytes']}B")
print(f"• Token Savings : {telemetry['reduction_percent']}%")
print(f"• KV-Cache Invariant: {telemetry['kv_cache_compatible']} (Hash: {telemetry['kv_cache_prefix_hash']})")
print(f"• Net Dual Savings : {telemetry['dual_synergy_saving_percent']}%")
7. Multi-Tool Calling & Pydantic v2 Schema Integration
Production recipe for complex multi-agent workflows with 50+ enterprise tool schemas.
from pydantic import BaseModel, Field
from cdx_a2a_slimmer import slim
# [1] Define Pydantic v2 Tool Argument Model
class SQLQueryArgs(BaseModel):
query: str = Field(description="ANSI SQL Query string to run")
timeout_sec: int = Field(default=30, description="Max query execution timeout in seconds")
# [2] Register standard OpenAI Function Tool
sql_tool = {
"type": "function",
"function": {
"name": "execute_sql",
"description": "Executes analytics queries on enterprise warehouse.",
"parameters": SQLQueryArgs.model_json_schema()
}
}
# [3] Slim payload with full Pydantic type constraint preservation
payload = {
"model": "deepseek-r1",
"messages": [{"role": "user", "content": "Retrieve Q3 financial metrics"}],
"tools": [sql_tool]
}
safe_slimmed = slim(payload)
8. Exception Handling & Zero-Downtime Fail-Open Fallback
Guarantees 100% dispatch continuity by automatically falling back to raw payload if a malformed schema is encountered.
import logging
from cdx_a2a_slimmer import slim, CDXException
# [Fail-Open Pattern] Guaranteed 100% uptime with instant raw payload fallback
try:
ready_payload = slim(raw_payload)
except CDXException as err:
logging.warning(f"CDX Exception caught, falling back to raw payload: {err}")
ready_payload = raw_payload
response = client.chat.completions.create(**ready_payload)
9. 100% KV-Cache Retention & 2-Tier Architecture
Detects prefix boundaries and 1,024-token minimum cache thresholds on OpenAI, Anthropic, DeepSeek, and vLLM, safely retaining 100% prompt cache hits while trimming dynamic payloads by 30% losslessly.
# [KV-Cache Invariance & 1,024-Token Threshold Telemetry]
slimmed_payload, metrics = slimmer.slim_payload(raw_payload)
print(f"• KV-Cache Compatible : {metrics['kv_cache_compatible']}")
print(f"• Prefix Invariant Hash: {metrics['kv_cache_prefix_hash']}")
print(f"• Prefix Token Estimate: {metrics['kv_cache_prefix_tokens']} tokens")
print(f"• 1,024T Threshold Met : {metrics['kv_cache_threshold_met']} (Status: {metrics['cache_discount_status']})")
print(f"• Net Dual Savings : {metrics['dual_synergy_saving_percent']}%")
200+ Agent Metadata Tagging & Modality Spec
Complete technical specifications on how CDX Proxy parses incoming HTTP requests into 【Node ID / Agent Name / Swarm Group / Target Model / Data Modality】 for real-time fleet observability.
1. Ingress Telemetry & Detection Matrix
| Fleet Field | Ingress Path | Auto Detected | Mechanism Description |
| Target Model |
HTTP Body JSON |
100% Automatic |
Directly extracted from {"model": "gpt-5-preview"} in sub-0.01ms |
| Data Modality |
CDX AST Engine |
100% Automatic |
Payload structure analyzed: CoT traces, Tool ASTs, Multi-Turn, Raw Markdown |
| Node ID & Name |
Header / User Field |
Explicit or Fallback |
Supplied via X-Agent-ID or user. Fallback assigns virtual session tag |
| Swarm Group |
Header / User Field |
Explicit or Fallback |
X-Swarm-Group (engineering, research, sre, crm). Auto-clustered if omitted |
2. Automatic Modality Classification Rules
• [CoT Reasoning Traces] (Avg 44.5% Reduction): Auto-tagged when reasoning_content or <thought> tags exist.
• [Tool Call & JSON AST] (Avg 32.1% Reduction): Auto-tagged when tools, tool_calls, function_call schemas are present.
• [Multi-Turn Context] (Avg 28.4% Reduction): Auto-tagged when messages length >= 6.
• [Raw Text & Markdown] (Avg 21.2% Reduction): Standard single-turn prompt and completion interactions.
3. Agent Fleet Tagging Code Examples
[Python SDK / OpenAI Client] Custom headers or user parameter injection:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8080/v1",
api_key="cdx_live_enterprise_..."
)
# Method 1: Explicit Swarm Tagging via extra_headers (Recommended)
response = client.chat.completions.create(
model="gpt-5-preview",
messages=[{"role": "user", "content": "Execute AST validation"}],
extra_headers={
"X-Agent-ID": "node-042",
"X-Agent-Name": "Code-Reviewer",
"X-Swarm-Group": "engineering" # engineering, research, sre, crm
}
)
# Method 2: Standard OpenAI user parameter (colon-delimited)
response = client.chat.completions.create(
model="claude-3.7-sonnet",
messages=[{"role": "user", "content": "SRE Health check"}],
user="sre:node-014:LeadSRE"
)
[cURL / HTTP Direct] Standard HTTP headers:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer cdx_live_enterprise_..." \
-H "X-Agent-ID: node-089" \
-H "X-Agent-Name: SQL-Synthesizer" \
-H "X-Swarm-Group: engineering" \
-d '{
"model": "deepseek-r1",
"messages": [{"role": "user", "content": "SELECT * FROM telemetry"}]
}'
Multi-Agent Framework Recipes
Ready-to-use recipes for LangGraph, CrewAI, AutoGen, LlamaIndex, and internal vLLM clusters with 1-line proxy routing.
1. LangGraph Enterprise StateGraph
from langchain_openai import ChatOpenAI
from langgraph.graph import StateGraph, END
llm = ChatOpenAI(
model="gpt-5-preview",
base_url="http://localhost:8080/v1",
api_key="sk-...",
default_headers={"X-Agent-ID": "node-001", "X-Swarm-Group": "engineering"}
)
2. CrewAI Multi-Agent Swarm
from crewai import Agent, Crew
from langchain_openai import ChatOpenAI
slimmer_llm = ChatOpenAI(
model="claude-3.7-sonnet",
base_url="http://localhost:8080/v1",
default_headers={"X-Agent-ID": "node-002", "X-Swarm-Group": "research"}
)
researcher = Agent(role="Senior Market Analyst", goal="Analyze trends", llm=slimmer_llm)
3. Microsoft AutoGen Multi-Agent Dialog
import autogen
llm_config = {
"config_list": [{
"model": "deepseek-r1",
"base_url": "http://localhost:8080/v1",
"api_key": "sk-...",
"extra_headers": {"X-Agent-ID": "node-003", "X-Swarm-Group": "sre"}
}]
}
4. Anthropic MCP (Model Context Protocol) 1-Click Drop-in Proxy
When connecting 3+ MCP servers (PostgreSQL, GitHub, Filesystem) to Claude Desktop, Cursor, or Windsurf, CDX intercepts JSON-RPC tools/list schemas and slims bloated metadata by up to 40%~50% in < 0.05ms (20%~30% overall end-to-end token reduction) losslessly.
Option A. Smithery.ai 1-Click Automated Install (Single Terminal Command)
# Claude Desktop Automated Registration
npx -y @smithery/cli install earthtbook/cdx-mcp-proxy --client claude
# Cursor Automated Registration
npx -y @smithery/cli install earthtbook/cdx-mcp-proxy --client cursor
Option B. Zero-Terminal Direct Copy-Paste Configuration (claude_desktop_config.json)
Config Paths: Windows(%APPDATA%\Claude\claude_desktop_config.json) / macOS(~/Library/Application Support/Claude/claude_desktop_config.json)
{
"mcpServers": {
"postgres-slimmer": {
"command": "python",
"args": [
"-m", "cdx_a2a_slimmer", "mcp-proxy", "--",
"npx", "-y", "@modelcontextprotocol/server-postgres", "postgresql://localhost/mydb"
]
},
"filesystem-slimmer": {
"command": "python",
"args": [
"-m", "cdx_a2a_slimmer", "mcp-proxy", "--",
"npx", "-y", "@modelcontextprotocol/server-filesystem", "/Users/workspace"
]
}
}
}
Option C. Zero-Install Cloud Remote Proxy via Smithery (run.tools)
{
"mcpServers": {
"cdx-mcp-proxy": {
"url": "https://cdx-mcp-proxy--earthtbook.run.tools"
}
}
}
Kubernetes & Docker Sidecar Deployment
Official enterprise deployment manifests for running CDX as a zero-code, network-transparent 50MB micro-container sidecar in Kubernetes Pods and cluster-wide DaemonSets.
🛡️ 100% Zero-Downtime Guarantee (Fail-Open Graceful Pass-Through Architecture)
If incoming requests contain non-standard JSON, binary multipart chunks, or malformed payloads, the CDX proxy instantly bypasses slimming in < 0.00ms and transparently forwards the untouched raw payload directly to upstream LLMs. Client workflows maintain uninterrupted uptime without failure risk.
1. Docker Single Container Run (Dev & Standalone)
Launch the CDX sidecar proxy instantly on any local developer workstation or VM instance.
docker run -d \
--name cdx-slimmer \
-p 8080:8080 \
-e TARGET_UPSTREAM="https://api.openai.com" \
-e CDX_LICENSE_KEY="cdx_live_sk_your_actual_key_here" \
-e LOG_LEVEL="info" \
--restart unless-stopped \
cdxno1/cdx-a2a-sidecar:latest
2. Kubernetes Multi-Container Pod Sidecar Manifest (deployment.yaml)
Shares localhost:8080 networking within the same Pod for sub-0.05ms in-memory proxying without code modifications.
apiVersion: apps/v1
kind: Deployment
metadata:
name: enterprise-agent-fleet
labels:
app: agent-worker
spec:
replicas: 3
selector:
matchLabels:
app: agent-worker
template:
metadata:
labels:
app: agent-worker
spec:
containers:
# [1] Primary AI Agent Application Container
- name: ai-agent-app
image: mycorp/agent-service:v2.4.0
env:
- name: OPENAI_BASE_URL
value: "http://127.0.0.1:8080/v1" # 👈 Transparently route to local sidecar loopback
- name: OPENAI_API_KEY
valueFrom:
secretKeyRef:
name: cdx-credentials
key: api-key
resources:
requests:
cpu: "500m"
memory: "512Mi"
# [2] CDX Ultralight AST Token Slimmer Sidecar (50MB)
- name: cdx-sidecar
image: cdxno1/cdx-a2a-sidecar:latest
imagePullPolicy: IfNotPresent
ports:
- containerPort: 8080
name: proxy-http
env:
- name: TARGET_UPSTREAM
value: "https://api.openai.com"
- name: CDX_LICENSE_KEY
valueFrom:
secretKeyRef:
name: cdx-credentials
key: license-key
- name: ENABLE_PROMETHEUS
value: "true"
resources:
requests:
cpu: "50m"
memory: "32Mi"
limits:
cpu: "200m"
memory: "64Mi"
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 3
periodSeconds: 10
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 2
periodSeconds: 5
3. Docker Compose Local Multi-Container Stack (docker-compose.yml)
One-click local orchestration of agent applications with the CDX proxy sidecar.
version: '3.8'
services:
# CDX Sidecar Proxy
cdx-proxy:
image: cdxno1/cdx-a2a-sidecar:v1.0.0-beta
container_name: cdx-slimmer-proxy
ports:
- "8080:8080"
environment:
- TARGET_UPSTREAM=https://api.openai.com
- CDX_LICENSE_KEY=${CDX_API_KEY}
- LOG_LEVEL=info
restart: always
# Internal AI Agent Service
agent-service:
build: .
depends_on:
- cdx-proxy
environment:
- OPENAI_BASE_URL=http://cdx-proxy:8080/v1
- OPENAI_API_KEY=${CDX_API_KEY}
4. Sidecar Environment Variables & Resource Reference
Standard configuration environment variables injected into the sidecar container runtime.
| Environment Variable | Default | Required | Description |
| TARGET_UPSTREAM | https://api.openai.com | Required | Upstream LLM provider URL (OpenAI, Anthropic, Gemini, internal vLLM) |
| CDX_LICENSE_KEY | None (Open Beta) | Optional | Enterprise license key (Defaults to Open Beta free tier if omitted) |
| CDX_PORT | 8080 | Optional | Sidecar inbound listening TCP port |
| ENABLE_PROMETHEUS | true | Optional | Exposes Prometheus time-series metrics on /metrics |
| LOG_LEVEL | info | Optional | Logging verbosity (debug, info, warn, error) |
5. Cloudflare Pages & Supabase Production Cloud Architecture
The CDX portal and global APIs run on a high-availability serverless topology powered by Cloudflare Anycast Edge CDN/WAF and Supabase Multi-Tenant PostgreSQL (RLS).
• Edge Layer (Global CDN & WAF): Cloudflare Pages with 330+ Anycast PoPs guarantees < 15ms TTFB globally, with Always-On L3/L4/L7 DDoS mitigation and Bot Fight Mode.
• IAM & Ledger Layer: Supabase Auth (FIDO2/Passkey, GitHub/Google OAuth, SAML SSO) with Strict PostgreSQL Row Level Security (RLS) guarantees 100% cryptographic tenant isolation.
• Zero Data Retention (ZDR): Customer prompts are processed in-memory and never written to disk or database. Only mathematical token metrics and HMAC billing receipts are securely recorded.
# [Step 1] Push Supabase PostgreSQL Schema with Strict RLS
supabase db push --file supabase_schema.sql
# [Step 2] 1-Click Global Edge Deployment to Cloudflare Pages (cdxengine.com)
npx wrangler pages deploy . --project-name=cdx-engine-portal --branch=main
Prometheus, Grafana & OpenTelemetry Observability
Official enterprise telemetry reference for piping real-time token savings, P99 in-memory latency, active sessions, and swarm health metrics to Grafana, Datadog, and OpenTelemetry via GET /metrics.
1. Prometheus Scrape Configuration (prometheus.yml)
Configure Prometheus to scrape the CDX sidecar proxy at port 8080 on a 10s interval.
scrape_configs:
- job_name: 'cantor-cdx-200-fleet'
scrape_interval: 10s
metrics_path: '/metrics'
static_configs:
- targets: ['localhost:8080']
labels:
environment: 'production'
region: 'us-east-1'
2. Core Exported Prometheus Metrics Reference
The 5 core time-series telemetry metrics exposed by the CDX in-memory engine.
| Metric Name | Type | Key Labels | Description & Unit |
| cdx_tokens_pruned_total |
Counter |
swarm, model, modality |
Cumulative count of lossless pruned tokens (Tokens) |
| cdx_reduction_ratio |
Gauge |
swarm, agent_id |
Real-time token reduction ratio (0.00 ~ 1.00, verified 0.20 ~ 0.3155) |
| cdx_proxy_latency_seconds |
Histogram |
le, endpoint |
CDX in-memory slimming latency (P50, P90, P99 in seconds) |
| cdx_active_sessions |
Gauge |
swarm |
Number of currently active concurrent agent sessions |
| cdx_requests_total |
Counter |
status, model |
Total ingress HTTP proxy requests and status codes (200, 4xx, 5xx) |
3. Ready-to-use Grafana PromQL Queries
Drop-in PromQL snippets for Grafana dashboard panels and chart visualization.
# 1) Token pruning velocity (Tokens / sec)
sum(rate(cdx_tokens_pruned_total[5m])) by (swarm)
# 2) P99 In-Memory Proxy Latency (Converted to ms)
histogram_quantile(0.99, sum(rate(cdx_proxy_latency_seconds_bucket[5m])) by (le)) * 1000
# 3) Average token reduction percentage by Swarm cluster (%)
avg(cdx_reduction_ratio) by (swarm) * 100
# 4) Fleet-wide HTTP Request Success Rate (%)
sum(rate(cdx_requests_total{status="200"}[5m])) / sum(rate(cdx_requests_total[5m])) * 100
4. OpenTelemetry (OTel) Collector & Datadog Pipeline
Forward CDX metrics to Datadog or Dynatrace using the OpenTelemetry Collector.
receivers:
prometheus:
config:
scrape_configs:
- job_name: 'cdx-slimmer'
scrape_interval: 10s
static_configs:
- targets: ['cdx-slimmer.production:8080']
exporters:
datadog:
api:
key: "${DATADOG_API_KEY}"
site: "datadoghq.com"
service:
pipelines:
metrics:
receivers: [prometheus]
exporters: [datadog]
5. Prometheus Alerting Rules (alerting_rules.yml)
Instant alerting rules for P99 latency spikes (>1ms) or proxy downtime via Slack & PagerDuty.
groups:
- name: cdx_alert_rules
rules:
- alert: CDXProxyHighLatency
expr: histogram_quantile(0.99, sum(rate(cdx_proxy_latency_seconds_bucket[5m])) by (le)) > 0.001
for: 1m
labels:
severity: warning
annotations:
summary: "CDX Sidecar P99 latency exceeded 1.0ms (Baseline: 0.08ms)"
- alert: CDXSidecarDown
expr: up{job="cantor-cdx-200-fleet"} == 0
for: 30s
labels:
severity: critical
annotations:
summary: "CDX Sidecar container is unreachable on port 8080"
6. Grafana 1-Click Dashboard Template (Import JSON)
Ready-to-use Grafana JSON dashboard snippet for immediate import via Dashboards > Import:
{
"title": "Cantor CDX Token Slimmer Observability",
"panels": [
{
"title": "Token Pruning Rate (Tokens/sec)",
"type": "timeseries",
"targets": [{"expr": "sum(rate(cdx_tokens_pruned_total[5m]))"}]
},
{
"title": "P99 In-Memory Latency (ms)",
"type": "gauge",
"targets": [{"expr": "histogram_quantile(0.99, sum(rate(cdx_proxy_latency_seconds_bucket[5m])) by (le)) * 1000"}]
},
{
"title": "Real-time Reduction Ratio (%)",
"type": "stat",
"targets": [{"expr": "avg(cdx_reduction_ratio) * 100"}]
}
]
}
Version Matrix & Verified Model Support
Semantic Versioning v2.0.0 standards and verified global LLM compatibility matrix.
1. Official LLM Compatibility Matrix
| Model / Framework | Verified Version | Supported Optimizations |
| OpenAI GPT-5 / GPT-4.5 / o3 / o1 / GPT-4o | v1.0.0-beta+ | 100% Compatible (Thinking & Tool Calling) |
| Anthropic Claude 3.7 Sonnet / 3.5 | v1.0.0-beta+ | 100% Compatible (Hybrid Reasoning & Tools) |
| Google Gemini 2.0 / 3.0 Flash & Pro | v1.0.0-beta+ | 100% Compatible (2M+ Multimodal Calling) |
| DeepSeek-V3 / DeepSeek-R1 | v1.0.0-beta+ | 100% Compatible (Private GPU VRAM Pruning) |
2. Product Release Roadmap
| Version | Target Schedule | Core Architecture & Scope | Target Reduction |
| v1.0 (Current) | 2026.09 Live | Tool JSON Schema AST Pruning, Prefix Normalization, KV-Cache Hit Preservation | 30% ~ 35% |
| v2.0 (Upcoming) | 2026.Q4 Scheduled | Tool Output & Multi-Turn Large Context Pruning Intelligent lossless filtering for large SQL result sets, web scrape data, and long agent dialogues | 60% ~ 77%+ |
| v2.1 | 2027.Q1 Scheduled | Ultra-low latency in-memory engine optimization (<10µs latency) and cluster auto-scaling | 60%+ Maintained |