OFFICIAL DEVELOPER DOCUMENTATION & API REFERENCE

Official Documentation & API Reference

Comprehensive technical guides for CDX A2A Slimmer SDK, 200+ Agent Fleet metadata tagging, AST modality classification, Kubernetes sidecar deployment, and Prometheus observability.

Doc Version: v1.0.0-beta Latest Release

Python SDK & API Reference

Official enterprise reference for the cdx-a2a-slimmer package, including installation, synchronous & async pipelines, constructor parameters, real-time telemetry, Pydantic integration, and fail-open exception handling.

1. Installation & Environment Requirements

Compatible with Python 3.9+ environments, powered by a zero-dependency, ultra-lightweight C/Python AST engine.

# Install official PyPI package pip install cdx-a2a-slimmer

2. API Key Authentication & Security Governance

CDX Sidecar Proxy and SDK enforce Zero-Trust client authentication via cryptographically generated Secret API Keys (cdx_live_sk_...).

• Key Issuance: Generated from Developer Console > 5. Proxy & API Key Governance with zero-downtime rolling.
• HTTP Auth Header: Standard RFC 6750 Bearer token format: Authorization: Bearer cdx_live_sk_....
• Security Best Practices: Never commit API keys into front-end code or public repositories. Store them in .env or an Enterprise Secret Vault (AWS Secrets Manager, HashiCorp Vault, Azure Key Vault).
# [1] Set Environment Variables (.env) export CDX_API_KEY="cdx_live_sk_your_actual_key_here" # 👈 Replace with actual Secret Key from Console export CDX_PROXY_URL="http://127.0.0.1:8090/v1" # 👈 Local Sidecar or VPC Private Endpoint # [2] Initialize Python SDK with Environment Variable import os from cdx_a2a_slimmer import CDXA2ATokenSlimmer slimmer = CDXA2ATokenSlimmer( license_key=os.getenv("CDX_API_KEY") )

3. Synchronous Base Invocation: slim(payload: dict) -> dict

Wrap existing agent request dictionaries with slim() for sub-0.08ms lossless pruning before dispatching to LLMs.

from cdx_a2a_slimmer import slim from openai import OpenAI client = OpenAI() # [1] Standard multi-agent request dictionary payload = { "model": "gpt-5-preview", "messages": [ {"role": "system", "content": "You are an enterprise code reviewer agent."}, {"role": "user", "content": "Analyze the pull request changes."} ], "tools": my_enterprise_tools # List of enterprise tool/MCP schemas } # [2] Wrap with slim() ➔ 0.08ms lossless pruning (20~30% token savings) slimmed_payload = slim(payload) response = client.chat.completions.create(**slimmed_payload) print(response.choices[0].message.content)

4. Asynchronous High-Throughput Pipeline: async_slim(payload: dict) -> dict

Non-blocking AST execution optimized for FastAPI, Asyncio, and LangGraph multi-agent clusters.

import asyncio from cdx_a2a_slimmer import async_slim from openai import AsyncOpenAI client = AsyncOpenAI() async def run_agent_workflow(payload: dict): # [1] Non-blocking asynchronous AST slimming slimmed = await async_slim(payload) # [2] Real-time streaming LLM dispatch stream = await client.chat.completions.create(**slimmed, stream=True) async for chunk in stream: content = chunk.choices[0].delta.content or "" print(content, end="", flush=True) # Run event loop asyncio.run(run_agent_workflow(payload))

5. Core Engine Class: CDXA2ATokenSlimmer

Initializes a high-concurrency in-memory parser instance with license verification and proprietary schema whitelisting.

from cdx_a2a_slimmer import CDXA2ATokenSlimmer # [1] Initialize dedicated in-memory slimmer instance slimmer = CDXA2ATokenSlimmer( license_key="cdx_live_enterprise_sk_...", # None for Open Beta / Developer Tier auto_canonicalize=True, # RFC 8259 deterministic schema sorting preserve_keys=["company_auth_id"], # Whitelisted keys exempt from pruning strip_docstring_redundancy=True # Lossless AST pruning of repetitive descriptions ) # [2] Execute slimming via instance method slimmed_payload = slimmer.slim(payload)
Constructor ParameterTypeDefaultDescription
license_keyOptional[str]NoneCommercial tier Ed25519 authentication key (Defaults to Free Tier if None)
auto_canonicalizeboolTrueSorts tool definitions deterministically in lexicographical order for RFC 8259 integrity
preserve_keysOptional[List[str]][]List of internal company parameter keys 100% exempt from docstring reduction
strip_docstring_redundancyboolTruePrunes conversational chatter in tool docstrings while preserving exact types & constraints

6. Real-Time Telemetry Retrieval: slimmer.slim_payload()

Returns both the slimmed payload and detailed telemetry (latency, raw/slimmed bytes, reduction ratio).

# [1] Extract slimmed payload and telemetry dictionary slimmed_payload, telemetry = slimmer.slim_payload(raw_payload) # [2] Print empirical telemetry logs print(f"• In-Memory Latency : {telemetry['latency_ms']} ms") print(f"• Payload Size : {telemetry['raw_bytes']}B ➔ {telemetry['slimmed_bytes']}B") print(f"• Token Savings : {telemetry['reduction_percent']}%") print(f"• KV-Cache Invariant: {telemetry['kv_cache_compatible']} (Hash: {telemetry['kv_cache_prefix_hash']})") print(f"• Net Dual Savings : {telemetry['dual_synergy_saving_percent']}%")

7. Multi-Tool Calling & Pydantic v2 Schema Integration

Production recipe for complex multi-agent workflows with 50+ enterprise tool schemas.

from pydantic import BaseModel, Field from cdx_a2a_slimmer import slim # [1] Define Pydantic v2 Tool Argument Model class SQLQueryArgs(BaseModel): query: str = Field(description="ANSI SQL Query string to run") timeout_sec: int = Field(default=30, description="Max query execution timeout in seconds") # [2] Register standard OpenAI Function Tool sql_tool = { "type": "function", "function": { "name": "execute_sql", "description": "Executes analytics queries on enterprise warehouse.", "parameters": SQLQueryArgs.model_json_schema() } } # [3] Slim payload with full Pydantic type constraint preservation payload = { "model": "deepseek-r1", "messages": [{"role": "user", "content": "Retrieve Q3 financial metrics"}], "tools": [sql_tool] } safe_slimmed = slim(payload)

8. Exception Handling & Zero-Downtime Fail-Open Fallback

Guarantees 100% dispatch continuity by automatically falling back to raw payload if a malformed schema is encountered.

import logging from cdx_a2a_slimmer import slim, CDXException # [Fail-Open Pattern] Guaranteed 100% uptime with instant raw payload fallback try: ready_payload = slim(raw_payload) except CDXException as err: logging.warning(f"CDX Exception caught, falling back to raw payload: {err}") ready_payload = raw_payload response = client.chat.completions.create(**ready_payload)

9. 100% KV-Cache Retention & 2-Tier Architecture

Detects prefix boundaries and 1,024-token minimum cache thresholds on OpenAI, Anthropic, DeepSeek, and vLLM, safely retaining 100% prompt cache hits while trimming dynamic payloads by 30% losslessly.

# [KV-Cache Invariance & 1,024-Token Threshold Telemetry] slimmed_payload, metrics = slimmer.slim_payload(raw_payload) print(f"• KV-Cache Compatible : {metrics['kv_cache_compatible']}") print(f"• Prefix Invariant Hash: {metrics['kv_cache_prefix_hash']}") print(f"• Prefix Token Estimate: {metrics['kv_cache_prefix_tokens']} tokens") print(f"• 1,024T Threshold Met : {metrics['kv_cache_threshold_met']} (Status: {metrics['cache_discount_status']})") print(f"• Net Dual Savings : {metrics['dual_synergy_saving_percent']}%")

200+ Agent Metadata Tagging & Modality Spec

Complete technical specifications on how CDX Proxy parses incoming HTTP requests into 【Node ID / Agent Name / Swarm Group / Target Model / Data Modality】 for real-time fleet observability.

1. Ingress Telemetry & Detection Matrix

Fleet FieldIngress PathAuto DetectedMechanism Description
Target Model HTTP Body JSON 100% Automatic Directly extracted from {"model": "gpt-5-preview"} in sub-0.01ms
Data Modality CDX AST Engine 100% Automatic Payload structure analyzed: CoT traces, Tool ASTs, Multi-Turn, Raw Markdown
Node ID & Name Header / User Field Explicit or Fallback Supplied via X-Agent-ID or user. Fallback assigns virtual session tag
Swarm Group Header / User Field Explicit or Fallback X-Swarm-Group (engineering, research, sre, crm). Auto-clustered if omitted

2. Automatic Modality Classification Rules

• [CoT Reasoning Traces] (Avg 44.5% Reduction): Auto-tagged when reasoning_content or <thought> tags exist.
• [Tool Call & JSON AST] (Avg 32.1% Reduction): Auto-tagged when tools, tool_calls, function_call schemas are present.
• [Multi-Turn Context] (Avg 28.4% Reduction): Auto-tagged when messages length >= 6.
• [Raw Text & Markdown] (Avg 21.2% Reduction): Standard single-turn prompt and completion interactions.

3. Agent Fleet Tagging Code Examples

[Python SDK / OpenAI Client] Custom headers or user parameter injection:

from openai import OpenAI client = OpenAI( base_url="http://localhost:8080/v1", api_key="cdx_live_enterprise_..." ) # Method 1: Explicit Swarm Tagging via extra_headers (Recommended) response = client.chat.completions.create( model="gpt-5-preview", messages=[{"role": "user", "content": "Execute AST validation"}], extra_headers={ "X-Agent-ID": "node-042", "X-Agent-Name": "Code-Reviewer", "X-Swarm-Group": "engineering" # engineering, research, sre, crm } ) # Method 2: Standard OpenAI user parameter (colon-delimited) response = client.chat.completions.create( model="claude-3.7-sonnet", messages=[{"role": "user", "content": "SRE Health check"}], user="sre:node-014:LeadSRE" )

[cURL / HTTP Direct] Standard HTTP headers:

curl http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer cdx_live_enterprise_..." \ -H "X-Agent-ID: node-089" \ -H "X-Agent-Name: SQL-Synthesizer" \ -H "X-Swarm-Group: engineering" \ -d '{ "model": "deepseek-r1", "messages": [{"role": "user", "content": "SELECT * FROM telemetry"}] }'

Multi-Agent Framework Recipes

Ready-to-use recipes for LangGraph, CrewAI, AutoGen, LlamaIndex, and internal vLLM clusters with 1-line proxy routing.

1. LangGraph Enterprise StateGraph

from langchain_openai import ChatOpenAI from langgraph.graph import StateGraph, END llm = ChatOpenAI( model="gpt-5-preview", base_url="http://localhost:8080/v1", api_key="sk-...", default_headers={"X-Agent-ID": "node-001", "X-Swarm-Group": "engineering"} )

2. CrewAI Multi-Agent Swarm

from crewai import Agent, Crew from langchain_openai import ChatOpenAI slimmer_llm = ChatOpenAI( model="claude-3.7-sonnet", base_url="http://localhost:8080/v1", default_headers={"X-Agent-ID": "node-002", "X-Swarm-Group": "research"} ) researcher = Agent(role="Senior Market Analyst", goal="Analyze trends", llm=slimmer_llm)

3. Microsoft AutoGen Multi-Agent Dialog

import autogen llm_config = { "config_list": [{ "model": "deepseek-r1", "base_url": "http://localhost:8080/v1", "api_key": "sk-...", "extra_headers": {"X-Agent-ID": "node-003", "X-Swarm-Group": "sre"} }] }

4. Anthropic MCP (Model Context Protocol) 1-Click Drop-in Proxy

When connecting 3+ MCP servers (PostgreSQL, GitHub, Filesystem) to Claude Desktop, Cursor, or Windsurf, CDX intercepts JSON-RPC tools/list schemas and slims bloated metadata by up to 40%~50% in < 0.05ms (20%~30% overall end-to-end token reduction) losslessly.

Option A. Smithery.ai 1-Click Automated Install (Single Terminal Command)
# Claude Desktop Automated Registration npx -y @smithery/cli install earthtbook/cdx-mcp-proxy --client claude # Cursor Automated Registration npx -y @smithery/cli install earthtbook/cdx-mcp-proxy --client cursor
Option B. Zero-Terminal Direct Copy-Paste Configuration (claude_desktop_config.json)

Config Paths: Windows(%APPDATA%\Claude\claude_desktop_config.json) / macOS(~/Library/Application Support/Claude/claude_desktop_config.json)

{ "mcpServers": { "postgres-slimmer": { "command": "python", "args": [ "-m", "cdx_a2a_slimmer", "mcp-proxy", "--", "npx", "-y", "@modelcontextprotocol/server-postgres", "postgresql://localhost/mydb" ] }, "filesystem-slimmer": { "command": "python", "args": [ "-m", "cdx_a2a_slimmer", "mcp-proxy", "--", "npx", "-y", "@modelcontextprotocol/server-filesystem", "/Users/workspace" ] } } }
Option C. Zero-Install Cloud Remote Proxy via Smithery (run.tools)
{ "mcpServers": { "cdx-mcp-proxy": { "url": "https://cdx-mcp-proxy--earthtbook.run.tools" } } }

Kubernetes & Docker Sidecar Deployment

Official enterprise deployment manifests for running CDX as a zero-code, network-transparent 50MB micro-container sidecar in Kubernetes Pods and cluster-wide DaemonSets.

🛡️ 100% Zero-Downtime Guarantee (Fail-Open Graceful Pass-Through Architecture) If incoming requests contain non-standard JSON, binary multipart chunks, or malformed payloads, the CDX proxy instantly bypasses slimming in < 0.00ms and transparently forwards the untouched raw payload directly to upstream LLMs. Client workflows maintain uninterrupted uptime without failure risk.

1. Docker Single Container Run (Dev & Standalone)

Launch the CDX sidecar proxy instantly on any local developer workstation or VM instance.

docker run -d \ --name cdx-slimmer \ -p 8080:8080 \ -e TARGET_UPSTREAM="https://api.openai.com" \ -e CDX_LICENSE_KEY="cdx_live_sk_your_actual_key_here" \ -e LOG_LEVEL="info" \ --restart unless-stopped \ cdxno1/cdx-a2a-sidecar:latest

2. Kubernetes Multi-Container Pod Sidecar Manifest (deployment.yaml)

Shares localhost:8080 networking within the same Pod for sub-0.05ms in-memory proxying without code modifications.

apiVersion: apps/v1 kind: Deployment metadata: name: enterprise-agent-fleet labels: app: agent-worker spec: replicas: 3 selector: matchLabels: app: agent-worker template: metadata: labels: app: agent-worker spec: containers: # [1] Primary AI Agent Application Container - name: ai-agent-app image: mycorp/agent-service:v2.4.0 env: - name: OPENAI_BASE_URL value: "http://127.0.0.1:8080/v1" # 👈 Transparently route to local sidecar loopback - name: OPENAI_API_KEY valueFrom: secretKeyRef: name: cdx-credentials key: api-key resources: requests: cpu: "500m" memory: "512Mi" # [2] CDX Ultralight AST Token Slimmer Sidecar (50MB) - name: cdx-sidecar image: cdxno1/cdx-a2a-sidecar:latest imagePullPolicy: IfNotPresent ports: - containerPort: 8080 name: proxy-http env: - name: TARGET_UPSTREAM value: "https://api.openai.com" - name: CDX_LICENSE_KEY valueFrom: secretKeyRef: name: cdx-credentials key: license-key - name: ENABLE_PROMETHEUS value: "true" resources: requests: cpu: "50m" memory: "32Mi" limits: cpu: "200m" memory: "64Mi" livenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 3 periodSeconds: 10 readinessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 2 periodSeconds: 5

3. Docker Compose Local Multi-Container Stack (docker-compose.yml)

One-click local orchestration of agent applications with the CDX proxy sidecar.

version: '3.8' services: # CDX Sidecar Proxy cdx-proxy: image: cdxno1/cdx-a2a-sidecar:v1.0.0-beta container_name: cdx-slimmer-proxy ports: - "8080:8080" environment: - TARGET_UPSTREAM=https://api.openai.com - CDX_LICENSE_KEY=${CDX_API_KEY} - LOG_LEVEL=info restart: always # Internal AI Agent Service agent-service: build: . depends_on: - cdx-proxy environment: - OPENAI_BASE_URL=http://cdx-proxy:8080/v1 - OPENAI_API_KEY=${CDX_API_KEY}

4. Sidecar Environment Variables & Resource Reference

Standard configuration environment variables injected into the sidecar container runtime.

Environment VariableDefaultRequiredDescription
TARGET_UPSTREAMhttps://api.openai.comRequiredUpstream LLM provider URL (OpenAI, Anthropic, Gemini, internal vLLM)
CDX_LICENSE_KEYNone (Open Beta)OptionalEnterprise license key (Defaults to Open Beta free tier if omitted)
CDX_PORT8080OptionalSidecar inbound listening TCP port
ENABLE_PROMETHEUStrueOptionalExposes Prometheus time-series metrics on /metrics
LOG_LEVELinfoOptionalLogging verbosity (debug, info, warn, error)

5. Cloudflare Pages & Supabase Production Cloud Architecture

The CDX portal and global APIs run on a high-availability serverless topology powered by Cloudflare Anycast Edge CDN/WAF and Supabase Multi-Tenant PostgreSQL (RLS).

• Edge Layer (Global CDN & WAF): Cloudflare Pages with 330+ Anycast PoPs guarantees < 15ms TTFB globally, with Always-On L3/L4/L7 DDoS mitigation and Bot Fight Mode.
• IAM & Ledger Layer: Supabase Auth (FIDO2/Passkey, GitHub/Google OAuth, SAML SSO) with Strict PostgreSQL Row Level Security (RLS) guarantees 100% cryptographic tenant isolation.
• Zero Data Retention (ZDR): Customer prompts are processed in-memory and never written to disk or database. Only mathematical token metrics and HMAC billing receipts are securely recorded.
# [Step 1] Push Supabase PostgreSQL Schema with Strict RLS supabase db push --file supabase_schema.sql # [Step 2] 1-Click Global Edge Deployment to Cloudflare Pages (cdxengine.com) npx wrangler pages deploy . --project-name=cdx-engine-portal --branch=main

Prometheus, Grafana & OpenTelemetry Observability

Official enterprise telemetry reference for piping real-time token savings, P99 in-memory latency, active sessions, and swarm health metrics to Grafana, Datadog, and OpenTelemetry via GET /metrics.

1. Prometheus Scrape Configuration (prometheus.yml)

Configure Prometheus to scrape the CDX sidecar proxy at port 8080 on a 10s interval.

scrape_configs: - job_name: 'cantor-cdx-200-fleet' scrape_interval: 10s metrics_path: '/metrics' static_configs: - targets: ['localhost:8080'] labels: environment: 'production' region: 'us-east-1'

2. Core Exported Prometheus Metrics Reference

The 5 core time-series telemetry metrics exposed by the CDX in-memory engine.

Metric NameTypeKey LabelsDescription & Unit
cdx_tokens_pruned_total Counter swarm, model, modality Cumulative count of lossless pruned tokens (Tokens)
cdx_reduction_ratio Gauge swarm, agent_id Real-time token reduction ratio (0.00 ~ 1.00, verified 0.20 ~ 0.3155)
cdx_proxy_latency_seconds Histogram le, endpoint CDX in-memory slimming latency (P50, P90, P99 in seconds)
cdx_active_sessions Gauge swarm Number of currently active concurrent agent sessions
cdx_requests_total Counter status, model Total ingress HTTP proxy requests and status codes (200, 4xx, 5xx)

3. Ready-to-use Grafana PromQL Queries

Drop-in PromQL snippets for Grafana dashboard panels and chart visualization.

# 1) Token pruning velocity (Tokens / sec) sum(rate(cdx_tokens_pruned_total[5m])) by (swarm) # 2) P99 In-Memory Proxy Latency (Converted to ms) histogram_quantile(0.99, sum(rate(cdx_proxy_latency_seconds_bucket[5m])) by (le)) * 1000 # 3) Average token reduction percentage by Swarm cluster (%) avg(cdx_reduction_ratio) by (swarm) * 100 # 4) Fleet-wide HTTP Request Success Rate (%) sum(rate(cdx_requests_total{status="200"}[5m])) / sum(rate(cdx_requests_total[5m])) * 100

4. OpenTelemetry (OTel) Collector & Datadog Pipeline

Forward CDX metrics to Datadog or Dynatrace using the OpenTelemetry Collector.

receivers: prometheus: config: scrape_configs: - job_name: 'cdx-slimmer' scrape_interval: 10s static_configs: - targets: ['cdx-slimmer.production:8080'] exporters: datadog: api: key: "${DATADOG_API_KEY}" site: "datadoghq.com" service: pipelines: metrics: receivers: [prometheus] exporters: [datadog]

5. Prometheus Alerting Rules (alerting_rules.yml)

Instant alerting rules for P99 latency spikes (>1ms) or proxy downtime via Slack & PagerDuty.

groups: - name: cdx_alert_rules rules: - alert: CDXProxyHighLatency expr: histogram_quantile(0.99, sum(rate(cdx_proxy_latency_seconds_bucket[5m])) by (le)) > 0.001 for: 1m labels: severity: warning annotations: summary: "CDX Sidecar P99 latency exceeded 1.0ms (Baseline: 0.08ms)" - alert: CDXSidecarDown expr: up{job="cantor-cdx-200-fleet"} == 0 for: 30s labels: severity: critical annotations: summary: "CDX Sidecar container is unreachable on port 8080"

6. Grafana 1-Click Dashboard Template (Import JSON)

Ready-to-use Grafana JSON dashboard snippet for immediate import via Dashboards > Import:

{ "title": "Cantor CDX Token Slimmer Observability", "panels": [ { "title": "Token Pruning Rate (Tokens/sec)", "type": "timeseries", "targets": [{"expr": "sum(rate(cdx_tokens_pruned_total[5m]))"}] }, { "title": "P99 In-Memory Latency (ms)", "type": "gauge", "targets": [{"expr": "histogram_quantile(0.99, sum(rate(cdx_proxy_latency_seconds_bucket[5m])) by (le)) * 1000"}] }, { "title": "Real-time Reduction Ratio (%)", "type": "stat", "targets": [{"expr": "avg(cdx_reduction_ratio) * 100"}] } ] }

Version Matrix & Verified Model Support

Semantic Versioning v2.0.0 standards and verified global LLM compatibility matrix.

1. Official LLM Compatibility Matrix

Model / FrameworkVerified VersionSupported Optimizations
OpenAI GPT-5 / GPT-4.5 / o3 / o1 / GPT-4ov1.0.0-beta+100% Compatible (Thinking & Tool Calling)
Anthropic Claude 3.7 Sonnet / 3.5v1.0.0-beta+100% Compatible (Hybrid Reasoning & Tools)
Google Gemini 2.0 / 3.0 Flash & Prov1.0.0-beta+100% Compatible (2M+ Multimodal Calling)
DeepSeek-V3 / DeepSeek-R1v1.0.0-beta+100% Compatible (Private GPU VRAM Pruning)

2. Product Release Roadmap

VersionTarget ScheduleCore Architecture & ScopeTarget Reduction
v1.0 (Current)2026.09 LiveTool JSON Schema AST Pruning, Prefix Normalization, KV-Cache Hit Preservation30% ~ 35%
v2.0 (Upcoming)2026.Q4 ScheduledTool Output & Multi-Turn Large Context Pruning
Intelligent lossless filtering for large SQL result sets, web scrape data, and long agent dialogues
60% ~ 77%+
v2.12027.Q1 ScheduledUltra-low latency in-memory engine optimization (<10µs latency) and cluster auto-scaling60%+ Maintained