Chat UI, curl, or coding agent — all through the same gateway. Same routing, guardrails, and logs. The only code is the front-door Worker.
Govern workforce AI, model traffic, and agent tools at the boundary each path actually crosses.
Model traffic flows through Cloudflare AI Gateway; workforce AI flows through the SASE / SWG layer; agents reach tools through MCP server portals. Same network, three different control points.
Note — AI Gateway (apps → models) and SASE / SWG (employees → SaaS AI) are different Cloudflare products — both say "gateway" in places, but they are distinct.
One front door for every model call — chat UI, curl, or coding agent. AI Gateway handles routing, safety, DLP, caching, budgets, and logs as config, not app code.
Human lane: the Chat UI Worker page is Access-protected by one-time PIN email and injects the Run token server-side. Machine lane: curl and coding agents skip the Worker and call the Authenticated Gateway directly with the Run token. The raw gateway URL is not a bypass — the token is enforced on every call.
Two request lanes reach the same AI Gateway. Human lane: the Chat UI passes through Access identity gate 1 — the Worker page protected by one-time PIN email — and the Front Door Worker injects the Run token server-side, never exposing it to the browser. Machine lane: curl / API and coding agents skip the Worker and call the Authenticated Gateway directly with the Run token. Both lanes reach the gateway entry — ai-direct.nobledemos.com behind Access for humans, or gateway.ai.cloudflare.com with the Run token for machines — then the ai-frontdoor AI Gateway core. Model destinations: the Workers AI open-source catalog (Llama, DeepSeek, Qwen, Mistral, GPT-OSS and more) and third-party providers (OpenAI, Google, Anthropic and more) via BYOK on the same gateway with identical controls.
AI Gateway capability stack — all dashboard config, no app rewrite. 1: Authenticated Gateway — Access or service token required on every call. 2: Caching — identical-response caching, skip upstream on cache hit. 3: Rate limiting — per-identity using metadata.useremail; 2 requests per minute; returns 429 on breach. 4: Guardrails — S1 through S13 hazards plus prompt injection; FLAG mode for all S1–S13 categories — flag or block. 5: DLP Firewall tab — Credentials and Secrets blocked in request and response. 6: Spend limits — cost cap per identity returning 429 as best-effort. 7: Dynamic routing — model equals dynamic route to target model with native fallback. 8: Retries and timeouts — auto-retry on upstream errors, configurable timeout per route. Observability: Logs, Analytics, and User Insights; spend per identity via metadata.userId and cf.user_id; Logpush and OpenTelemetry. Unified Billing: pay Cloudflare in credits; provider inference passed through with no markup; Provider Keys and Secrets Store; WebSockets and realtime streaming.
Chat UI, curl, or coding agent — all through the same gateway. Same routing, guardrails, and logs. The only code is the front-door Worker.
Routing, rate limits, DLP, Guardrails, caching, budgets — all dashboard toggles. Identity decides the model roster via metadata.useremail.
Identity, model, cost, cache status, Guardrails result — all live in AI Gateway logs, with payload logging off by default.
1 click Create the gateway, then point every client at it. The only code change is the base URL:
api.openai.com → gateway.ai.cloudflare.com/v1/…/compat
toggles Routing, fallback, Guardrails, DLP, caching, budgets — dashboard config, no code, no redeploy.
instant Identity, model, cost, cache, Guardrails result — live in AI Gateway logs and User Insights.
SASE Block direct AI access so employees can only use this sanctioned path. See workforce AI ↑
Gateway shows which AI SaaS employees use. Administrators classify those applications, then apply the least disruptive control to the action and data at risk.
Open Zero Trust insights ↗A single left-to-right pipeline shows one employee request travelling through four Cloudflare One stages to a sanctioned AI application. Discover: a shadow AI app (ChatGPT) is identified from Shadow IT SaaS analytics. Sanction: its Application Status moves from Unreviewed to Approved, and status alone is not enforcement. Inspect: after Gateway TLS inspection, an AI-prompt DLP profile detects and redacts credentials in a SendPrompt operation. Constrain: the approved app stays usable for chat while only the risky upload action is blocked — the least disruptive control. The request then reaches the approved AI app.
Enable Gateway HTTP filtering, install the Cloudflare root certificate, and apply TLS decryption to supported AI applications.
Create an HTTP policy where application is Claude, operation is SendPrompt, and the DLP profile matches Credentials and Secrets.
Submit harmless synthetic data, then confirm the block page and matched DLP profile in Gateway HTTP logs.
Gateway governs traffic in motion. Connect CASB separately for sanctioned OpenAI, Anthropic, Gemini, or Copilot tenant posture and data at rest.
MCP server portals govern agent entry and tools. Code Mode is a portal capability—not a network hop; each upstream uses OAuth, Gateway routes upstream traffic, and response DLP watches data returning from tools. Use standard sensitive-data profiles because AI prompt profiles do not apply to portal traffic.
Open MCP dashboard ↗OpenCode connects through the opencode-team-access Access application to the Access-protected OpenCode Worker that hosts the opencode-team-mcp-portal. Code Mode is enabled on the portal. The portal connects to READY Cloudflare DEX MCP and Atlassian upstream servers over OAuth. Upstream traffic routes through Gateway and response-side DLP blocks credentials and secrets.
OpenCode connects through the opencode-team-access Access application to the Access-protected OpenCode Worker that hosts the portal. Code Mode is enabled in the dashboard.
Cloudflare DEX MCP is READY with 18 tools; Atlassian is READY with 31 tools.
Gateway routing is enabled. Response-side DLP blocks Credentials and Secrets from the configured upstream hosts.
Add remote HTTP MCP servers, create a portal hostname, and require Access authentication for the portal.
Select only the tools and prompts each audience needs instead of exposing every upstream capability by default.
Connect an MCP client, call an allowed tool, then confirm server, capability, status, and duration in portal logs.
A portal policy does not protect an upstream server's direct URL. Secure that endpoint with Access or its own OAuth layer.