Anyray ← back to site

About Anyray

Anyray is a self-hosted AI gateway that cuts the LLM-inference spend an organization's own employees generate — without prompt or response content ever leaving the organization's environment.

01What we build

Anyray sits on the path of every LLM request a company makes — from coding assistants like Claude Code, Cursor, and Windsurf to in-house agents, SDK jobs, and scripts. It exposes an OpenAI-compatible API in front of your existing providers (OpenAI, Anthropic, Google Vertex AI, AWS Bedrock, and Azure), serves each request as cheaply as it can with a real-time optimizer, and gives the organization full spend visibility and governance. Typical deployments cut inference spend by around 60% with no code changes for developers and no degradation to outputs.

02Why it exists

Most teams adopt LLMs faster than they can measure or control the cost, and the usual answer — a cloud-hosted router — means sending every prompt and response to a third party. Anyray was built for organizations that need both: aggressive cost control and the guarantee that prompt and response content never leaves their own environment. Content is encrypted at rest and never logged in plaintext; logs and the spend store carry metadata only.

03How we run

Anyray runs entirely in your environment via Docker, Kubernetes, or Railway. There is no multi-tenant cloud that sees your traffic. We publish our documentation, OpenAPI specification, and install guides openly so engineering teams can evaluate and deploy without a sales gate.

04Learn more