Staff Software Engineer

Engineering Tel Aviv · On-site / hybrid Full-time Staff level

We're hiring a staff software engineer to build and own the systems that run Anyray's optimizer on real enterprise traffic. This is production-grade work on the hot path of every LLM request.

01The role

Anyray is a self-hosted, pre-model middleware that optimizes the cost and performance of every LLM request. That optimization lives across every layer of a request: input, context, tool output, files, dev commands, and caching.

It's a hands-on engineering role. You own real production systems end to end, and you're on the hook for their latency, reliability, and correctness under real customer load.

02Why this role

LLM optimization is a young field with no settled playbook, so the optimizations you invent here are genuinely new, and they ship into production rather than a paper. And they matter: real enterprises run enormous volumes of LLM traffic and feel every wasted token, so the work you do lands on live customer requests and solves a problem they actually have. It's about as close to the frontier, and as close to the product, as engineering gets.

03What you'll do

  • Invent and ship optimizations. Design new strategies across every layer of a request: prompt and context compression, tool pruning, file and dev-command trimming, semantic caching, cache-prefix stabilization. Then take them all the way to the hot path. Each one has to cut tokens on real traffic without degrading output.
  • Prove every win. Build the evals and the token/quality instrumentation that separate a genuine saving from a plausible-sounding one. An optimization only ships when the numbers back it.
  • Live in the caching trade-offs. Provider prompt caches only pay off on byte-stable prefixes, and a naive rewrite can cost more than it saves. A lot of the job is knowing exactly when an optimization actually wins: deterministic transforms, per-session decision pins, semantic dedup.
  • Build and own the gateway that runs them. The real-time proxy that carries production LLM traffic, with streaming, retries and backpressure, executing every optimization on the live request path.
  • Keep it fast and deploy anywhere. Defend a tight latency budget on the hot path, and package the whole thing so it installs cleanly into any customer's cloud, whether that's AWS, GCP, Azure, containers or Kubernetes, and upgrades without drama.
  • Own systems end to end. From first commit to code serving live customer traffic, plus the spend-visibility and governance product built on top, with a direct line to the founders.

04What we're looking for

  • You've built and run backend or infrastructure systems in production, owned them end to end, and care about latency, correctness, and observability.
  • You're comfortable in the request hot path. Concurrency, streaming, tail latency and failure modes are things you reason about by default.
  • You've shipped to cloud environments and can build something that installs cleanly in someone else's.
  • You're strong in TypeScript/Node, which is our stack, or fluent enough in a systems language that picking it up is quick.
  • You already build AI-native. Coding agents and LLM tooling are part of how you ship.
  • You understand, or want to go deep on, LLMs, tokens, and how model APIs actually behave.

05Nice to have

  • You've built proxies, gateways, or high-throughput streaming and networking systems.
  • Experience with LLM inference, model APIs, or ML infrastructure.
  • You've packaged software for customer self-hosting, with Helm, operators or air-gapped installs.
  • You've been early at a company before and liked it.

06How to apply

If this sounds like you, email hi@anyray.ai. Tell us about a system you built and owned end to end, or a nasty production problem you tracked down and fixed.