Staff Research Engineer

Engineering Tel Aviv · On-site / hybrid Full-time Staff level

We're hiring a staff research engineer to push Anyray's inference optimizations to the frontier — genuinely new work that lands on real enterprise traffic.

01The role

Anyray is a self-hosted, pre-model middleware that optimizes the cost and performance of every LLM request. That optimization lives across every layer of a request — input, context, tool output, files, dev commands, and caching.

The role is simple to state and hard to do well: make that optimization save more tokens without ever costing output quality or breaking a provider's cache. It's a research-heavy engineering role — you form theories from real traffic and ship the ones that prove out on live customer requests.

02Why this role

LLM optimization is a young field with no settled playbook, so the optimizations you invent here are genuinely new — and they ship into production, not a paper. And they matter: real enterprises run enormous volumes of LLM traffic and feel every wasted token, so the work you do lands on live customer requests and solves a problem they actually have. It's about as close to the frontier, and as close to the product, as engineering gets.

03What you'll do

  • Find where the tokens go and cut them. You start from real enterprise traffic, form a theory about what's wasteful, and turn it into an optimization that ships to production and provably saves tokens.
  • Prove the token savings. An idea that sounds good is worth nothing until the numbers back it. You'll build the evals and the token and quality measurement that tell a real win from a plausible one.
  • Live in the caching trade-offs. Provider caches only help when the prefix is byte-for-byte stable, and a clever rewrite can easily burn more tokens than it saves. Knowing which is which is a lot of the job.
  • Use AI to do the work of a team. We expect you to build coding agents, tooling and pipelines for yourself. The good ones end up as tools everyone here runs on.
  • Set the bar. You take an optimization all the way, from a notebook to code serving live customer traffic, and the way you work rubs off on how the rest of us do.

04What we're looking for

  • Strong engineering fundamentals, and the judgment to tell when something is actually correct, fast and safe rather than just green in CI.
  • A real research instinct. You're fine with ambiguity, you form theories and drop them when the data says no, and you'd rather be right than clever.
  • You already build AI-native. Coding agents and LLM pipelines are part of your day, and you've used them to do work that used to take a team.
  • You know LLMs at a low level: tokenization, context windows, how caching actually behaves, streaming, tool calls, and how the various providers bill.
  • Staff-level range. You've carried big, vague pieces of work on your own and shaped how a team builds.

05Nice to have

  • You've worked on inference infra, an LLM gateway or proxy, or model-serving efficiency before.
  • You've built retrieval, embeddings, compression, or caching systems.
  • You're solid in TypeScript/Node, which is what the gateway and optimizer are written in, and comfortable in Python for research.
  • You've been early at a company before and liked it.

06How to apply

If this sounds like you, email hi@anyray.ai. Tell us about an optimization you'd try first, or a time the data made you drop an idea you liked.