We're hiring a staff research engineer to push Anyray's inference optimizations to the frontier — genuinely new work that lands on real enterprise traffic.
01The role
Anyray is a self-hosted, pre-model middleware that optimizes the cost and performance of every LLM request. That optimization lives across every layer of a request — input, context, tool output, files, dev commands, and caching.
The role is simple to state and hard to do well: make that optimization save more tokens without ever costing output quality or breaking a provider's cache. It's a research-heavy engineering role — you form theories from real traffic and ship the ones that prove out on live customer requests.
02Why this role
LLM optimization is a young field with no settled playbook, so the optimizations you invent here are genuinely new — and they ship into production, not a paper. And they matter: real enterprises run enormous volumes of LLM traffic and feel every wasted token, so the work you do lands on live customer requests and solves a problem they actually have. It's about as close to the frontier, and as close to the product, as engineering gets.
03What you'll do
- Find where the tokens go and cut them. You start from real enterprise traffic, form a theory about what's wasteful, and turn it into an optimization that ships to production and provably saves tokens.
- Prove the token savings. An idea that sounds good is worth nothing until the numbers back it. You'll build the evals and the token and quality measurement that tell a real win from a plausible one.
- Live in the caching trade-offs. Provider caches only help when the prefix is byte-for-byte stable, and a clever rewrite can easily burn more tokens than it saves. Knowing which is which is a lot of the job.
- Use AI to do the work of a team. We expect you to build coding agents, tooling and pipelines for yourself. The good ones end up as tools everyone here runs on.
- Set the bar. You take an optimization all the way, from a notebook to code serving live customer traffic, and the way you work rubs off on how the rest of us do.
04What we're looking for
- Strong engineering fundamentals, and the judgment to tell when something is actually correct, fast and safe rather than just green in CI.
- A real research instinct. You're fine with ambiguity, you form theories and drop them when the data says no, and you'd rather be right than clever.
- You already build AI-native. Coding agents and LLM pipelines are part of your day, and you've used them to do work that used to take a team.
- You know LLMs at a low level: tokenization, context windows, how caching actually behaves, streaming, tool calls, and how the various providers bill.
- Staff-level range. You've carried big, vague pieces of work on your own and shaped how a team builds.
05Nice to have
- You've worked on inference infra, an LLM gateway or proxy, or model-serving efficiency before.
- You've built retrieval, embeddings, compression, or caching systems.
- You're solid in TypeScript/Node, which is what the gateway and optimizer are written in, and comfortable in Python for research.
- You've been early at a company before and liked it.
06How to apply
If this sounds like you, email hi@anyray.ai. Tell us about an optimization you'd try first, or a time the data made you drop an idea you liked.