Member of Technical Staff — AI Inference Systems

Palo Alto, CA | $230K–$350K + Equity | Full-Time | Onsite

Help Build the Infrastructure AI Runs On

The next leap in AI won’t come from models alone.

It will come from making those models faster, more efficient, more scalable, and dramatically less expensive to run.

We’re partnering with a well-funded AI infrastructure startup building a next-generation inference platform from the ground up. The founding team comes from the heart of the AI infrastructure ecosystem, and they’re assembling a small group of exceptional systems engineers to rethink how LLMs are served at scale.

This is not an AI application role.

You’ll be building the engine underneath it.

What You’ll Build

You’ll have the opportunity to shape an entirely new inference runtime and tackle problems at the frontier of AI infrastructure:

  • Build a high-performance inference runtime from scratch in Rust

  • Rethink batching, scheduling, request routing, and KV cache management

  • Push the limits of latency, throughput, GPU utilization, and cost per token

  • Scale inference across multiple GPUs and nodes

  • Profile and optimize the entire serving pipeline

  • Make foundational architecture decisions alongside a small founding team

There is no mature platform to maintain and no legacy architecture dictating how things have to work.

You get to help decide how it should work.

Who We're Looking For

You’re a systems engineer who has spent real time inside LLM inference and serving infrastructure.

You understand what happens between a model request and the GPU — and you've worked on the machinery that makes that path fast.

You likely bring:

  • 2+ years of strong backend or distributed systems engineering experience

  • Hands-on experience building or optimizing LLM inference / model-serving systems

  • Deep understanding of attention, KV cache, batching, scheduling, and inference bottlenecks

  • Production experience with vLLM, SGLang, TensorRT-LLM, or comparable infrastructure

  • Strong systems programming skills in Rust, C++, Go, Python/PyTorch, or similar

Even better if you’ve worked with CUDA, Triton, NCCL, multi-GPU/multi-node serving, prefix caching, speculative decoding, or contributed to inference-focused open source.

Rust experience is a major plus — but deep inference expertise matters more. If you're exceptional in this space and can ramp quickly into Rust, we want to talk.

Why This One Is Different

AI is moving incredibly fast, but the infrastructure underneath it still has enormous room to improve.

Better inference means models can respond faster, serve more people, use hardware more efficiently, and become economically viable for entirely new categories of products.

That’s the layer you’ll be working on.

You’ll join early enough to influence the architecture, the engineering culture, and ultimately how the platform is built.

If you’re the kind of engineer who obsesses over shaving latency, squeezing more work out of GPUs, understanding why a system behaves the way it does, and building things other engineers depend on—

this is a chance to do that work at the center of where AI is going.

Compensation: $230K–$350K base + competitive equity
Location: Palo Alto, CA — 5 days/week onsite
Visa: Transfers and new sponsorship available