Software Engineer, AI Distributed Systems


Software Engineer — Distributed AI Systems

San Francisco, CA | Five Days On-Site | $180K–$250K + Equity

AI does not change the world when it works in a demo.

It changes the world when it can operate continuously, reliably, and accurately inside the complex environments where consequential work actually happens.

We’re partnering with a well-funded AI startup building the infrastructure required to make that possible. Its technology is already operating in demanding enterprise environments, processing enormous volumes of information and turning AI-generated outputs into reliable systems people can use to make better decisions and drive meaningful change.

The company has fewer than 20 employees.

That means the systems you build, the standards you establish, and the decisions you make will materially influence the product and the company.

The Opportunity

As a Software Engineer focused on Distributed AI Systems, you’ll build and own the data, orchestration, and infrastructure layer supporting a continuously running AI platform.

You’ll work on genuinely difficult engineering problems involving large-scale data processing, distributed compute, AI-driven transformations, and production reliability.

This is not an analytics, dashboarding, or warehouse-modeling role. It is not a narrow infrastructure seat where you maintain systems designed by someone else.

You’ll own company-critical technology from architecture through production—and help determine what the business is capable of building next.

What You’ll Own

  • Build and evolve large-scale, AI-driven data and transformation pipelines

  • Turn probabilistic AI outputs into reliable, structured, usable data

  • Design resilient systems for retries, backfills, partial failures, and recovery

  • Improve platform accuracy, reliability, throughput, latency, and cost

  • Build observability, evaluation, and debugging tools for distributed workloads

  • Develop the internal platform capabilities used by other engineers

  • Make architectural decisions with direct product and business impact

  • Participate in a separately compensated on-call rotation

What We’re Looking For

  • 3+ years of experience in data platform, backend, infrastructure, or distributed-systems engineering

  • Strong production Python experience

  • End-to-end ownership of production data platforms or ETL pipelines

  • Experience operating systems at terabyte-to-petabyte scale

  • Strong understanding of distributed-system reliability and failure modes

  • Hands-on experience with AWS and infrastructure as code

  • Ability to explain not only what you built, but why you built it that way

  • Comfort solving problems without an existing blueprint

  • Willingness to work in the San Francisco office five days per week

Especially Relevant

  • AI-enabled ETL or production LLM systems

  • DAG and workflow-orchestration platforms

  • Distributed compute, queues, sharding, partitioning, or streaming

  • Spark, Ray, Flink, Kafka, Airflow, Dagster, Prefect, or Temporal

  • PostgreSQL, vector databases, or high-scale data systems

  • Infrastructure ownership through a major company-growth inflection

  • Quantitative, research, or other unusually strong problem-solving experience

Why This Role Matters

  • You’ll build the foundation. The company’s AI capabilities depend on the systems you own.

  • You’ll solve real production problems. This technology is already operating in demanding customer environments.

  • You’ll work at meaningful scale. These are continuously running distributed workloads—not occasional experiments.

  • You’ll have a seat at the table. With fewer than 20 people, one exceptional engineer can materially influence the architecture, product, and engineering culture.

  • You’ll build without a playbook. The hardest problems here do not have obvious answers or established solutions.

  • You’ll see your impact. Your work will directly determine how reliably the product performs and how much the company can accomplish.

Compensation and Logistics

  • $180K–$250K base salary

  • Competitive equity

  • Comprehensive benefits

  • Additional compensation for on-call participation

  • Visa transfers considered for qualified candidates

  • Five days per week in the San Francisco office

The world has plenty of AI products that look impressive when everything goes right.

This team is building the systems that keep AI accurate, reliable, and useful when the data is enormous, the environment is complicated, and the outcome actually matters.

If that is the kind of problem you want to own, we should talk.


*This organization participates in E-Verify.