Echelon
Senior/Staff AI Infrastructure Engineer
- Location
- San Francisco, CA, US
- Arrangement
- On-site
- Employment type
- Full-time
- Level
- Staff
- Posted
- 30 September 2026 (today)
Checked todayApplications go to the employer, never to RoleSprint
About this role
ABOUT ECHELON
Echelon is building the AI platform for Business Operations. Our goal is to automate the knowledge work of BizOps so one exceptional operator can deliver the leverage of an entire team, as Ramp and Rippling have done for Finance and HR.
Our platform combines passive process mining with a living ontology that maps how work happens across people, systems, documents, decisions, and outcomes. It connects structured and unstructured data trapped in fragmented, duplicative, and legacy enterprise systems, then turns that context into automated workflows and AI agents.
We are an early-stage company tackling a difficult technical problem at high speed. Engineers work directly with founders and customers, make decisions with incomplete information, ship production systems, and own the results. The pace, rate of change, and standards are high.
THE MANDATE
Build the secure execution substrate for Echelon's agents. You will scale ephemeral sandboxes, optimize filesystem and startup performance, and make long-running agent work durable under load without weakening tenant isolation or operational control.
WHAT YOU'LL OWN
- Design and operate sandbox lifecycle systems for agent code execution, browser work, document processing, and tool use.
- Improve sandbox startup time, density, scheduling, warm pools, caching, and resource utilization.
- Build fast, reliable filesystem primitives for ephemeral and persistent agent state, large artifacts, and concurrent workloads.
- Enforce tenant isolation, network policy, secrets boundaries, quotas, and least-privilege access.
- Make long-running agent jobs durable through checkpointing, retries, idempotency, cancellation, and recovery.
- Scale orchestration and control-plane services through rapid workload growth and unpredictable bursts.
- Build observability for resource pressure, execution failures, queue health, noisy neighbors, cost, and end-to-end latency.
- Run load tests, capacity plans, failure drills, and incident reviews; fix root causes rather than adding fragile workarounds.
WHAT YOU BRING
- 5+ years building production infrastructure, distributed systems, developer platforms, or execution runtimes.
- Hands-on ownership of containerized or virtualized workloads in a multi-tenant production environment.
- Strong Linux systems knowledge across processes, filesystems, networking, resource isolation, and performance debugging.
- Experience with Kubernetes or a comparable scheduler, infrastructure as code, and cloud primitives on AWS or Azure.
- Strong programming ability in Go, Rust, TypeScript, Python, or another systems-oriented language.
- Experience designing for retries, idempotency, backpressure, load shedding, observability, and safe rollouts.
- Security instincts appropriate for executing untrusted or model-generated work.
USEFUL EXPERIENCE
- Firecracker, gVisor, Kata Containers, namespaces/cgroups, seccomp, or eBPF.
- Sandbox products such as E2B, Modal, Fly Machines, or custom ephemeral compute platforms.
- FUSE, overlay filesystems, content-addressed storage, snapshotting, distributed caches, or object storage.
- Agent runtimes, code interpreters, browser automation, remote development environments, or CI execution systems.
- BYOC, private networking, customer-managed deployments, or enterprise security reviews.
- Inngest, Temporal, Kafka, NATS, or other durable workflow and event systems.
BENEFITS
- Team lunches and dinners in the office.
- On-site Gym access for all employees.
- Fully covered health, dental, and vision insurance.
- 401(k) plan.
- Access to AI Tooling of your choice.
- Unlimited PTO.
Work location
- San Francisco, CA, US
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Ashby under the board identifier Echelon, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 30 September 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.