Avra
Member of Technical Staff | Inference Platform
- Location
- São Paulo, BR
- Arrangement
- Remote
- Employment type
- Full-time
- Level
- Staff
- Posted
- 23 September 2026 (8 days ago)
Checked 6 days agoApplications go to the employer, never to RoleSprint
About this role
ABOUT THE ROLE
At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area.
In this role, you'll join the Platform team to own where our models execute. Customers consume our models through large batches of millions of records and through real-time APIs, and they make business decisions on every response. You'll run governed model releases reliably and efficiently — in our cloud and on customer-hosted Kubernetes — and make inference fast, predictable, and cheap enough to serve both enterprise and mid-market customers.
WHAT YOU'LL DO
- Evolve Sophos, our online and batch inference runtime, built on Kubernetes.
- Run large batch inference on ephemeral jobs, with multi-dimensional admission control (CPU, memory, GPU).
- Build and extend the controller and its Kubernetes custom resources.
- Optimize each model's inference engine and feature processing.
- Serve graphs and data efficiently.
- Own execution of training, post-training, and fine-tuning jobs, in our cloud and in customer dataplanes / on-premisse cloud.
- Drive autoscaling, GPU serving, performance, and cost optimization, with telemetry for every model we run.
HOW WE MEASURE SUCCESS
- 99.9% serving availability.
- p95/p99 latency for online inference and throughput for batch.
- Cost per prediction and per training job.
- GPU utilization: paid capacity versus capacity actually used.
- Training and batch jobs that finish on time and succeed without manual retries.
WHAT WE'RE LOOKING FOR
- Experience running model serving or large-scale batch compute on Kubernetes.
- Experience building Kubernetes controllers or operators.
- Skill at profiling and optimizing data-heavy Python pipelines.
- A clear sense of cost: you treat compute efficiency as a product feature.
- Production-quality code and reviews, and a willingness to operate what you build.
NICE TO HAVE
- Ray, Ray Serve, or KubeRay in production.
- Admission-control systems.
- GPU serving and performance optimization.
- Arrow, Parquet, Lance, or other columnar formats.
- Shipping software to customer-hosted Kubernetes.
- GCP/AWS and GKE/EKS, and financial services or regulated environments.
Work location
- São Paulo, BR
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Ashby under the board identifier Avra, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 23 September 2026, last checked 6 days ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.