Avra
Member of Technical Staff | Observability & Reliability
- Location
- São Paulo, BR
- Arrangement
- Remote
- Employment type
- Full-time
- Level
- Staff
- Posted
- 23 September 2026 (8 days ago)
Checked 6 days agoApplications go to the employer, never to RoleSprint
About this role
ABOUT THE ROLE
At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area.
In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.
WHAT YOU'LL DO
- Evolve our observability stack for logs, metrics, traces, and alerting.
- Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.
- Bring telemetry into customer clusters within a model where agents only make outbound connections.
- Detect drift between the desired state and what's actually running in each environment.
- Monitor the health of our deployment and runtime agents.
- Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.
- Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.
- Reduce telemetry cost: less redundant data, more useful signal.
HOW WE MEASURE SUCCESS
- 99.9% serving availability, with incidents trending down.
- MTTR, including on-premise incidents.
- Near-zero drift between desired and actual state.
- All agents active and reporting, across every dataplane.
WHAT WE'RE LOOKING FOR
- Deep experience with OpenTelemetry and observability backends.
- Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.
- Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).
- Experience operating software in environments you don't fully control.
- Production-quality code and reviews, and a willingness to operate what you build.
NICE TO HAVE
- Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).
- GCP or GKE, AWS or EKS.
- ML multi-node/multi-cluster workloads in production.
- Financial services or regulated environments.
Work location
- São Paulo, BR
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Ashby under the board identifier Avra, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 23 September 2026, last checked 6 days ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.