Skip to content

Megaport

Senior Site Reliability Engineer

Location
Sao Paulo, BR
Arrangement
Hybrid
Employment type
Other
Level
Senior
Posted
13 August 2026 (about a month ago)

Checked 5 days agoApplications go to the employer, never to RoleSprint

About this role

The Role At Latitude.sh, the Reliability team is responsible for the health and resilience of the infrastructure that powers our global bare metal cloud. As a Senior Site Reliability Engineer (SRE), you’ll focus on building reliable, observable, and self-healing systems at scale.

SREs at Latitude.sh work at the intersection of software engineering and infrastructure. You’ll design and implement tools that automate operations, improve incident response, and enhance system observability—ensuring our platform is always ready for the workloads of our customers.

This might be a good opportunity if you’re passionate about reliability, automation, and creating cloud-like experiences for bare metal infrastructure.

What You'll Be Doing • Continuously improve Latitude.sh’s platform reliability and performance

• Design, build, and maintain tools to automate operational tasks and incident response

• Implement and improve observability solutions, including monitoring, alerting, and tracing

• Collaborate with engineering and platform teams to design scalable and resilient systems

• Participate in on-call rotations and lead post-incident reviews with a focus on learning

• Develop and document processes and runbooks that ensure operational excellence

• Contribute to SLOs/SLIs definition and reliability metrics adoption across teams

What We're Looking For • Strong verbal and written English communication skills

• Advanced knowledge of Linux/Unix systems in production environments

• Experience with Kubernetes and container orchestration

• Proficiency with infrastructure automation tools (e.g., Terraform, Ansible)

• Experience with observability stacks (e.g., Prometheus, Grafana, Loki, ELK)

• Familiarity with scripting and programming languages such as Bash, Python, Go, or Ruby

• Working knowledge of Git and CI/CD pipelines

• Solid understanding of incident management and root cause analysis processes

• Knowledge of cloud-native reliability and security best practices

What We Offer • Contractor (PJ)

• Paid Time Off

• Competitive Compensation

• Wellhub (former Gympass)

• Annual Bonus based on company and team performance

• Flexible work hours

• Opportunities for professional growth and development

Work location

  • Sao Paulo, BR

Related jobs

Ready to make a decision?

This role is either worth your time or it isn’t.

Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.

Nothing is submitted automatically. You choose what happens next.

About this listing

Advertised by Megaport and published on Lever, the applicant tracking system they use.

RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.

Published 13 August 2026, last checked 5 days ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.

Browse all current openings

No credit card requiredStart free