Jobgether
Lead Software Engineer, Cloud Site Reliability (SRE)
- Location
- India, IN
- Arrangement
- Remote
- Employment type
- Full-time
- Level
- Lead
- Posted
- 24 September 2026 (7 days ago)
Checked 6 days agoApplications go to the employer, never to RoleSprint
About this role
Accountabilities: • Lead 24x7 NOC and site reliability operations through rotational shifts, ensuring system availability, operational stability, and adherence to service-level agreements.
• Act as Major Incident Manager for P1 and P2 incidents, coordinating triage activities, war rooms, technical teams, and stakeholder communications through resolution.
• Manage and troubleshoot Azure infrastructure, including virtual machines, networking, storage, and related cloud services.
• Administer Azure Kubernetes Service (AKS), Kubernetes, and Docker environments, including scaling, troubleshooting, performance optimization, and reliability improvements.
• Establish and enhance observability practices across logs, metrics, and traces using platforms such as Datadog and Azure Monitor.
• Drive proactive monitoring, alert optimization, anomaly detection, and AIOps initiatives to identify and address potential reliability issues before they impact users.
• Build automation and self-healing workflows using Terraform, ARM templates, Helm, Power Automate, PowerShell, Python, Bash, and other appropriate technologies.
• Collaborate with engineering teams to strengthen deployment pipelines, improve system reliability, and advance cloud-native architecture and operational practices.
• Develop operational dashboards and reports using Power BI and ServiceNow to provide visibility into incidents, reliability, and service performance.
• Lead monthly business reviews and provide clear operational reporting and insights to leadership and stakeholders.
• Mentor team members, promote knowledge sharing, standardize operational processes, and drive continuous improvements across the reliability function.
• Support broader initiatives involving multi-cloud environments, predictive monitoring, and self-healing systems where appropriate.
Requirements
• 7–12 years of professional experience in CloudOps, Site Reliability Engineering, NOC, or comparable 24x7 operations environments.
• Strong hands-on expertise with Azure infrastructure, particularly virtual machines, networking, storage, and related IaaS services.
• Extensive experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker, including troubleshooting, scaling, and performance tuning.
• Strong experience with monitoring and observability platforms such as Datadog and Azure Monitor, including integrations, alerting, dashboards, and operational analysis.
• Proven experience in incident management and major incident handling, including P1/P2 coordination, stakeholder communication, root-cause analysis, and operational reporting.
• Experience with Infrastructure as Code technologies such as Terraform, ARM templates, and Helm.
• Strong scripting capabilities using PowerShell, Python, Bash, or similar automation technologies.
• Experience working with ServiceNow, particularly Incident, Problem, and Change Management modules and associated dashboards.
• Good understanding of distributed systems, cloud-native architectures, and reliability engineering principles.
• Excellent communication, leadership, coordination, and problem-solving skills, with the ability to work effectively across technical and business teams.
• Experience working in multi-cloud environments, particularly Azure and AWS, is a plus.
• Exposure to AIOps, predictive monitoring, anomaly detection, or self-healing systems is desirable.
• Relevant certifications in Azure, Datadog, Kubernetes, or related cloud and reliability technologies are advantageous.
• Bachelor’s degree or equivalent technical education and professional experience.
Benefits
• Fully remote opportunity based in India, with a rotational shift structure supporting 24x7 operations.
• Leadership responsibility across cloud reliability, infrastructure operations, incident management, and operational excellence.
• Hands-on exposure to Azure IaaS, AKS, Kubernetes, Docker, Datadog, Azure Monitor, and Infrastructure as Code.
• Opportunity to develop advanced observability, AIOps, predictive monitoring, automation, and self-healing capabilities.
• Collaboration with engineering teams on cloud-native architecture, deployment pipelines, scalability, and reliability initiatives.
• Opportunities to mentor team members and influence operational standards and engineering practices.
• Exposure to multi-cloud technologies and large-scale distributed systems.
• An inclusive work environment focused on teamwork, openness, respect, fairness, and continuous improvement.
• Support for professional development through exposure to modern cloud and reliability technologies and relevant certification paths.
• Opportunities to participate in leadership reporting, business reviews, and cross-functional initiatives with broad organizational visibility.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Work location
- IN
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Lever under the board identifier Jobgether, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 24 September 2026, last checked 6 days ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.