Skip to content

Gauss Labs

Senior Site Reliability Engineer (KR)

Location
Seoul, KR
Arrangement
Hybrid
Employment type
Full-time
Level
Senior
Posted
24 July 2026 (about 2 months ago)

Checked about a month agoApplications go to the employer, never to RoleSprint

About this role

Gauss Labs is an industrial AI company on a mission to revolutionize manufacturing with AI, starting with the semiconductor sector. Panoptes is an AI-based virtual metrology solution deployed in high-volume manufacturing fabs, helping customers improve yield, reduce costs, and accelerate production. Our software runs in our customers' own managed environments, and we're seeking a Site Reliability Engineer to own the reliability of the infrastructure and platform that Panoptes runs on. You will keep the platform available, performant, and scalable; own monitoring, alerting, incident first-response, and the on-call rotation; and build the automation and observability that let engineering teams operate their services safely.

Responsibilities • Platform reliability and operations: Own platform-layer reliability across both environments. In our internal cloud environment: full ownership — cluster health, resource management (CPU/memory/OOM), scheduling, autoscaling, Kubernetes/EKS lifecycle. In the customer environment: operate directly at the application-namespace level and for the customer-controlled cluster/node layer, diagnose and clearly communicate what's needed, and operate the platform within their setup, decisions, and constraints.

• Monitoring and Alerting: Build and maintain robust monitoring and alerting for the infrastructure and platform layer to proactively identify and resolve issues before they impact the platform.

• Incident Response: Own incident first-response for the platform layer and participate in the on-call rotation to minimize downtime and restore service quickly.

• Automation: Develop automation tools and scripts to streamline operations, reduce manual effort, and enable engineering teams to operate their own services safely.

• Capacity Planning: Forecast resource needs, optimize resource utilization, and ensure the platform infrastructure can handle increasing workloads.

• Deployment infrastructure: Build and maintain CI/CD pipelines and deployment infrastructure for the platform.

• Continuous Improvement: Drive a culture of continuous improvement by identifying opportunities to enhance platform reliability, performance, and efficiency.

Basic Qualifications • Bachelor's degree in computer science, engineering, or a related discipline

• 5+ years of industry experience as a Site Reliability Engineer or in platform/infrastructure engineering

• Hands-on experience operating Kubernetes in production (EKS preferred): cluster lifecycle, scheduling, autoscaling, resource management

• Experience with cloud platforms (AWS preferred) and containerization technologies (Docker, Kubernetes)

• Experience with observability and alerting tools (Prometheus, Grafana, ElasticSearch, Jaeger)

• Experience with scripting languages (Python, Bash)

• Working knowledge of GitHub, GitHub Actions, and CI/CD concepts

• Strong problem-solving and troubleshooting skills

• Working proficiency in English for internal documentation and technical coordination

Preferred Qualifications • Knowledge of AI/ML infrastructure and workloads.

• Knowledge of database technologies (MongoDB, PostgreSQL)

• Experience operating software in customer-managed (on-prem or customer-cloud) environments

• Exposure to manufacturing, semiconductor, or enterprise B2B customer environments

[Interview process] Application reivew - Phone interview - Virtual onsite interview - VP interview/Core Value interview - CEO interview

Work location

  • Seoul, KR

Related jobs

Ready to make a decision?

This role is either worth your time or it isn’t.

Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.

Nothing is submitted automatically. You choose what happens next.

About this listing

Published on Lever under the board identifier Gausslabs, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.

RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.

Published 24 July 2026, last checked about a month ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.

Browse all current openings

No credit card requiredStart free