Glia
Senior Software Engineer, SRE / Observability Tooling
- Location
- Estonia - Remote, EE
- Arrangement
- Remote
- Employment type
- Full-time
- Level
- Senior
- Posted
- 15 July 2026 (about 2 months ago)
Checked about a month agoApplications go to the employer, never to RoleSprint
About this role
About Glia
Glia is the #1 Banking AI platform, empowering community and regional financial institutions to create efficiencies, accelerate loan growth, drive deposits, and deliver experiences that win against megabanks and fintechs.
Glia's Banking AI Operating System is a central intelligence layer on top of existing tech stacks, activating an AI workforce of specialized agents that draw from banking data, interaction history, and integrated systems of record. These banking-trained agents automate workflows across voice and digital–from front office to back office–resulting in decreased operational costs and the Universal Banker model.
Trusted by 700+ banks and credit unions for its ironclad security and reliability, Glia delivers the industry’s first contractual no-hallucination guarantee. It’s why Glia customers quickly and confidently put Banking AI to work with measurable results from day one. More information about Glia can be found at glia.com http://glia.com.
The Team
You'll be joining our dedicated Observability Team, which builds the observability platform and SRE tooling used to monitor Glia’s cloud-native core infrastructure serving the conversational AI. Our team focuses on enablement and automation: we provide the standards, tooling, and platform that engineering teams use to keep their own systems available and performing optimally.
The Mission & The Work
As a Site Reliability Engineer on this team, your focus will be on building SRE and observability tooling — the platform, automation, and standards other teams use to keep their services healthy. This is a tooling and enablement role, not a production operations role: you will not be directly operating Glia’s production services. Responsibilities will include:
- Developing standards, infrastructure and automation for dashboards, alerts, and monitors as code.
- Partnering with development teams to establish production readiness and operational readiness.
- Building the tooling and templates teams use to define, measure, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for their services.
- Developing tooling to automate observability and operational workflows, eliminating manual toil for engineering teams.
- Building and improving the incident response tooling and workflows that help teams resolve outages faster and learn from them.
Our Collaboration Model
Glia Engineering is remote-first, spanning Canada, Portugal, Poland, and Estonia. The Observability team is based in Estonia, with optional offices in Tallinn and Tartu. We thrive on flexible remote collaboration, but we still bring the whole team together in Estonia twice a year for in-person innovation and connection.
Our tech stack
- Infrastructure: AWS, Kubernetes (AWS EKS), Istio, EFK
- Persistence: Amazon Aurora Serverless for Postgres, RabbitMQ, Amazon RDS
- Cache: Amazon ElastiCache
- Monitoring & Observability: DataDog with a focus on dashboards and alerts for system health
- CI/CD: Github Actions, ArgoCD, Jenkins, Helm, with a focus on automation and pipeline optimization.
- Infrastructure as Code: Terraform
- Additionally, our Engineering teams use:
- Backend: Python, Elixir, Node.js http://node.js, Ruby, Go
- Frontend: Javascript and React.js
- Native mobile SDKs: Java and Swift
What We’re Looking for
- Expert-level proficiency with AWS and Kubernetes (EKS), particularly in areas of observability, networking, and auto-scaling.
- Experience with modern observability platforms (e.g., DataDog, Prometheus) and a deep understanding of metrics, logging, and tracing.
- Deep, practical understanding of Site Reliability Engineering (SRE) principles (SLOs, error budgets, toil reduction).
- Demonstrable experience analyzing and troubleshooting large-scale distributed systems.
- Strong software development skills in a language like Python or Go, used to build operational tools, services, or automation.
- Expertise in designing and operating robust CI/CD pipelines for a microservices architecture (e.g., using ArgoCD, Github Actions, Helm).
- A systematic, data-driven approach to problem-solving and root cause analysis.
- Proficiency in using AI tools thoughtfully, maintaining ownership of the final output while recognizing the tools' limitations.
Glia is an equal-opportunity employer. Glia does not discriminate against any employee or applicant because of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), or any other basis protected by law.
The Glia Talent Acquisition team uses @glia.com http://glia.com and @ darina.danchenko@gliatalent.comgliatalent.com http://gliatalent.com email addresses for coordinating interviews, providing updates, and sending documents.
Our hiring process involves an introduction, practical and team interviews, and a decision and offer. For more information, visit our Recruitment Privacy Notice page https://www.glia.com/eu-recruitment-privacy-notice or contact our talent team via talent@glia.com
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Ashby under the board identifier Glia, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 15 July 2026, last checked about a month ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.