Jobgether
Senior AI Systems Quality Engineer
- Location
- US
- Arrangement
- Remote
- Employment type
- Full-time
- Level
- Senior
- Posted
- 30 September 2026 (yesterday)
Checked todayApplications go to the employer, never to RoleSprint
About this role
Accountabilities: • Build and deploy production-grade automated validation frameworks, test harnesses, and evaluation pipelines across the full AI development lifecycle.
• Design and evolve an AI testing platform integrated with Databricks and MLflow to enable repeatable testing, traceability, lineage, and auditability.
• Create large-scale, scenario-based test suites covering hundreds or thousands of cases, including edge cases, long-tail scenarios, and system failure modes.
• Validate agentic orchestration behaviors such as tool usage, memory, decision logic, and non-deterministic outputs before production deployment.
• Embed quality-by-design principles by defining system contracts, guardrails, safe-degradation patterns, and validation requirements at key system boundaries.
• Define measurable quality signals for LLM systems, including grounding, hallucination rates, relevance, latency, cost, accuracy, and explainability.
• Integrate automated quality gates into CI/CD pipelines and ensure validation runs continuously following model, prompt, or code changes.
• Build reusable testing libraries, frameworks, and components that enable engineering teams to adopt consistent AI quality practices.
• Establish measurable release-readiness criteria and support go/no-go decisions based on defined quality thresholds.
• Partner with AI, platform, security, and delivery teams to translate business and mission requirements into clear quality criteria, trade-offs, and confidence levels.
• Evaluate system behavior, reliability, security, privacy, and operational risk in regulated and mission-critical environments.
Requirements:
• 7+ years of software engineering experience, primarily focused on backend or platform systems.
• Proven experience designing and implementing automated AI testing and validation solutions in production environments.
• Demonstrated ability to build custom testing, validation, or evaluation frameworks for complex and distributed systems.
• Strong proficiency in Python and/or TypeScript within modern AI engineering environments.
• Hands-on experience with AI-powered systems, including LLM-based or agentic workflows and non-deterministic behavior.
• Experience designing AI testing at scale, including regression frameworks, long-tail evaluations, and broad test coverage.
• Deep understanding of CI/CD practices and experience embedding automated tests and quality gates into deployment pipelines.
• Solid knowledge of AWS cloud-native architectures.
• Strong track record of engineering for quality, reliability, governance, safety, and operational resilience as core system principles.
• Working knowledge of security, privacy, and operational risk within regulated or mission-critical environments, including failure modes and recovery strategies.
• Experience with AI testing methodologies such as non-deterministic output evaluation, drift detection, bias and fairness testing, and robust regression strategies.
• Ability to establish measurable trust thresholds and operationalize metrics such as query accuracy, hallucination limits, explainability, and PHI-safe behavior as release criteria.
• Experience collaborating with domain experts to define correctness and real-world validation scenarios that reflect genuine production use cases.
• Experience with Databricks and Medallion architecture is preferred but not required.
• Familiarity with MLflow for model evaluation, lineage, and auditability is a plus.
• Exposure to observability tools such as Datadog, Prometheus, or Grafana is desirable.
• Familiarity with LLM evaluation techniques, guardrails, and policy enforcement frameworks is beneficial.
• Experience evaluating AI performance, latency, and cost regressions is a plus.
• Ability to clearly communicate system behavior and quality trade-offs to both technical and business audiences.
• Formal AI/ML training or certifications, such as ISTQB AI Testing, AWS ML Specialty, or Google ML Engineer, are welcome.
• Experience designing prompts, agent behaviors, and orchestration logic as versioned, testable artifacts is advantageous.
• Familiarity with using AI systems to generate and expand diverse, adversarial, and large-scale test scenarios is a plus.
Benefits:
• Compensation based on experience, skills, and location, including base salary, performance bonus eligibility, and equity grants.
• Unlimited paid time off.
• Work-from-anywhere flexibility.
• Comprehensive health coverage with multiple plan options.
• Equity for every employee.
• Growth-focused environment with opportunities for professional development.
• One-time home office setup allowance.
• Monthly cell phone allowance.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Work location
- US
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Lever under the board identifier Jobgether, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 30 September 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.