Jobgether
Software Engineering Director, Agentic Evaluations
- Compensation
- $260K–$320K / yr
- From job posting
- Location
- US
- Arrangement
- Remote
- Employment type
- Full-time
- Level
- Director
- Posted
- 26 September 2026 (4 days ago)
Checked 4 days agoApplications go to the employer, never to RoleSprint
About this role
Accountabilities: • Lead the development and continuous improvement of evaluations that run against live software agents, identifying practical approaches for producing useful, credible, and repeatable results.
• Own the full technology stack supporting the evaluation platform, including administrative tooling, APIs, evaluation workflows, orchestration, and production systems.
• Design reusable evaluation primitives and architecture that can be applied across multiple software categories and reduce the effort required to onboard new agent types.
• Turn recurring integration, onboarding, and maintenance processes into reusable AI skills, agents, or automated workflows that increase evaluation velocity.
• Monitor emerging frameworks, methodologies, and industry practices for agent evaluation, identifying opportunities to improve measurement quality and distribution of results.
• Partner closely with data science teams to develop proprietary benchmarks and richer approaches to evaluating agent performance.
• Lead, mentor, and develop a core engineering team, providing technical direction and subject-matter expertise in AI agent evaluations.
• Promote the effective use of evaluations across agent-focused products and initiatives throughout the wider organization.
Requirements
• 10+ years of professional software development experience, primarily in backend or full-stack environments.
• 2+ years of direct engineering management experience, including team leadership, mentoring, and technical direction.
• Expert-level backend development skills in languages such as Python, Java/Kotlin, TypeScript/JavaScript, or Go, with strong experience in frameworks such as FastAPI or Node.js.
• Hands-on experience building evaluations for customer-facing AI agents, including the use of agent trajectory traces and evaluation rubrics to measure task completion, accuracy, correctness, and/or policy adherence.
• Direct experience using frontier models from providers such as OpenAI, Anthropic, or Google in LLM-as-a-judge applications.
• Regular experience using coding-agent tools such as Claude Code, Codex, OpenCode, or Pi as part of a modern software development workflow.
• Bachelor’s degree in computer science, engineering, or a related field.
• Experience integrating agent tool use directly or through MCP servers is a plus.
• Familiarity with benchmark frameworks such as STATE-Bench, tau2-bench, or similar evaluation approaches is desirable.
• Experience using browser automation and agent tooling such as Playwright, browser-use, or Chrome DevTools MCP is advantageous.
• Experience with durable execution and workflow platforms such as Temporal, DBOS, Cloudflare Workflows, or Vercel Workflows is helpful.
• Familiarity with agent sandboxing technologies such as AWS E2B, Daytona, Cloudflare Containers, or Vercel Containers is also valuable.
• Strong communication, technical leadership, mentoring, and cross-functional collaboration skills are essential.
Benefits
• Total earnings of approximately $260,000–$320,000, combining base salary and bonus.
• Equity participation.
• Performance-based bonus opportunities.
• Fully remote position within the United States.
• Flexible working environment designed to support distributed teams.
• Unlimited paid time off.
• Generous parental leave.
• Comprehensive benefits designed to support employee well-being and flexibility.
• Inclusive and diverse workplace with employee-led community initiatives and professional growth opportunities.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Work location
- US
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Lever under the board identifier Jobgether, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 26 September 2026, last checked 4 days ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.