Skip to content

METR

Task Development Engineer

Compensation
$260.9K–$385.5K / yr
Location
Berkeley, US
Arrangement
On-site
Employment type
Other
Posted
1 August 2026 (about 2 months ago)

Checked todayApplications go to the employer, never to RoleSprint

About this role

[Due to capacity constraints, we may be slow to get back to you about this role.]

About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment. We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.

About the role • As part of informing the world about risk from frontier AI systems, METR often runs and p ublishes evaluations of frontier models.

• Time Horizons is a central tool the world uses to understand AI progress. Our methodology has been included in system cards , called an "obsession" by the NYT , has wide reach online, and is used by governments to inform national policy. It is essential to our broader risk assessment work to have good capability evaluations.

• Task Development Engineers contribute to METR’s expanding ambition of our evaluations with high quality tasks supporting the Time Horizons methodology. We expect our results to be seen by policymakers, frontier labs, national security stakeholders, and other key decisionmakers influencing society’s response to AI progress.

Note: We recently changed this role to be a full time, in-person role by default (though we are happy to discuss contractor/remote setups if you prefer that).

What this role looks like • (Primarily, and most importantly) Developing difficult, novel tasks for models. You will build well-scoped tasks that remain challenging as model time horizons grow, potentially to hundreds of hours.

• Quality assurance for existing tasks. Once a task has been developed, you will verify that it's actually solvable as specified, and that the model is given (only) the information it needs.

• Baselining and scoring tasks. Where helpful, you may be asked to baseline tasks within your domain of expertise, and/or score task completions from AIs or human baseliners.

• Improving task development infrastructure. We're always improving our processes. Strong candidates will notice when existing workflows are inefficient or produce low-quality output, and take responsibility for improving them.

Skills we're looking for • Software engineering: You have several years of experience working on complex projects and codebases.

• Evaluations: You have experience building hard (ideally agent-based) AI evaluations (e.g. RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA), ideally using the Inspect framework.

• High attention to detail: You read closely, spot misspecifications and ambiguity, and pay attention to fiddly minutiae.

• (Nice to have) Familiarity with METR infrastructure: Prior experience with Hawk , and familiarity with the methodology behind our Time Horizons work, is a plus.

Our Culture

METR is a mission-driven organization. We believe our work can meaningfully shape humanity's future for the better, and we want to be the best people in the world doing this work. We have a tight-knit, collaborative research culture rooted in truth-seeking and integrity. We're fiercely committed to producing high-quality, trustworthy science. We're honest and transparent about our results, especially when they may go against the grain. We've earned trust as reliable partners who handle confidential information with care. We maintain a low-ego, drama-free environment focused on what matters.

We encourage you to apply even if your background may not seem like the perfect fit! We would rather review a larger pool of applications than risk missing out on a promising candidate for the position.

We are committed to diversity and equal opportunity in all aspects of our hiring process. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We welcome and encourage all qualified candidates to apply for our open positions.

Work location

  • Berkeley, US

Related jobs

Ready to make a decision?

This role is either worth your time or it isn’t.

Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.

Nothing is submitted automatically. You choose what happens next.

About this listing

Published on Lever under the board identifier Metr, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.

RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.

Published 1 August 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.

Browse all current openings

No credit card requiredStart free