Task Development Engineer

24 minutes ago
Contract
Mid Level
Software Development
METR

METR

METR, or Model Evaluation and Threat Research, is a nonprofit research institute located in Berkeley, California. Founded in August 2022, METR focuses on evaluating advanced AI models to identify capabilities that may pose significant risks to society. The organization conducts pre-deployment empirical evaluations of AI systems, assessing dangerous capabilities such as autonomous replication and cybersecurity threats. METR's mission is to develop scientific methods for assessing the risks associated with AI systems' autonomous capabilities. It provides services to leading AI companies, including OpenAI, Anthropic, and Google DeepMind, helping them understand AI capabilities and risks before deploying new models. The organization also contributes to the development of standardized evaluation methodologies and publishes research to enhance public understanding of AI risks. With a dedicated team, METR aims to promote safe AI development and informed decision-making.

nonprofit organization management
51-200
Founded 2022
$71M raised

Description

  • Develop novel, difficult, well-scoped tasks that remain challenging as AI model time horizons increase.
  • Perform quality assurance on existing tasks to verify solvability, clarity, and appropriate information constraints.
  • Baseline tasks within areas of expertise when useful.
  • Score task completions produced by AI systems or human baseliners.
  • Identify inefficient or low-quality workflows and improve task-development infrastructure and processes.
  • Contribute to evaluations whose results inform policymakers, frontier AI labs, national security stakeholders, and other decision-makers.

Requirements

  • Several years of software engineering experience with complex projects and codebases.
  • Experience building difficult AI evaluations, ideally agent-based evaluations.
  • Familiarity with evaluations such as RE-Bench, HCAST, SWE-bench Verified, Cybench, or GPQA.
  • High attention to detail, including the ability to identify ambiguity, misspecifications, and small errors.
  • Experience with the Inspect evaluation framework preferred.
  • Prior experience with METR’s Hawk infrastructure preferred.
  • Familiarity with the methodology behind METR’s Time Horizons work preferred.
  • Ability to overlap with Pacific Coast Time for at least 1 hour daily, ideally 4 hours.

Benefits

  • Remote, worldwide contract or freelance work.
  • Flexible schedule of 20–40 hours per week.
  • Compensation of $150–$300 per hour, with the top of the range reserved for exceptional candidates.
  • Contributors working more than 80 hours may receive acknowledgment in the final research output, if desired.
  • Collaborative, mission-driven research culture focused on truth-seeking, integrity, and high-quality science.
  • Opportunity to contribute to research informing major decisions about frontier AI risks and progress.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI Strategy & Transformation Architect

NeuraFlash 251-1K IT Services

The AI Strategy & Transformation Architect at NeuraFlash, part of Accenture, guides enterprise clients from AI strategy through Salesforce-based implementation by converting ambiguous business challenges into actionable roadmaps and working software.

AWS Azure Generative AI Salesforce
9 minutes ago

AI Tutor - Italian

x.ai 51-250 Internet Software & Services

SpaceXAI is seeking a multilingual AI Tutor specializing in Italian audio to curate and annotate speech data that improves Grok’s voice interactions, speech recognition, and auditory experiences across languages and accents.

macOS
24 minutes ago

Ex-MBB Strategy Consultant - AI Training (Remote)

Mindrift.ai: Be the “I” in AI Internet Software & Services

Mindrift, powered by Toloka, is seeking experienced consultants from top-tier strategy firms to design realistic consulting environments, AI tasks, and evaluation frameworks that teach advanced models business reasoning.

Generative AI Machine Learning Reinforcement Learning
1 hour, 9 minutes ago

AI Tutor - Hindi

x.ai 51-250 Internet Software & Services

SpaceXAI is hiring an AI Tutor focused on multilingual audio to help train Grok for better voice interactions, speech recognition, and audio understanding across languages and accents.

1 day ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers