Task Development Engineer

4 weeks, 1 day ago
Contract
Mid Level
Software Development
METR

METR

METR, or Model Evaluation and Threat Research, is a nonprofit research institute located in Berkeley, California. Founded in August 2022, METR focuses on evaluating advanced AI models to identify capabilities that may pose significant risks to society. The organization conducts pre-deployment empirical evaluations of AI systems, assessing dangerous capabilities such as autonomous replication and cybersecurity threats. METR's mission is to develop scientific methods for assessing the risks associated with AI systems' autonomous capabilities. It provides services to leading AI companies, including OpenAI, Anthropic, and Google DeepMind, helping them understand AI capabilities and risks before deploying new models. The organization also contributes to the development of standardized evaluation methodologies and publishes research to enhance public understanding of AI risks. With a dedicated team, METR aims to promote safe AI development and informed decision-making.

nonprofit organization management
51-200
Founded 2022
$71M raised

Description

  • Develop novel, difficult, well-scoped tasks that remain challenging as AI model time horizons increase.
  • Perform quality assurance on existing tasks to verify solvability, clarity, and appropriate information constraints.
  • Baseline tasks within areas of expertise when useful.
  • Score task completions produced by AI systems or human baseliners.
  • Identify inefficient or low-quality workflows and improve task-development infrastructure and processes.
  • Contribute to evaluations whose results inform policymakers, frontier AI labs, national security stakeholders, and other decision-makers.

Requirements

  • Several years of software engineering experience with complex projects and codebases.
  • Experience building difficult AI evaluations, ideally agent-based evaluations.
  • Familiarity with evaluations such as RE-Bench, HCAST, SWE-bench Verified, Cybench, or GPQA.
  • High attention to detail, including the ability to identify ambiguity, misspecifications, and small errors.
  • Experience with the Inspect evaluation framework preferred.
  • Prior experience with METR’s Hawk infrastructure preferred.
  • Familiarity with the methodology behind METR’s Time Horizons work preferred.
  • Ability to overlap with Pacific Coast Time for at least 1 hour daily, ideally 4 hours.

Benefits

  • Remote, worldwide contract or freelance work.
  • Flexible schedule of 20–40 hours per week.
  • Compensation of $150–$300 per hour, with the top of the range reserved for exceptional candidates.
  • Contributors working more than 80 hours may receive acknowledgment in the final research output, if desired.
  • Collaborative, mission-driven research culture focused on truth-seeking, integrity, and high-quality science.
  • Opportunity to contribute to research informing major decisions about frontier AI risks and progress.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI/ML Architect

NewRocket 251-1K Internet Software & Services

NewRocket is hiring an AI/ML Architect in India to support a large global client by designing, deploying, and maintaining scalable machine learning and AI applications.

Apache Spark AWS Azure BERT GCP Generative AI GPT Hugging Face LLM Machine Learning Matplotlib MongoDB MySQL NLP NumPy Pandas PostgreSQL Power BI Python PyTorch Scikit-learn Seaborn Snowflake SQL Tableau TensorFlow Vertex AI
5 hours, 46 minutes ago

Dermatologist (AI Evaluation)

Gramian Consultancy Group Professional Services

Gramian Consultancy is seeking India-based Dermatologists for a four-week remote contractor assignment evaluating dermatology images and AI-generated clinical content to improve healthcare AI reliability and safety.

5 hours, 46 minutes ago

AI Voice Evaluation Specialist

Innodata 1K-5K IT Services

Innodata is hiring Voice Specialists to conduct consistent, real-time conversations with AI models and evaluate which model provides the stronger conversational experience.

Generative AI
1 day, 5 hours ago

AI Conversation Specialist - (Khaleeji/Gulf) Saudi Arabic

RWS Group 5K-10K Internet Software & Services

RWS is seeking freelance AI Business Conversation Analysts fluent in Saudi Arabic (Khaleeji Gulf) to conduct structured bilingual conversations that support the development and improvement of AI models in business and analytical domains.

1 day, 5 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers