Head of AI Safety

1 hour, 12 minutes ago
Full-time
Executive
Artificial Intelligence and Machine Learning
Moonshot

Moonshot

Moonshot develops innovative technologies and methodologies to identify and mitigate online threats, aiming to protect communities and promote safety by addressing issues such as violent extremism and disinformation through a commitment to evidence, et...

Diversified Consumer Services
51-250
Founded 2015
$7M raised

Description

  • Lead and quality-assure applied AI safety projects across violence, extremism, CSEA, abuse and grooming, mental health, crisis, and child-safety risks.
  • Design evaluation frameworks and lead hands-on red teaming and adversarial testing of AI systems.
  • Advise AI companies on model safety, product protections, policies, and intervention systems.
  • Translate subject-matter expertise into actionable guidance for model, policy, product, research, and engineering teams.
  • Identify safety failures, edge cases, and improvement opportunities while maintaining rigorous documentation.
  • Manage client and partner relationships with AI companies, governments, regulators, researchers, foundations, and civil society organizations.
  • Lead, coach, and support the wellbeing and professional development of the AI safety team.
  • Develop the portfolio through partnerships, funding, proposals, service offerings, publications, and thought leadership.
  • Oversee project staffing, budgets, forecasting, timelines, delivery quality, and operational risks.

Requirements

  • Experience in trust and safety, online harms, violence prevention, safeguarding, public health, or a related field.
  • Ability to build technical fluency in AI and engage credibly with technical counterparts.
  • Experience designing research, evaluation frameworks, or interventions for violent extremism, CSEA, self-harm, crisis, or targeted violence.
  • Experience managing projects, teams, budgets, partners, and clients.
  • Excellent written communication for government, foundation, or enterprise audiences.
  • Resilience and demonstrated comfort working with sensitive or graphic content, including CSEA, extremist, and crisis material.
  • Strong judgment, discretion, diplomacy, and ability to work through ambiguity and sensitive stakeholder environments.
  • Willingness to travel, work occasional nonstandard hours, and complete required security clearances.
  • Eligibility to work in the United States; candidates must be based in MA, CO, NY, VA, GA, PA, MD, WI, TN, TX, OR, NJ, or DC.
  • Experience with business development, grants, or procurement; model safety, LLM red teaming, child-safety evaluation, regulatory engagement, intervention design, or classifier development is preferred.

Benefits

  • Salary of $110,000–$120,000, depending on skills and experience.
  • 15 days of paid vacation, federal holidays, and additional paid leave options.
  • Private healthcare for employees and dependents, plus dental and vision insurance.
  • Life and disability insurance.
  • 24/7 counseling through an Employee Assistance Program.
  • 3% matched 401(k) contributions and Roth contributions.
  • Paid parental leave: 26 weeks maternity and 8 weeks paternity.
  • Share options granted to all permanent employees.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI Trainer - Armenian - Poland

Prolific 51-250 Professional Services

Prolific is seeking fluent Armenian speakers to evaluate AI-generated written content for naturalness, authenticity, tone, and cultural accuracy.

12 minutes ago

Freelance Software Localization and AI Language Evaluator | English to Czech

Acclaro 251-1K Professional Services

Acclaro is seeking remote freelance English-to-Czech Software Localization and AI Language Evaluators to assess and improve localized software and technical content for Czech users across technology projects.

57 minutes ago

Fractional, AI Delivery Engine Leader

Servant 11-50 Internet Software & Services

Servant is seeking a fractional AI Delivery Engine Leader to design and implement an AI-first delivery organization that moves work from strategy to production faster, more reliably, and at greater scale.

CI/CD
57 minutes ago

Freelance Agent Evaluation Engineer

Mindrift.ai: Be the “I” in AI Internet Software & Services

Mindrift is seeking experienced software developers for project-based remote work creating and evaluating realistic coding tasks and tests for AI coding agents.

Cybersecurity Docker FastAPI JavaScript Kafka Machine Learning NLP PostgreSQL Python React Redis TypeScript
57 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers