Machine Learning Engineer - Model Evaluation & Experimentation

1 week, 4 days ago
Contract
Junior
Data Science and Analytics
Weekday

Weekday

Weekday helps companies hire engineers who are vouched by other software engineers, enabling passive income for engineers. They offer services like drafting outreach messages, shortlisting candidates, and conducting reference checks. Backed by Y Combin...

Construction & Engineering
11-50
Founded 2020

Description

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
  • Implement machine learning solutions in Python and produce reference implementations that demonstrate correct methodology and expected outcomes.
  • Execute experiments, run training pipelines, and analyze model behavior and results.
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
  • Collaborate with AI researchers and subject matter experts to improve benchmark quality, technical rigor, and evaluation consistency.
  • Document experimental methodologies and technical findings clearly.

Requirements

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
  • Practical experience conducting machine learning experiments, including setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Familiarity with reinforcement learning concepts, including reward functions, policy optimization, and training behavior, is preferred.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and ability to solve complex, open-ended technical problems independently.
  • Ability to commit approximately 35 hours per week on a consistent basis.
  • Experience developing or evaluating large language models, foundation models, or generative AI systems is preferred.
  • Background in reinforcement learning, deep learning, distributed training, or model optimization is preferred.
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation is preferred.
  • Experience contributing to research publications, open-source machine learning projects, or advanced AI systems is preferred.
  • Strong written communication skills for documenting experimental methodologies and technical findings.

Benefits

  • Compensation of $60-$90 per hour.
  • Fully remote engagement.
  • Flexible working hours.
  • Approximately 35 hours per week.
  • Weekly payments based on approved work completed.
  • Potential for project extension depending on requirements and performance.
  • Reasonable accommodations available throughout the application and engagement process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

SWE Fellow - Human Frontier Collective (UK)

Scale AI 251-1K Diversified Consumer Services

The Human Frontier Collective Fellowship at Scale is a fully remote contract role supporting AI research by designing, evaluating, and interpreting advanced generative AI systems with interdisciplinary experts and partner labs.

C++ Generative AI Java JavaScript Machine Learning Python Rust Swift
15 hours, 47 minutes ago

Norwegian Post-Editor | Technology Company

OLIVER 1K-5K Media

A tech company is seeking a Norwegian Post-Editor to support MTPE projects by refining marketing copy and product descriptions for AI-driven content development.

16 hours, 17 minutes ago

Founding BDR — AI Consulting (Outbound + Pipeline)

PHIZENIX 11-50 information technology & services

Phizenix is hiring its first outbound sales rep to generate and qualify leads for AI agent services and keep a strong pipeline moving to founder-led closing.

CRM Salesforce
1 day, 15 hours ago

GenAI Motion Graphics Designer

OLIVER 1K-5K Media

OLIVER is hiring a remote GenAI Motion Designer in Toronto to create and localize high-volume motion and static marketing assets for a leading vacuum and floorcare brand across the U.S. and other markets.

Copywriting Design Systems E-commerce Generative AI Illustrator Photoshop
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers