Machine Learning Engineer - Model Evaluation & Experimentation

1 week ago
Contract
Junior
Data Science and Analytics
Weekday

Weekday

Weekday helps companies hire engineers who are vouched by other software engineers, enabling passive income for engineers. They offer services like drafting outreach messages, shortlisting candidates, and conducting reference checks. Backed by Y Combin...

Construction & Engineering
11-50
Founded 2020

Description

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
  • Implement machine learning solutions in Python and produce reference implementations that demonstrate correct methodology and expected outcomes.
  • Execute experiments, run training pipelines, and analyze model behavior and results.
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
  • Collaborate with AI researchers and subject matter experts to improve benchmark quality, technical rigor, and evaluation consistency.
  • Document experimental methodologies and technical findings clearly.

Requirements

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
  • Practical experience conducting machine learning experiments, including setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Familiarity with reinforcement learning concepts, including reward functions, policy optimization, and training behavior, is preferred.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and ability to solve complex, open-ended technical problems independently.
  • Ability to commit approximately 35 hours per week on a consistent basis.
  • Experience developing or evaluating large language models, foundation models, or generative AI systems is preferred.
  • Background in reinforcement learning, deep learning, distributed training, or model optimization is preferred.
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation is preferred.
  • Experience contributing to research publications, open-source machine learning projects, or advanced AI systems is preferred.
  • Strong written communication skills for documenting experimental methodologies and technical findings.

Benefits

  • Compensation of $60-$90 per hour.
  • Fully remote engagement.
  • Flexible working hours.
  • Approximately 35 hours per week.
  • Weekly payments based on approved work completed.
  • Potential for project extension depending on requirements and performance.
  • Reasonable accommodations available throughout the application and engagement process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI Trainer - Mechanical Engineers - CAD Expertise - (Remote Advisory - France)

Prolific 51-250 Professional Services

Prolific is seeking Mechanical Engineers to join its remote expert network for CAD-focused AI research projects that capture real-world design and validation workflows.

Linux Python
10 hours, 1 minute ago

Senior AI/ML Engineer

ALX Africa 1K-5K Diversified Consumer Services

ALX Africa is hiring a Senior AI/ML Engineer to build the core learner competency navigation engine for its AI learning platform, Project A.

LLM Machine Learning MLflow Python
10 hours, 1 minute ago

LLMOps Engineer

ALX Africa 1K-5K Diversified Consumer Services

ALX Africa is hiring a full-time LLMOps Engineer to measure and improve the performance of its AI learning platform through evaluation, reporting, and collaboration with product and AI teams.

Python Statistics
10 hours, 16 minutes ago

AI Trainer - Mechanical Engineers - CAD Expertise - (Remote Advisory - Germany)

Prolific 51-250 Professional Services

Prolific is seeking mechanical engineers to join a remote advisory talent pool supporting AI research by documenting real-world CAD design and validation workflows.

Linux Python
10 hours, 16 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers