Faire

Faire

Faire is an online wholesale marketplace connecting independent retailers with unique merchandise from around the world. With flexible payment terms, free returns, and personalized recommendations, Faire empowers small businesses to compete with larger...

Textiles, Apparel & Luxury Goods
1K-5K
Founded 2017
$1500M raised

Description

  • Design and operate ML infrastructure, including workspaces, clusters, jobs, and workflows.
  • Productionize ML workloads using Spark, Delta Lake, MLflow, and Databricks Workflows.
  • Teach data scientists how to move models from notebooks into production on the ML platform.
  • Implement Unity Catalog for data governance, lineage, access control, and secure multi-tenant usage.
  • Build CI/CD pipelines for machine learning using Terraform and Git-based workflows such as GitHub Actions.
  • Optimize performance, reliability, and cost across training and inference workloads.
  • Configure IAM and RBAC for sensitive datasets.
  • Establish observability for data quality, model performance, and platform health.
  • Build and maintain technical documentation for the ML platform.

Requirements

  • 8+ years of experience building production ML or data platforms.
  • A degree, preferably graduate level, in Computer Science, Engineering, Statistics, or a related technical field.
  • Strong hands-on expertise with Databricks, Spark, Delta Lake, and MLflow.
  • Proficiency in Python, SQL, and distributed systems concepts.
  • Experience with cloud platforms and infrastructure-as-code.
  • Solid understanding of MLOps best practices, including CI/CD, monitoring, reproducibility, and security.
  • Experience supporting multiple ML teams in a shared platform environment.
  • Experience with Kotlin, PyTorch, Kafka, Snowflake, Fivetran, Iceberg, Datadog, Airflow, Cockroach DB, or MySQL is preferred.
  • Experience with AWS, S3, SageMaker, Kubernetes, Docker, GitHub Actions, or Terraform is preferred.
  • Familiarity with generative AI tools such as Claude Sonnet 4.5 and ChatGPT 5.2 is listed in the tech stack.

Benefits

  • Salary range of $224,000 to $308,000 per year in San Francisco.
  • Eligibility for equity.
  • Hybrid work schedule with 3 days per week in the office.
  • Flexibility to work remotely up to 4 weeks per year in hybrid roles.
  • Reasonable accommodation support during the recruitment process.
  • Equal employment opportunity commitment.
  • Access to benefits, though specific plan details are not listed.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

ML Infrastructure Engineer

x.ai 51-250 Internet Software & Services

SpaceXAI is hiring an ML Infrastructure Engineer to build and optimize the machine learning platform that powers recommendations on X.

Ansible C++ Linux Puppet Python PyTorch Rust
16 hours, 16 minutes ago

Senior Machine Learning Engineer, Safety

Reddit 1K-5K Internet Software & Services

Reddit is hiring a remote-friendly Machine Learning Engineer to build and improve safety systems that support enforcement of Reddit rules using large language models.

Deep Learning LLM Machine Learning NLP Python PyTorch TensorFlow
16 hours, 46 minutes ago

Senior, Machine Learning Engineer - 3D Perception

Torc 251-1K Road & Rail

Torc is hiring a Senior Machine Learning Engineer – 3D Perception to develop and deploy production perception models for autonomous trucks and improve Bird's Eye View understanding across its autonomy stack.

C++ Computer Vision Deep Learning Machine Learning Python PyTorch
17 hours, 1 minute ago

Applied AI Engineer

Unframe Inc. 51-200 Technology, Information and Internet

Unframe is hiring an Applied AI Engineer to work with enterprise customers and internal teams on deploying AI-driven solutions that connect business problems to production systems.

Machine Learning Python React TypeScript
1 day, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers