Beauty For All Industries (BFA)

Beauty For All Industries (BFA)

Beauty For All Industries (BFA), also known as Beauty for All, is a leading beauty innovation platform that integrates technology and community to make beauty accessible to everyone. The company is dedicated to inspiring self-expression and operates at the intersection of technology and beauty, utilizing extensive member data to create personalized shopping experiences. BFA serves over 30 million subscribers and generates over $1 billion in annual revenue. Headquartered in San Mateo, California, with additional offices in New York, Miami, Santa Monica, and Argentina, BFA is a digitally native company that functions as a beauty subscription powerhouse and brand incubator. Its portfolio includes notable brands such as IPSY, the largest beauty subscription service in the world, BoxyCharm, Madeby Collective, and Refreshments. BFA focuses on inclusivity and aims to create a welcoming community for all, while also attracting entrepreneurial thinkers and creative problem solvers to drive its mission forward.

information technology & services
51-200
Founded 2020
$333M raised

Description

  • Build and maintain Datadog observability, including dashboards, monitors, APM, log pipelines, and low-noise alerts.
  • Define and track SLIs, SLOs, and error budgets with service owners.
  • Participate in on-call rotations and support incident triage, prioritization, escalation, and resolution.
  • Lead incident response documentation, status updates, handoffs, and ownership transfers.
  • Manage alerting and escalation through Opsgenie and incident communication through Slack.
  • Conduct blameless post-incident reviews, identify root causes, and track preventive actions.
  • Automate operational toil through scripts, tooling, self-healing, and automated remediation.
  • Use AI tools such as Claude and Cursor to accelerate debugging, runbook creation, RCA drafting, and automation.
  • Improve reliability across AWS and third-party services including Netlify, CommerceTools, Auth0, and Contentful.
  • Contribute to CI/CD reliability, deployment safety, infrastructure-as-code, runbooks, and triage workflows.

Requirements

  • Experience with observability and monitoring tools, ideally Datadog; Grafana, Prometheus, New Relic, or CloudWatch experience also applies.
  • Experience participating in on-call rotations and incident response, including triage, prioritization, escalation, and post-incident reviews.
  • Working knowledge of SRE principles, including SLIs, SLOs, error budgets, and toil reduction.
  • Working knowledge of AWS and distributed, microservice, and API-gateway architectures.
  • Scripting and automation skills in Python, Bash, or similar languages, with the ability to read and reason about code.
  • Familiarity with CI/CD pipelines and infrastructure-as-code tools such as Terraform.
  • Strong written and verbal communication during incidents and in RCA documentation.
  • Ability to collaborate across distributed, multi-time-zone teams.
  • Comfort using AI tools for debugging, documentation, RCAs, and reliability workflows.
  • Preferred: experience with high-traffic consumer or e-commerce platforms, Opsgenie or PagerDuty, CommerceTools, Auth0, or Amplitude.
  • Must work remotely from Mexico or Colombia while covering the PST time zone.
  • Resume/CV must be submitted in English.

Benefits

  • Competitive salary paid in USD.
  • Paid time off.
  • Work-from-home flexibility.
  • Remote-first work environment with virtual activities and company-wide offsites.
  • Professional development and learning sessions.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

[Job 31894] AI ORCHESTRATOR (APP SRE)

CI&T 5K-10K Internet Software & Services

A CI&T busca um AI Orchestrator para governar a entrega ponta a ponta de um projeto de monitoramento SRE em jornadas de aplicativos de saúde, coordenando equipes e agentes de IA desde o upstream até a produção.

Agile
2 days, 1 hour ago

Senior Site Reliability Engineer (SRE)

UJET 251-1K Professional Services

UJET is seeking a Senior Site Reliability Engineer to build and scale its SRE function, improving reliability, reducing operational toil, and establishing production engineering best practices for its cloud-based contact center platform.

AWS Azure Go Java Kubernetes Microservices Python Terraform
3 days ago

[31893] AI ENGINEER SRE (APP)

CI&T 5K-10K Internet Software & Services

Engenheiro(a) SRE/Software Sênior na CI&T, atuando na rearquitetura e melhoria de performance de aplicativos de saúde, desde a definição da arquitetura até a entrega de soluções testadas para produção.

Angular AWS CI/CD Confluence Java JIRA Kanban Microservices Node.js Scrum Spring Boot
4 days, 1 hour ago

Graduate Site Reliability Engineer

Axon 1K-5K Professional Services

Axon is seeking a junior Site Reliability Engineer in Australia to improve the reliability, scalability, security, and operational performance of mission-critical cloud-native services.

AWS Azure C# CI/CD Go Java Kubernetes Linux Python Terraform
4 days, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers