Multi Media

Multi Media

Multi Media LLC is a leading live streaming technology company that provides innovative live streaming video solutions, engaging community forums, data insights, encryption, custom payment technology, global compliance, and smart in-app messages. Their...

Internet Software & Services
51-250
Founded 2011

Description

  • Analyze system performance using APM and distributed telemetry to identify instability sources.
  • Improve scalability, reliability, and performance through software enhancements and patching.
  • Develop tools and automation to streamline the DevOps pipeline.
  • Design and manage infrastructure across data center metal environments and public cloud platforms.
  • Conduct predictive failure analysis and disaster planning.
  • Administer and configure databases and key-value stores with a focus on uptime and performance.
  • Analyze complex systems to reduce operational surprises and minimize downtime.
  • Participate in incident response and write postmortem reports.
  • Collaborate with other engineering teams on reliability and infrastructure initiatives.
  • Help shape reliability strategy across systems and teams at different levels of ownership.

Requirements

  • STEM degree and/or relevant experience as a Site Reliability Engineer, DevOps Engineer, or Software Engineer.
  • Proficiency in Python or Golang, or another compiled/high-level language such as C, C#, C++, Java, or Rust.
  • Experience running web applications at scale.
  • Experience with web application concepts and frameworks such as ORM, MVC, Django, Flask, or Laravel.
  • Strong Linux administration skills, including Bash and knowledge of Linux internals such as filesystems and system calls.
  • Strong networking knowledge, including routing, switching, TCP stack, and cloud networking concepts such as VPCs and Security Groups.
  • Experience in database administration and configuration.
  • Experience with DevOps tools such as Terraform, Ansible, Docker, Kubernetes, ArgoCD, or Helm.
  • Willingness to participate in on-call rotation and respond to monitoring and alerting for core website functions.
  • Experience in production environments at an intermediate, senior, or staff level (preferred).

Benefits

  • Fair and competitive base salary of $169,000 - $215,000 USD.
  • Fully remote optional.
  • Health, vision, dental, and life insurance for you and dependents, with premiums covered by the company.
  • Long- and short-term disability insurance.
  • Unlimited PTO and 12 paid holidays.
  • Annual year-end company closure.
  • Optional 401(k) with 5% matching.
  • Paid lunches in-office or a $125/week remote stipend via Sharebite.
  • Employee Assistance and Employee Recognition Programs.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Lodgify 251-1K Internet Software & Services

Lodgify, a Barcelona-based vacation-rental technology company, is hiring a Senior Site Reliability Engineer to improve the reliability, scalability, observability, and operational ownership of its cloud platform and critical product services.

Datadog Grafana Kubernetes Microservices Prometheus Python
10 hours, 2 minutes ago

Senior Site Reliability Engineer

PointClickCare 1K-5K Health Care Providers & Services

PointClickCare is seeking a Senior Site Reliability Engineer to provide technical leadership and improve the reliability, automation, observability, and operational efficiency of cloud-based healthcare applications.

Agile Ansible AWS Azure C C++ Chef Docker Go Java Kubernetes Linux Perl Puppet Python Ruby TCP/IP Terraform Windows Server
10 hours, 17 minutes ago

AWS - Incident Handler

Caseware 251-1K Internet Software & Services

Caseware is hiring a fully remote Incident Commander in Colombia to lead incident response for its 24/7 SaaS operations, coordinating resolution, communication, root-cause analysis, and post-incident improvements.

AWS JIRA New Relic PagerDuty
2 days, 9 hours ago

Principal Site Reliability Engineer, Platform

Blue River Technology 251-1K Industrial Conglomerates

Blue River Technology, a John Deere company developing AI and robotics for agriculture and construction, is seeking a Principal Site Reliability Engineer to shape and scale the Kubernetes-based platform that enables reliable delivery of autonomous systems and products.

Argo CD AWS CI/CD GitHub Actions Go JavaScript Kubernetes Python Rust Terraform
2 days, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers