Galaxy

Galaxy

Galaxy is a leading financial services firm in digital assets and blockchain, providing a full range of solutions to institutions and clients worldwide.

Capital Markets
251-1K
Founded 2018
$505M raised

Description

  • Oversee an SRE team focused on designing, deploying, and maintaining automation toolsets and the systems they support.
  • Establish and enforce infrastructure-as-code standards for consistent, repeatable, and secure deployments across the infrastructure ecosystem.
  • Lead automated configuration and state management using Ansible and Packer across Windows, Linux, and ESXi platforms.
  • Manage monitoring and observability for the automation platforms, including defining SLIs and SLOs to ensure high availability and performance.
  • Drive the automated lifecycle of physical and virtual assets, including template creation, deployment, patching, scaling, and decommissioning.
  • Develop custom scripts and internal tooling using Python, Go, PowerShell, and Bash to improve insights and operational efficiency.
  • Collaborate with the Datacenter team and other partners to support shared workflows and broader team needs.
  • Analyze system behavior and resource utilization in virtual environments to optimize automated deployment performance.
  • Provide technical guidance and career mentorship to SREs while promoting an automate-first culture and continuous improvement.

Requirements

  • 6-10 years of experience in Infrastructure, SRE, or DevOps with a focus on infrastructure automation at scale.
  • Deep proficiency with Terraform, including providers, modules, and state management.
  • Deep proficiency with Ansible, including roles, playbooks, and Tower/AWX.
  • Hands-on experience with image creation tools such as Packer, Ansible, and SCCM for standardized Windows and Linux images in hybrid environments.
  • Strong experience managing and automating virtual platforms such as VMware vSphere/vCenter and cloud platforms such as Azure and AWS.
  • High-level scripting skills in Python, Go, PowerShell, and Bash.
  • Experience with observability tools such as Splunk, ELK, Prometheus, or Grafana.
  • Strong understanding of network topology and design, with experience using platforms such as Juniper Networks or Palo Alto.
  • Strong mastery of Git branching strategies and PR workflows, plus CI/CD platforms such as Jenkins, GitLab CI, or GitHub Actions.
  • Experience troubleshooting and tuning performance for both Windows Server and Linux.
  • Previous team leadership or management experience is preferred.
  • Experience with IAM platforms such as Entra ID, Active Directory, and Okta is preferred.
  • Experience with block and object storage solutions on-prem or in cloud is preferred.
  • Storage backup and disaster recovery administration experience with Commvault or Veeam is preferred.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 36 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable through incident management, observability, production support, and resilience work across services.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 36 minutes ago

PnP Pipeline On-call Support Engineer

PHIZENIX 11-50 information technology & services

This role at an internal engineering organization focuses on monitoring and triaging PnP (Power and Performance) pipelines to keep post-silicon work stable and moving efficiently.

CI/CD Jest
16 hours, 51 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers