Axle Informatics

Axle Informatics

Axle Informatics specializes in providing innovative technology solutions and research support to advance biomedical research, collaborating with prominent organizations like the NIH to accelerate discovery and enhance organizational success.

Pharmaceuticals
51-250
Founded 2002

Description

  • Design and implement monitoring and observability frameworks across distributed systems using Splunk, Grafana, and OpenTelemetry.
  • Establish and manage SLIs, SLOs, and error budgets to improve reliability.
  • Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments.
  • Automate workload onboarding and offboarding with standardized governance.
  • Track system ownership, dependencies, and lifecycle states for operational transparency.
  • Build proactive detection and intelligent alerting using AIOps.
  • Design and operate scalable, resilient, and secure infrastructure platforms.
  • Implement automated compliance tracking and enforcement aligned with standards such as NIST, FISMA, and FedRAMP.
  • Embed ITIL processes into SRE workflows.
  • Build and maintain automated deployment pipelines and golden-path platform templates for consistent workload delivery.
  • Automate provisioning, patching, configuration management, and environment lifecycle tasks.
  • Enable AI/ML infrastructure for data pipelines, storage, inference, GPU, and high-performance workloads.
  • Support cloud migration, modernization, and platform standardization initiatives.
  • Promote DevOps, SRE, and platform engineering best practices across developer communities.

Requirements

  • 6+ years of experience in DevOps or SRE roles.
  • 4+ years of hands-on Linux experience with Ubuntu, CentOS, or Red Hat, including containers and dependency management.
  • 4+ years of experience automating Infrastructure as Code deployments on AWS, GCP, or Azure.
  • 4+ years of experience with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, or GitHub Actions.
  • Strong scripting skills in Python, Bash, PowerShell, or similar languages.
  • Proficiency using vibe coding and coding assistants to develop DevOps and SRE scripts, tools, and applications.
  • Experience monitoring on-prem and cloud-hosted workloads with tools such as Prometheus, Grafana, ELK, or cloud-native equivalents.
  • Ability to troubleshoot or deploy SQL and NoSQL databases, object storage, web servers, and open-source stacks such as Node.js, R, Python, .NET Core, or Java.
  • Willingness to learn and adapt to new technologies and changing project needs.
  • Cloud certifications preferred.
  • Certifications in Grafana, Splunk, Docker, or Kubernetes preferred but optional.

Benefits

  • 100% medical, dental, and vision coverage for employees.
  • Paid time off and paid holidays.
  • 401(k) match up to 5%.
  • Educational benefits for career growth.
  • Employee referral bonus.
  • Flexible spending accounts for healthcare, parking reimbursement, dependent care, and transportation.
  • Market-competitive salary with a base range of $140,000 to $155,000 USD.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Developer (python/java) / SRE

WatchGuard Technologies 1K-5K Internet Software & Services

WatchGuard is hiring a remote Site Reliability Developer in Spain to support the reliability, security, and operational excellence of its production cloud environments alongside application teams.

Apache Spark AWS Azure CloudFormation Docker Elasticsearch Flink GitHub Go Java Jenkins JIRA Kubernetes Microservices New Relic Python Serverless Terraform
50 minutes ago

Site Reliability Engineer (6266)

Dan.com - a GoDaddy brand Internet Software & Services

itD is hiring a Site Reliability Engineer for a 100% remote U.S.-based contract focused on automating and improving the reliability of large-scale cloud infrastructure and production environments.

Ansible AWS CI/CD GitLab CI Kubernetes Linux RSpec Ruby
50 minutes ago

Sr. Site Reliability Engineer, tvScientific

Pinterest 5K-10K Internet Software & Services

Pinterest is hiring a Senior Site Reliability Engineer to operate and improve tvScientific’s cloud-native CTV advertising platform on AWS, Kubernetes, and GitOps workflows.

Argo CD AWS Bash CI/CD GCP GitHub Actions GitOps Helm Kubernetes Linux Python Secrets Management Terraform
23 hours, 35 minutes ago

Site Reliability Engineering (SRE) Leader

PatSnap 251-1K Internet Software & Services

PatSnap is hiring a Site Reliability Engineering (SRE) Leader to lead its UK SRE team and drive the reliability, scalability, security, and performance of its global SaaS platform.

AWS Docker Kubernetes
1 day ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers