Aspenview Technology Partners

Aspenview Technology Partners

Aspenview Technology Partners empowers IT organizations by providing agile, expert-staffed nearshore IT teams that deliver scalable capacity and advanced capabilities in software development, data engineering, business intelligence, artificial intellig...

Internet Software & Services

Description

  • Design and implement reliability strategies for distributed systems across AWS and GCP.
  • Define and track SLIs, SLOs, and other reliability metrics.
  • Build and improve observability solutions using monitoring, logging, tracing, and alerting tools.
  • Lead incident response, root cause analysis, and postmortem processes.
  • Collaborate with engineering teams to improve performance, resiliency, scalability, and operational readiness.
  • Automate operational processes and reduce toil through engineering solutions.
  • Guide teams on reliability-focused architecture, capacity planning, and non-functional requirements.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 7+ years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong experience supporting production systems in AWS and/or GCP.
  • Experience operating and troubleshooting Kubernetes platforms such as EKS and/or GKE.
  • Experience supporting large-scale cloud migration or modernization programs.
  • Expertise in incident management and production operations for high-availability systems.
  • Experience working in Agile, DevOps, or DevSecOps environments.
  • Strong knowledge of observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Strong scripting and automation skills using Python, Bash, or comparable languages.
  • Solid understanding of networking, distributed systems, cloud security, and performance optimization.
  • Experience implementing chaos engineering or resilience testing practices.
  • Knowledge of service mesh technologies such as Istio.
  • AWS and/or GCP certifications are a plus.
  • Must be currently authorized to work in the United States on a permanent basis without visa sponsorship now or in the future.

Benefits

  • Competitive base salary.
  • Comprehensive benefits and wellness support.
  • Flexible work model with hybrid, remote, or in-office options.
  • Real growth opportunities and leadership visibility.
  • Inclusive, respectful culture.
  • A people-first company that listens, invests in employees, and celebrates wins together.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

AssureSoft 51-250 Internet Software & Services

AssureSoft is hiring a remote Site Reliability Engineer to support production cloud infrastructure and platform reliability for long-term client projects.

Argo CD AWS Bash DNS Docker GCP GitHub Actions Go Grafana HTTP Kafka Kubernetes Linux Load Balancing Prometheus Python RabbitMQ Snowflake TCP/IP TLS TypeScript Unix
36 minutes ago

Senior Site Reliability Engineer

Megaport 251-1K Diversified Telecommunication Services

Megaport is hiring a Senior Platform Engineer to support secure, reliable, and maintainable global production systems within its SRE-focused platform team.

AWS Bash Cassandra CI/CD ClickHouse Git GitHub Go Kubernetes Linux PostgreSQL Python Terraform
1 hour, 6 minutes ago

Site Reliability Engineer - Azure, Observability and Scripting

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to support the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes.

Azure Bash Grafana Kubernetes OpenTelemetry Oracle PowerShell Prometheus Python Terraform
1 day, 1 hour ago

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Senior Service Reliability Engineer at Thoughtworks, focused on improving infrastructure reliability, observability, and incident response for customer-facing production systems.

Azure Bash Datadog ELK Stack GitOps Go Grafana Kubernetes Microservices Network Security New Relic Nomad Python REST API Serverless Terraform
1 day, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers