Coderio

Coderio

Coderio specializes in providing on-demand, enterprise-level software development talent, enabling businesses to quickly assemble and scale their tech teams with experienced developers aligned to their time zones.

Internet Software & Services
51-250
Founded 2018

Description

  • Define and execute the observability strategy aligned with SRE and DevOps best practices.
  • Design end-to-end monitoring architectures for infrastructure, applications, and services.
  • Configure business-impact-driven alert thresholds and reduce alert noise.
  • Develop and maintain real-time operational dashboards in tools such as Grafana and Kibana.
  • Implement monitoring automation, including agent deployment and automated incident response.
  • Administer, patch, and maintain monitoring platforms while optimizing costs.
  • Create and maintain service maps, monitoring runbooks, and troubleshooting procedures.
  • Support incident resolution through root cause analysis and event correlation.

Requirements

  • 3+ years of experience in Monitoring, IT Operations, SRE, or Systems Administration.
  • Advanced experience with observability platforms such as Prometheus, Grafana, ELK Stack, New Relic, or Datadog.
  • Hands-on experience monitoring cloud environments such as AWS, Azure, or GCP.
  • Experience with containerized workloads using Docker and Kubernetes.
  • Knowledge of log aggregation tools such as Fluentd, Logstash, or Loki.
  • Knowledge of distributed tracing tools such as Jaeger, Zipkin, or OpenTelemetry.
  • Proficiency in Python or Bash for automation and custom checker creation.
  • Strong Linux administration skills and root-cause analysis capabilities.
  • Bachelor’s degree in Computer Science, Systems Engineering, or equivalent practical experience.
  • Preferred: official cloud certifications in AWS, Azure, or GCP.
  • Preferred: tooling certifications in Datadog, Dynatrace, Elastic, or Prometheus.
  • Preferred: SRE or DevOps certifications/foundational knowledge.
  • Preferred: solid grasp of networking concepts such as TCP/IP, DNS, and load balancing.

Benefits

  • 100% remote, long-term role with autonomy and impact.
  • Strategic, high-visibility position in a modern engineering culture.
  • Collaborative international team with strong technical leadership.
  • Clear path to growth and leadership within Coderio.
  • Inclusive, challenging environment with fair compensation.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
2 days ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
2 days ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
2 days, 23 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
3 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers