Coforge

Coforge

Coforge is a global digital services provider specializing in transforming businesses through technology and industry expertise, powering growth with innovative solutions and platforms.

IT Services
10K-50K
Founded 1992

Description

  • Design, build, and operate scalable and highly available cloud platforms.
  • Ensure the reliability, performance, and stability of distributed production systems.
  • Implement and maintain infrastructure as code using Terraform or similar tools.
  • Manage Kubernetes-based and containerized environments.
  • Define and operate SLOs, SLIs, error budgets, dashboards, runbooks, and alerting standards.
  • Implement observability, monitoring, and incident response practices.
  • Participate in on-call rotations and respond to production incidents.
  • Collaborate with engineering teams to improve automation, scalability, and platform resilience.
  • Conduct postmortem reviews and drive continuous reliability improvements.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Systems, Software Engineering, or a related technical field, or equivalent practical experience.
  • 6+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps Engineering, Backend Engineering, or Production Engineering.
  • Strong software engineering skills in at least one language such as Python, Go, Java, TypeScript, or C#.
  • Strong understanding of distributed systems, microservices, APIs, asynchronous processing, queues, databases, caching, retries, idempotency, and failure modes.
  • Experience with cloud infrastructure on AWS, Azure, or GCP.
  • Experience with Kubernetes, containers, Terraform or similar IaC tooling, CI/CD pipelines, and Linux-based systems.
  • Experience with observability tools such as Datadog, Prometheus, Grafana, OpenTelemetry, CloudWatch, New Relic, Splunk, or Sentry.
  • Experience defining and operating SLOs, SLIs, error budgets, alerting standards, dashboards, runbooks, and incident response practices.
  • Strong communication skills and experience working across cross-functional teams.
  • Cloud, Kubernetes, Infrastructure, Reliability Engineering, Security, or DevOps certifications are preferred.
  • Experience in logistics, transportation, final-mile delivery, field-service software, routing, dispatch, or fleet operations is preferred.
  • Experience working with operational SaaS or marketplace platforms is preferred.
  • Experience driving automation, platform reliability, and operational excellence initiatives is preferred.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AMS:US:SL:Service Reliability Engineer:Lead

Thoughtworks 10K-50K Professional Services

Thoughtworks is hiring a Service Reliability Engineer to improve the reliability, resilience, and performance of client infrastructure and production systems.

Ansible CircleCI CloudFormation ELK Stack GitLab GitOps Go Grafana Jaeger Java Jenkins Kubernetes Nomad Prometheus Python Ruby Shell Scripting Terraform Zipkin
5 hours, 43 minutes ago

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

Bitly 51-250 Internet Software & Services

Shield AI is seeking a Sr. Staff Platform / Data Reliability Engineer to make its Databricks platform reliable, secure, scalable, and operationally mature for enterprise and regulated use.

CI/CD Databricks Git
2 days, 5 hours ago

Senior Manager Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Manager of Site Reliability Engineering in India to lead its SRE Center of Excellence and build a global, platform-focused operations capability that improves reliability, developer productivity, and scale.

AWS CI/CD Datadog GitHub Actions Go Grafana Kafka Kubernetes Microservices PostgreSQL Prometheus Python Terraform
3 days, 5 hours ago

Vice President, Global Production Operations & Reliability

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Vice President, Global Production Operations & Reliability to lead the company’s global production operations for its cloud-native SaaS platform and drive reliability, scalability, security, and operational excellence.

AWS CI/CD Kubernetes
4 days, 4 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers