PhonePe

PhonePe

PhonePe is India's top digital wallet and payment app, providing UPI payments, investments, insurance, recharges, and more to over 44 crore users.

Capital Markets
5K-10K
Founded 2015
$2500M raised

Description

  • Configure, maintain, and manage Ubuntu virtual machines and Azure infrastructure components.
  • Design and manage Azure services for log storage, alerting, monitoring, and operational management.
  • Configure and maintain complex networking components including Azure Firewall, route tables, virtual network gateways, ExpressRoute, and IPsec connectivity.
  • Automate BAU tasks and infrastructure provisioning using Terraform, Saltstack, Ansible, and scripting.
  • Set up, manage, and replicate high-availability databases and data stores such as MySQL and Aerospike.
  • Implement and maintain monitoring, logging, and visualization systems using Prometheus, Victoria Metrics, Riemann, Loki, and Grafana.
  • Manage firewalls, coordinate with SOC and Infosec teams, and remediate security vulnerabilities.
  • Perform capacity planning and support critical services including Nginx, HAProxy, Docker, and RabbitMQ.
  • Participate in on-call rotation, lead incident response, conduct RCA, and support post-mortems and DR failovers.
  • Document procedures, create runbooks, and collaborate with technical and non-technical stakeholders.

Requirements

  • Deep hands-on experience with Microsoft Azure, including Virtual Machines, Storage Accounts, CosmosDB, and Azure Data Explorer.
  • Expert-level Azure networking knowledge, including Azure Firewall, Route Tables, Virtual Network Gateways, ExpressRoute, and Azure Private DNS.
  • Proficiency in BGP routing and troubleshooting connectivity with on-prem data centers.
  • Strong Linux administration experience, specifically with Ubuntu/Linux.
  • Deep expertise in at least one high-level language such as Python, Go, or Java.
  • Mastery of Bash shell scripting for automation and operational tasks.
  • Strong experience with monitoring tools such as Prometheus, Victoria Metrics, and Riemann.
  • Proficiency with centralized logging using Loki and dashboarding with Grafana.
  • Strong experience with Terraform and configuration management tools such as Saltstack or Ansible.
  • Hands-on experience managing high-availability MySQL and Aerospike, with familiarity in Elasticsearch and InfluxDB.
  • Experience with Nginx, HAProxy, RabbitMQ, Docker, and core networking services like DNS.
  • Experience defining and meeting SLOs and SLIs, reducing toil through automation, and optimizing cloud costs.
  • Excellent written and verbal communication skills; mentoring experience is preferred for senior roles.

Benefits

  • Medical, critical illness, accidental, and life insurance.
  • Employee Assistance Program, onsite medical center, and emergency support system.
  • Maternity, paternity, adoption, and day-care support programs.
  • Relocation benefits, transfer support policy, and travel policy.
  • Employee PF contribution, flexible PF contribution, gratuity, NPS, and leave encashment.
  • Higher education assistance, car lease, and salary advance policy.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
8 hours, 41 minutes ago

Incident Commander

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Interactive is hiring an Incident Commander to join its site reliability team and lead cross-functional incident response for its online and physical platforms.

Ansible AWS Docker Elasticsearch GCP Helm JIRA Kafka Kubernetes Linux MySQL PostgreSQL Prometheus Python Redis Terraform
8 hours, 56 minutes ago

Site Reliability Engineer

VantageScore 11-50 Banks

Site Reliability Engineer at a growing engineering team, focused on DevSecOps for maintaining the reliability, security, and compliance of cloud infrastructure, APIs, and software supply chains.

Agile AWS AWS CDK Bash CI/CD CloudFormation CodePipeline Datadog DevSecOps Docker EC2 GitHub Actions Grafana HashiCorp Vault Kong Kubernetes Microservices Python REST API Scrum Terraform
1 day, 8 hours ago

Application Site Reliability Engineer (SRE)

CXM Direct 51-250 Capital Markets

Application Site Reliability Engineer at a trading technology company, responsible for keeping .NET/C# Windows-based trading and back-office services highly reliable, observable, and resilient.

AWS Bash C# CI/CD Docker Grafana Kubernetes Microservices .NET OpenTelemetry PowerShell Prometheus Python Terraform Windows Server
1 day, 8 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers