Symmetrio

Symmetrio

Symmetrio is a top Staffing and Recruiting company in the Philadelphia region, specializing in recruiting qualified full-time candidates, providing staff augmentation services, and offering advisory services to help clients meet their corporate objecti...

Professional Services

Description

  • Serve as the primary technical owner for production reliability across U.S. customer environments.
  • Investigate and resolve complex issues across web applications, APIs, backend services, data pipelines, cloud infrastructure, and customer integrations.
  • Lead production incident response efforts and coordinate cross-functional teams to restore service and reduce customer impact.
  • Perform root cause analysis and drive corrective actions that improve long-term system stability and resilience.
  • Partner with software engineering and platform teams to identify recurring reliability risks and implement sustainable solutions.
  • Design, configure, and validate secure customer connectivity solutions, including Site-to-Site VPNs, Transit Gateway integrations, routing configurations, and secure network paths.
  • Support customer onboarding by troubleshooting connectivity issues and ensuring consistent implementation processes.
  • Improve platform observability through monitoring, logging, alerting, tracing, and operational dashboards.
  • Contribute to CI/CD, infrastructure automation, and deployment processes that improve release safety and operational consistency.
  • Develop operational tooling for incident response, troubleshooting, onboarding, and system monitoring.
  • Collaborate with engineering leadership to improve cloud architecture, scalability, security, and operational readiness.
  • Partner with customer-facing teams to communicate technical issues, remediation plans, and reliability improvements clearly.
  • Support compliance, security, and risk management initiatives in regulated healthcare environments.

Requirements

  • 6+ years of hands-on experience supporting and managing AWS-based production environments.
  • 4+ years of experience supporting web applications and backend services; Python/Django experience is strongly preferred.
  • Experience with AWS networking technologies including VPCs, Site-to-Site VPNs, Transit Gateways, routing, NAT gateways, and security groups.
  • Strong experience with Terraform and infrastructure-as-code deployment practices.
  • Experience with containerized environments including ECS, Fargate, Kubernetes, or similar technologies.
  • Experience building and supporting CI/CD pipelines and release automation processes.
  • Familiarity with monitoring and observability platforms such as Datadog, CloudWatch, Sentry, Grafana, or similar tools.
  • Experience leading production incidents, outage management, and root cause analysis initiatives.
  • Exposure to Windows Server environments, Active Directory, Kerberos, and enterprise infrastructure concepts is preferred.
  • Healthcare technology, healthcare SaaS, clinical software, or other regulated industry experience is highly preferred.
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field is preferred.

Benefits

  • Health Care Plan (Medical, Dental & Vision).
  • Retirement Plan (401k, IRA).
  • Paid Time Off (Vacation, Sick & Public Holidays).

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer - AWS and Azure

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to build, maintain, and improve reliable, scalable, and secure cloud infrastructure across Windows and Linux environments.

Ansible AWS Azure Bash CI/CD Docker GitHub GitHub Actions Grafana Kubernetes PowerShell Prometheus Python TeamCity Terraform
19 hours, 29 minutes ago

[Job-31445] Sênior Software Engineer | SRE & Software Architecture, Brazil

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa especialista em observabilidade para atuar como facilitadora técnica junto aos times de desenvolvimento, apoiando a confiabilidade, a performance e a evolução das aplicações.

Agile Angular Azure CI/CD Datadog Docker Git GitFlow GitHub Actions Grafana Java Kanban Kubernetes OpenShift OpenTelemetry Prometheus Scrum WAF
1 day, 18 hours ago

Senior Site Reliability Engineer

Sports Academy Education Services

Texas Sports Academy is seeking a part-time Senior Site Reliability Engineer consultant to audit, improve, and scale the infrastructure supporting its AI-first K–12 school.

AWS CI/CD Datadog Grafana Prometheus
1 day, 19 hours ago

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
3 days, 19 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers