Smile Digital Health

Smile Digital Health

Smile Digital Health provides cutting-edge HL7 FHIR-based clinical data repository solutions for healthcare organizations, enabling rapid digital transformation and connectivity with features like multiple FHIR versions and full text indexing.

IT Services
251-1K
Founded 2016
$20M raised

Description

  • Collaborate with Security Operations teams to define and implement cloud configuration best practices for Azure and other providers.
  • Develop and coordinate a multi-tenant approach for cloud service offerings such as databases, container platforms, authentication, certificates, and product registries.
  • Design and maintain cloud performance testing strategies, frameworks, and environments.
  • Develop and maintain cost and utilization tracking and attribution processes across cloud service providers.
  • Create documentation for cloud service offerings, including use cases, best practices, and implementation details.
  • Build and maintain technical relationships with core cloud service providers.
  • Implement and maintain secure, scalable infrastructure for delivering cloud services applications.
  • Ensure internal and external SLAs are met and continuously monitor and improve system KPIs.
  • Create tools to automate deployment, monitoring, and operations for the platform.
  • Participate in on-call rotation for application support, incident management, and troubleshooting.
  • Maintain internal tools and improve system health and reliability.
  • Support customer on-site deployments when needed.
  • Implement and manage observability tools for performance insights, including logging, metrics, and tracing.

Requirements

  • Experience with cloud service providers and best practices for implementation and configuration, preferably managing Azure for multiple teams in a SaaS environment.
  • Experience with microservices architecture, with a strong focus on Java-based services.
  • Experience applying chaos engineering practices to improve system resiliency.
  • Strong troubleshooting skills for performance issues, including time analysis, resource allocation, and optimization recommendations.
  • Familiarity with performance testing methodologies and tools for assessing behavior under load.
  • Experience with observability tools such as Prometheus and the Grafana suite.
  • Experience designing and executing performance test plans, including load, stress, soak, and spike testing, for services sustaining 500+ TPS within defined latency and error-rate thresholds.
  • Hands-on experience with performance/load testing tools such as JMeter, Gatling, or Azure Load Testing.
  • Experience tuning and validating autoscaling in Kubernetes/OpenShift HPA and Azure scale sets under variable load.
  • Experience tuning Kafka and other messaging/queueing components to sustain target transaction rates.
  • Experience with Azure-native monitoring and diagnostics tools such as Azure Monitor, Application Insights, and Log Analytics.
  • Experience with security and compliance best practices, including SOC2, HIPAA, and ISO27001.
  • Proficiency with Terraform, Ansible, or Chef.
  • Experience with troubleshooting, support escalation, on-call process optimization, and knowledge documentation.
  • Experience operating and maintaining production systems in Linux and public cloud environments.
  • Prior experience in high-performance or distributed systems, with openness to a range of experience levels.
  • Working knowledge of information security best practices.
  • Experience building or maintaining a large-scale cloud service.
  • Ability to prioritize and track multiple projects in parallel.

Benefits

  • Remote work environment.
  • Flexible time away from work, including PTO, personal days, and sick days.
  • Competitive salary and health/medical benefits.
  • RRSP/TFSA/401K employee contribution.
  • Life and disability coverage.
  • Employee Assistance Program.
  • FHIR Study Program and Skillsoft Learning.
  • Super HAPI Fun Club.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 8 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable through incident management, observability, production support, and resilience work across services.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 8 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 15 hours ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its production document workflow platform reliable, resilient, and low-downtime for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers