AlphaSense

AlphaSense

AlphaSense develops an artificial intelligence-based search platform that enables investment and corporate professionals to quickly access and analyze extensive financial data and market insights from over 500 million documents, enhancing decision-maki...

Internet Software & Services
251-1K
Founded 2011
$770M raised

Description

  • Architect reliability frameworks and self-service tooling that enable teams to own the reliability of their services.
  • Drive AIOps initiatives to automate diagnostics, remediation, and proactive failure prevention.
  • Embed SRE practices across engineering through design reviews, production readiness, and operational standards.
  • Serve as Incident Commander during critical incidents and ensure blameless postmortems lead to durable improvements.
  • Deliver end-to-end monitoring, tracing, and profiling to improve system performance proactively.
  • Mentor engineers across SRE and product teams through technical guidance and knowledge sharing.
  • Influence architectural decisions and set the technical bar for reliability across the organization.
  • Lead by example in incident response and help scale a “You Build It, You Run It” culture.

Requirements

  • 8+ years of experience in Site Reliability Engineering, DevOps, or a similar role.
  • At least 3+ years of experience in a Senior+ SRE position.
  • Experience running production SaaS systems at scale.
  • Proficiency in at least one programming or scripting language such as Python or Go.
  • Hands-on experience with cloud platforms such as AWS, GCP, or Azure and Kubernetes.
  • Deep understanding of networking fundamentals, including TCP/IP, DNS, HTTP/S, and load balancing.
  • Experience with monitoring and alerting tools such as Prometheus, Grafana, Datadog, or ELK.
  • Familiarity with advanced observability tooling such as OTEL and continuous profiling.
  • Proven incident management experience, including leading high-severity incidents and postmortems.
  • Strong troubleshooting, communication, and collaboration skills.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

AssureSoft 51-250 Internet Software & Services

AssureSoft is hiring a remote Site Reliability Engineer to support production cloud infrastructure and platform reliability for long-term client projects.

Argo CD AWS Bash DNS Docker GCP GitHub Actions Go Grafana HTTP Kafka Kubernetes Linux Load Balancing Prometheus Python RabbitMQ Snowflake TCP/IP TLS TypeScript Unix
1 day, 11 hours ago

Senior Site Reliability Engineer

Megaport 251-1K Diversified Telecommunication Services

Megaport is hiring a Senior Platform Engineer to support secure, reliable, and maintainable global production systems within its SRE-focused platform team.

AWS Bash Cassandra CI/CD ClickHouse Git GitHub Go Kubernetes Linux PostgreSQL Python Terraform
1 day, 12 hours ago

Site Reliability Engineer - Azure, Observability and Scripting

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to support the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes.

Azure Bash Grafana Kubernetes OpenTelemetry Oracle PowerShell Prometheus Python Terraform
2 days, 12 hours ago

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Senior Service Reliability Engineer at Thoughtworks, focused on improving infrastructure reliability, observability, and incident response for customer-facing production systems.

Azure Bash Datadog ELK Stack GitOps Go Grafana Kubernetes Microservices Network Security New Relic Nomad Python REST API Serverless Terraform
2 days, 12 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers