Parallel Works

Parallel Works

Parallel Works is a next-generation technical computing platform that empowers engineers and scientists to scale simulation and modeling workflows across high-performance computing systems in the cloud. With the Swift parallel scripting language at its...

Internet Software & Services
1-10
Founded 2015
$0M raised

Description

  • Build and operate production Slurm clusters, including controllers, partitions, QoS, accounting, GPU GRES, prolog/epilog, and cgroup enforcement.
  • Federate customer-owned clusters with the ACTIVATE control plane while aligning schedulers, storage, identity, accounts, and allocations.
  • Provision and maintain on-premises bare-metal systems, out-of-band management, firmware, rack networking, and vendor coordination.
  • Validate multi-node GPU environments, including drivers, CUDA, DCGM, XID triage, Fabric Manager, NVLink, InfiniBand, and NCCL.
  • Tune parallel and high-throughput storage and develop reproducible Ansible, Terraform, and image-building pipelines.
  • Perform STIG hardening, vulnerability remediation, FIPS-related security work, and security package artifact development.
  • Resolve Tier 3 escalations and participate in the on-call rotation.
  • Train junior engineers and work directly with customers from issue identification through resolution.

Requirements

  • 10+ years operating production Linux systems across multiple distribution families, including RHEL/Rocky/Alma and Debian/Ubuntu.
  • Production experience configuring, debugging, and upgrading Slurm.
  • Production experience with at least one parallel or high-throughput filesystem and InfiniBand or RoCE fabrics.
  • Multi-node NVIDIA GPU operations, including driver stack management and fault triage.
  • Experience supporting both customer-owned or on-premises clusters and public cloud environments, including bare-metal provisioning and out-of-band management.
  • Infrastructure-as-code experience with tools such as Ansible or Terraform.
  • Proficiency with Bash, Python, or similar scripting languages.
  • U.S. citizenship and eligibility for a Secret clearance; an active clearance is helpful but not required, and eligible candidates may be sponsored.
  • Preferred: Experience at a government supercomputing center, national laboratory, or university research computing center.
  • Preferred: Experience with STIG, SCAP, Tenable, eMASS, RMF, FedRAMP, Impact Level environments, Kubernetes/OpenShift GPU workloads, HPC containers, PBS Pro, Prometheus, or Grafana.

Benefits

  • Medical, vision, and dental coverage.
  • 401(k) with company match.
  • Short-term disability coverage.
  • Generous paid vacation and sick time.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Mainframe Systems Programmer - ISV

Ensono 1K-5K IT Services

Ensono is seeking a Mainframe Systems Programmer to support enterprise z/OS environments by installing, configuring, maintaining, and resolving issues with ISV and supporting systems software.

1 day, 10 hours ago

L3 Systems Engineer

Dijital Team 11-50 Internet Software & Services

Join a growing Australian managed services provider as a Level 3 Systems Engineer, owning complex escalations and delivering infrastructure and security projects across cloud and on-premises environments.

Azure Cybersecurity Datadog IoT
1 week, 6 days ago

TechOps Engineer

Ocrolus 251-1K Banks

Ocrolus is hiring a TechOps (IT) Engineer to maintain and improve the stable, reliable infrastructure supporting its AI-powered fintech platform and lending customers.

AWS Cloudflare GCP JIRA Linux
2 weeks, 1 day ago

Principal Data Center Controls Engineer

Montera is hiring a Principal Data Center Controls Engineer to define and govern portfolio-wide controls architecture, data exposure, and commissioning outcomes across hyperscale data center sites.

Microservices MQTT REST API
2 weeks, 4 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers