PointClickCare

PointClickCare

PointClickCare provides a leading cloud-based healthcare software platform that enables long-term and post-acute care providers to effectively manage the complete lifecycle of resident care while enhancing operational efficiency and improving resident ...

Health Care Providers & Services
1K-5K
Founded 2000
$232M raised

Description

  • Own the gold data layer by transforming silver tables into curated, semantically rich, documented gold datasets for AI development.
  • Reverse-engineer data semantics by working with product engineers, clinical experts, and workflow experts to understand how data is created and represented.
  • Bridge researcher needs with data design by translating AI applied research requirements into reusable gold data products and documentation.
  • Curate datasets across modalities, including structured tables, unstructured content, features, labels, and chunked/tagged data for different AI use cases.
  • Build reusable silver-to-gold data pipelines in Databricks/Spark as scheduled and observable workloads.
  • Automate data quality, filtering, synthesis, labeling, and weak supervision workflows for AI data preparation.
  • Maintain reproducible dataset snapshots, lineage, and semantic definitions for downstream AI R&D reuse.
  • Collaborate with AI researchers, data platform, product, clinical, and workflow teams throughout the R&D lifecycle.
  • Support model development, evaluation, experimentation, and operational sustaining across classical ML, generative AI, RAG, and agentic approaches.

Requirements

  • 5+ years building production data systems, including at least 2 years supporting ML or AI workloads.
  • Advanced Python, SQL, and PySpark/Databricks experience for working with large, messy data.
  • Expert-level SQL and the ability to read complex stored procedures and reverse-engineer business logic from queries.
  • Strong Databricks ecosystem experience, including Delta Lake, Unity Catalog, Spark/PySpark tuning, and MLflow.
  • Working knowledge of AI concepts such as embeddings, tokenization, feature engineering, point-in-time correctness, train/validation/test splits, and data drift.
  • Experience transforming unstructured data such as text, PDFs, transcripts, and logs into model-ready forms.
  • Familiarity with AI-friendly storage and formats such as Parquet and Hugging Face datasets, plus partitioning, sharding, and caching concepts.
  • Experience with data quality and synthesis techniques such as programmatic labeling, weak supervision, MinHash/LSH, and LLM-generated synthetic data.
  • Experience with pipeline orchestration and dataset versioning tools such as Airflow, Databricks Workflows, Dagster, Prefect, and Unity Catalog.
  • Experience handling regulated or sensitive data under controlled access, including HIPAA or equivalent, and familiarity with de-identification concepts.
  • Git-based version control and CI/CD experience for data and code.
  • Strong written documentation skills and the ability to elicit requirements from technical and non-technical experts.
  • Bachelor’s degree in computer science, data science, engineering, statistics, or a related field, or equivalent practical experience.
  • Preferred: Hands-on EHR data experience in skilled nursing, long-term care, post-acute care, or senior living.
  • Preferred: Working knowledge of clinical terminologies and data standards such as ICD-10, SNOMED CT, LOINC, HL7v2, FHIR, and CCDA.
  • Preferred: dbt experience for transformation and testing.
  • Preferred: Familiarity with training-side ML frameworks such as PyTorch to debug data-side bottlenecks.
  • Preferred: Experience supporting LLM or foundation-model training or fine-tuning data pipelines.
  • Preferred: Clinical NLP, OCR, document parsing, or ASR/transcript pipeline experience.
  • Preferred: Experience with data lineage and catalog tools.
  • Preferred: Prior experience embedded inside an AI or ML research team.
  • Preferred: Master’s degree in a relevant quantitative or computer science field.

Benefits

  • Benefits starting from day 1.
  • Retirement plan matching.
  • Flexible paid time off.
  • Wellness support programs and resources.
  • Parental and caregiver leaves.
  • Fertility and adoption support.
  • Continuous development support program.
  • Employee assistance program.
  • Allyship and inclusion communities.
  • Employee recognition and more.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Data Engineer

Cortex 11-50 Internet Software & Services

Cortex is hiring a Senior Data Engineer to own internal data systems and pipelines that power company-wide reporting, product analytics, and AI-driven operations from a fully remote US-based team.

ClickHouse dbt GCP OpenTelemetry Python Segment SQL
16 hours, 20 minutes ago

Data Solutions Engineer

Standard Metrics 11-50 Capital Markets

Standard Metrics is hiring a Data Solutions Engineer to own customer-facing data workflows, integrations, and AI-driven automation for a financial data platform serving investors and portfolio companies.

dbt LLM Python REST API Snowflake SQL
1 day, 16 hours ago

Senior Security Research Engineer

PlayStation 100K+ Household Durables

Sony Interactive Entertainment is hiring a Senior Security Research Engineer to lead Global Vulnerability Management efforts that advance its Threat Exposure Management program and strengthen security risk reduction across the PlayStation ecosystem.

Bash CI/CD Domo Git JavaScript PowerShell Python Snowflake SQL
2 days, 16 hours ago

Senior Data Engineer (R13923)

Oportun 1K-5K Banks

Oportun is hiring a Senior Data Engineer to design and deliver scalable data platforms and solutions that support business and analytical needs across cross-functional initiatives.

Agile Apache Airflow Apache Spark AWS Azure Databricks GCP Hadoop Java Jenkins Kafka Kanban MariaDB PostgreSQL Python Scala Scrum SQL
2 days, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers