RevStar

RevStar

RevStar Consulting is a client-centric cloud consulting firm that specializes in developing modern, user-focused, cloud-native web and mobile applications. They offer custom integrations, implementations, and solutions using the latest cloud technologi...

Internet Software & Services
51-250
Founded 2009

Description

  • Develop and optimize data pipelines using Apache Spark and Delta Lake within Databricks.
  • Implement ETL/ELT workflows for data ingestion, transformation, and storage.
  • Design scalable Lakehouse architecture solutions across structured and unstructured data sources.
  • Integrate Databricks with cloud storage platforms such as Azure Data Lake, AWS S3, and Google Cloud Storage.
  • Optimize Spark jobs for scalability, cost efficiency, and low latency.
  • Implement monitoring, alerting, automated validation, and data quality processes.
  • Support ML model training and deployment within Databricks using MLflow for tracking and versioning.
  • Collaborate with data scientists, ML engineers, architects, and business stakeholders to deliver aligned solutions.
  • Implement feature engineering pipelines and help integrate models into production environments.
  • Ensure data security, access control, governance, compliance, and documentation best practices.

Requirements

  • 3+ years of hands-on experience in data engineering with big data processing and cloud-native architectures.
  • 2+ years of hands-on experience with Databricks, including Apache Spark, Delta Lake, and MLflow.
  • Databricks Certified Data Engineer Associate or higher certification is mandatory.
  • Proficiency in Python, SQL, and Spark-based frameworks.
  • Experience developing and optimizing large-scale ETL/ELT pipelines.
  • Strong understanding of Lakehouse architecture and cloud-agnostic data solutions.
  • Familiarity with CI/CD pipelines and Infrastructure-as-Code tools for Databricks, such as Terraform and Databricks CLI.
  • Knowledge of data governance, security, and compliance best practices.
  • Experience working in Agile environments and following DevOps/MLOps best practices.
  • Preferred: additional Databricks certifications, such as Databricks Certified Machine Learning Associate.
  • Preferred: experience with real-time streaming tools such as Kafka, Kinesis, or Event Hub.
  • Preferred: familiarity with orchestration tools such as Apache Airflow or Prefect.
  • Preferred: background in AI/ML integration within Databricks, including feature engineering and model deployment.
  • Preferred: experience in client-facing roles or consulting environments.

Benefits

  • Paid time off.
  • Remote-first working environment.
  • Comprehensive health coverage including medical, dental, and vision.
  • 401(k) retirement plan.
  • Annual learning and development stipend for conferences, certifications, or courses.
  • Peer mentorship and coaching.
  • Professional growth opportunities with exposure to AWS GenAI, data, and cloud technologies.
  • Company outings and volunteer opportunities.
  • Collaborative, innovative culture.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Boomi Integration Developer

Coforge 10K-50K IT Services

Coforge is seeking a Senior Boomi Integration Developer in Brazil to build, deploy, and support enterprise integrations across cloud applications, business systems, databases, APIs, and file-based interfaces in a remote role.

Agile Java JSON JWT REST API SAP SQL XML
2 minutes ago

Manager, Data Platform

ZoomInfo 1K-5K Professional Services

ZoomInfo is seeking a Manager, Data Platform to lead the architecture, delivery, and reliability of enterprise data infrastructure supporting analytics, machine learning, operational systems, and AI applications while managing and developing a team of platform engineers.

Apache Airflow Apache Spark AWS Databricks dbt Prototyping Python Snowflake SQL
1 day, 23 hours ago

Data Engineer, AI & Analytics

Power Digital is hiring a Data Engineer to build and operate unified, AI-ready data foundations—including pipelines, models, and semantic layers—for its marketing agency, clients, products, and AI initiatives.

CI/CD dbt GCP Git Google Ads LinkedIn Ads Python Shopify Snowflake SQL TikTok
2 days ago

Data Engineer, AI & Analytics

Power Digital is hiring a Data Engineer/Analytics Engineer to own the end-to-end data foundation supporting agency teams, clients, products, and AI initiatives across a complex multi-client marketing data environment.

CI/CD dbt GCP Git Google Ads LinkedIn Ads Python Shopify Snowflake SQL TikTok
2 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers