Senior Data Engineer - Databricks, Healthcare, HIPAA - REMOTE
CyberCoders - United States
Hiring: Senior Data Engineer - Databricks, Healthcare, HIPAA - REMOTE Company: CyberCoders Location: United States Job Posted Time: 2026-09-16 22:20:16 Employment Type: Contract / Remote Target Skills & Keywords : AI, AWS, Airflow, CI/CD, Data Lake, Data Pipeline, Databricks, Delta Lake, Docker, DynamoDB, EKS, ELT, EMR, ETL, Event-Driven, Feature Store, HIPAA, Java, Kafka, Kinesis, Kubernetes, Lambda, Machine Learning, NLP, PySpark, Python, RAG, Redshift, Regulatory Compliance, S3, SQL, Scala, Spark, Step Functions, dbt About the job Experience: •7+ years of professional data engineering experience. •3+ years working with cloud-native data lake or lakehouse architectures. Required Skills: •Design and develop scalable batch and real-time data pipelines using Databricks, Spark, Delta Lake, and AWS services. •Lead the evolution of a centralized lakehouse platform, including data modeling, governance, optimization, metadata management, and metrics layers. •Build and support ETL/ELT workflows utilizing Airflow, dbt, AWS Step Functions, and modern orchestration frameworks. •Develop streaming data solutions using Kafka and Kinesis for high-volume event processing. •Partner with Data Science and AI teams to support feature stores, model scoring pipelines, and machine learning enablement. •Build infrastructure supporting NLP and generative AI applications, including vector databases, document ingestion pipelines, knowledge repositories, and RAG-based architectures. •Develop APIs and data integrations that support business applications, analytics platforms, CRM systems, and external data exchange. •Implement data security, governance, privacy controls, audit logging, and regulatory compliance standards. Qualifications: •Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, or a related technical discipline. •Prior experience in healthcare, health technology, insurance, financial services, or other regulated industries preferred. •, including Delta Lake, Unity Catalog, Spark, and Workflows. •Advanced experience with Apache Spark using PySpark, Scala, or Java. •Strong AWS experience including S3, Redshift, Athena, Glue, Lambda, Step Functions, Kinesis, DynamoDB, EMR, Bedrock, Transcribe, and Comprehend Medical. •Proficiency in Python and SQL; Java or Scala experience is a plus. •Knowledge of Kafka, Kinesis, and event-driven architectures. •Operational familiarity with vector databases, embedding stores, document processing pipelines, and RAG architectures. Compensation: •$135,000 - $190,000 / year •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!