Sr Data Engineer

McGraw Hill - United States

Hiring: Sr Data Engineer Company: McGraw Hill Location: United States Job Posted Time: 2026-09-16 10:24:26 Employment Type: Remote Target Skills & Keywords : AWS, Agile, Airflow, CI/CD, Data Lakehouse, Data Pipeline, Databricks, Delta Lake, ELT, ETL, Git, IAM, Iceberg, Jira, Kanban, Lambda, MLflow, PySpark, Python, Redshift, S3, SQL, Scala, Shell, Snowflake, Spark, Step Functions, Terraform About the job Experience: •3+ years of experience working with cloud platforms — primarily AWS — architecting and operating Databricks environments including workspace configuration, cluster policies, instance profiles, and cost optimization strategies. •1+ years of experience with workflow automation and pipeline orchestration using Databricks Workflows, Apache Airflow (with the Databricks provider), or equivalent cloud-native schedulers, replacing traditional Unix shell scripting with scalable, observable pipeline management. Required Skills: •Senior Data Engineer must have prior hands-on experience designing and delivering data solutions on Databricks, including building and maintaining lakehouses using Delta Lake with a medallion (Bronze/Silver/Gold) architecture. •Strong knowledge working with data from financial and operational systems, with proven experience implementing Slowly Changing Dimensions (SCD Types 1, 2, and 3) using Delta Lake MERGE operations and Databricks SQL within a unified lakehouse model. •Strong experience with Git-based version control integrated into Databricks (Databricks Repos / Git folders) and project management tools such as Jira, operating within Agile/Kanban delivery frameworks. •Strong experience with modern data architecture principles, including Unity Catalog for data governance, Delta Sharing, and cloud-native lakehouse design patterns on AWS with Databricks. •Demonstrated capacity to translate business requirements into technical designs and deliver production-grade data solutions within Databricks, from initial scoping through deployment. •Design and develop parallel and distributed ETL/ELT pipelines using Apache Spark (PySpark/Scala) on Databricks, applying partitioning, caching, and broadcast join strategies for optimal resource efficiency and throughput. •Understand data mapping and transformation requirements and implement them using Databricks-native constructs including Spark transformations (aggregations, joins, unions, window functions, lookups, and pivot/unpivot operations) and Delta Live Tables (DLT) for declarative pipeline development. •Develop and maintain Databricks Workflows and job orchestration logic (including dependency management, retry policies, and alerting), replacing traditional shell-based wrapper patterns with cloud-native, maintainable pipeline automation. Qualifications: •Deep expertise in modern data lakehouse architecture, including Delta Lake, medallion design patterns, Unity Catalog governance, and the transition from traditional data warehousing to cloud-native lakehouse solutions on Databricks. •Databricks — Delta Live Tables (DLT), Databricks Workflows, Unity Catalog, Delta Lake (MERGE, OPTIMIZE, VACUUM, Z-ordering), Databricks SQL, and MLflow •AWS services — S3, Redshift, Glue, Lambda, EMR, Athena (with Iceberg), Step Functions, and IAM — integrated with Databricks as the primary compute and transformation layer •Scripting and programming languages — Python (PySpark), Scala (Spark), or SQL as primary languages for pipeline development and data transformation within Databricks Compensation: •$135,000 - $160,000 / year •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!