Staff Software Engineer, Data Platform
Harvey - New York, United States
Hiring: Staff Software Engineer, Data Platform Company: Harvey Location: New York, United States Job Posted Time: 2026-09-16 16:50:30 Target Skills & Keywords : AWS, Airbyte, Airflow, Azure, BigQuery, Dagster, Data Warehouse, Databricks, Delta Lake, Fivetran, Flink, GCP, Iceberg, Kafka, Kubernetes, Pulumi, Python, REST, Redshift, SQL, Snowflake, Spark, Terraform, dbt About the job Experience: •10+ years building and operating production data infrastructure, with ownership of systems other teams depend on Required Skills: •Harvey is generating far more data than we currently know how to use well. Product telemetry, agent execution traces, model usage, customer engagement, financial and operational systems — the volume and the number of teams who need to work with it are both growing faster than any single team can serve by hand. •The near-term foundation is ingestion and the warehouse — reliable streaming and batch paths into Snowflake, CDC off production systems, orchestration, and schema evolution that absorbs upstream change instead of breaking under it, and factors in the hard data sensitivity requirements our domain requires. •You'll sit between Analytics, Data Engineering, product teams, and Infrastructure. Today this work is distributed and improvised. You'll make it a system, set the technical direction, and help build the team around you. •This role is based in San Francisco, CA or New York, NY •Own the data platform's architecture and technical direction — treating data infrastructure as a software product built from reusable frameworks, and making deliberate build-vs-buy tradeoffs as the platform grows •Build and operate the ingestion layer across streaming, batch, CDC, and third-party connectors, including schema evolution that absorbs upstream change safely rather than silently breaking consumers, so onboarding a new source is a paved path instead of a project •Land data into Snowflake with the freshness, completeness, and cost characteristics downstream consumers can plan around, and define a clean handoff for Analytics Engineering •Own the orchestration platform — scheduling, retries, backfills, and dependency management across the full data graph Qualifications: •Deep experience with cloud data warehouses — Snowflake strongly preferred (BigQuery, Databricks, or Redshift experience transfers well) — including performance tuning and cost management •Hands-on experience building CDC and streaming pipelines with technologies like Kafka, Debezium, Flink, or Spark Streaming •Strong fluency with workflow orchestration — Temporal, Airflow, Dagster, or similar — operated at scale, not just configured •Strong programming skills in Python and advanced SQL •Practical experience with data quality, observability, and lineage tooling, and with schema evolution in systems that can't afford downtime •Solid functional working knowledge of data governance in a regulated environment: PII classification, masking, access control, retention, and data residency •Operational familiarity with cloud data services (Azure, AWS, GCP), Kubernetes, and infrastructure-as-code (Terraform, Pulumi) •Comfort operating in ambiguity and defining scope where none exists •Exposure to data infrastructure for AI products Compensation: •$231,000 - $340,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!