Data Engineer

Skills
Computer ScienceApache SparkBig DataData ArchitectureDatabasesEasily AdaptableElastic Stack
Role

What the job involves

The main requirements, responsibilities and hiring steps.

Requirements

  • Around five years of professional experience in software engineering data engineering or a closely related field
  • At least three years of hands-on experience with Spark and Scala using DataFrame and Dataset APIs
  • Proven experience tuning Spark applications at multi-terabyte or petabyte scale
  • Strong debugging and problem-solving experience in the big data ecosystem
  • Solid relational database experience
  • Working knowledge of AWS EMR S3 Lambda Kafka Snowflake Grafana Hadoop Elastic Stack and Docker
  • Strong grasp of the full software development lifecycle and solution architecture data structures and data modeling
  • Enthusiasm for using generative AI to accelerate data engineering work

Nice to have

  • Gen AI experience
  • Probabilistic data structures
  • High cardinality systems
  • Bachelor's degree

Day to day

  • Design and maintain ETL and ELT pipelines that power the business at terabyte and petabyte scale on AWS EMR.
  • Build and tune Spark and Scala data workflows across S3 and Snowflake, solving performance and cost issues that only appear at very large volumes.
  • Monitor pipeline health, investigate data quality issues, and support production operations including incident response and on-call coverage.