Data Engineer
Skills
Computer ScienceApache SparkBig DataData ArchitectureDatabasesEasily AdaptableElastic Stack
What the job involves
The main requirements, responsibilities and hiring steps.
Requirements
- Around five years of professional experience in software engineering data engineering or a closely related field
- At least three years of hands-on experience with Spark and Scala using DataFrame and Dataset APIs
- Proven experience tuning Spark applications at multi-terabyte or petabyte scale
- Strong debugging and problem-solving experience in the big data ecosystem
- Solid relational database experience
- Working knowledge of AWS EMR S3 Lambda Kafka Snowflake Grafana Hadoop Elastic Stack and Docker
- Strong grasp of the full software development lifecycle and solution architecture data structures and data modeling
- Enthusiasm for using generative AI to accelerate data engineering work
Nice to have
- Gen AI experience
- Probabilistic data structures
- High cardinality systems
- Bachelor's degree
Day to day
- Design and maintain ETL and ELT pipelines that power the business at terabyte and petabyte scale on AWS EMR.
- Build and tune Spark and Scala data workflows across S3 and Snowflake, solving performance and cost issues that only appear at very large volumes.
- Monitor pipeline health, investigate data quality issues, and support production operations including incident response and on-call coverage.
