How Remoteville checks and expires listings
Sr./Staff Data Engineer
Skills
SQLAirflowApache SparkCustomer DataData PipelinesData ScienceDistributed Computing
What the job involves
The main requirements, responsibilities and hiring steps.
Requirements
- Deep experience as a hands-on Data Engineer building production data pipelines
- Experience managing the delivery of complex data
- Experience in ETL orchestration and workflow management tools preferably Apache Airflow
- Experience in Spark or other distributed computing frameworks
- SQL and Python experience
- Advanced SQL performance tuning
- Knowledge of Kubernetes and building Docker images
- Experience in AWS & GCP
- Experience working with APIs to collect or ingest data
- Manage SLA for all pipelines in allocated areas of ownership
- Experience with streaming technologies like Kafka, Spark streaming
- Experience with ELK stack, Grafana
Day to day
- Understand all aspects of a business problem including those unrelated to their area of expertise, weigh pros and cons of different approaches and suggest ones likely to succeed
- Work with cross-functional organization including engineering, delivering, subject-matter experts, product managers, as well as platform engineers to deliver a scalable framework
- Map customer data into Machinify canonical form, identify and ingest non-canonical fields and generalize the process to a minimal level of customization
- Proactively design and adapt the canonical form to suit changing query patterns and needs
- Ultimately own data availability and quality for the Data Science organization
