How Remoteville checks and expires listings

Lead Data Scientist

Skills
Computer ScienceApache SparkDatabasesDocument ProcessingDriftISO StandardsSemantic Analysis
Role

What the job involves

The main requirements, responsibilities and hiring steps.

Requirements

  • 7+ years in data science ML engineering or related roles
  • 3+ years building NLP or generative AI applications and implementing MLOps in production
  • Bachelor's or Master's degree in Data Science Computer Science Statistics or related field
  • Strong Python experience with Pandas NumPy scikit-learn XGBoost TensorFlow PyTorch Hugging Face Transformers FastAPI Flask MLflow and pytest
  • Advanced SQL proficiency with complex queries window functions and optimization
  • Strong foundation in supervised and unsupervised learning deep learning document understanding text classification and semantic analysis
  • Hands-on experience with foundation models prompt engineering RAG architectures and vector databases
  • End-to-end experience with ML pipelines experiment tracking model versioning feature stores drift detection CI/CD for ML and Docker containerization
  • Experience with evaluation frameworks custom metrics benchmark datasets and human-in-the-loop validation
  • Experience with AWS services including SageMaker Bedrock S3 Lambda EC2 and CloudWatch
  • Strong foundation in statistics A/B testing causal inference and experimental design
  • Proficiency with Tableau Power BI or Python visualization libraries
  • Track record of deploying ML systems processing large-scale datasets with monitoring and governance
  • Must be able to work without visa sponsorship

Nice to have

  • Agentic AI frameworks
  • Life Sciences regulated industries
  • Big data tools
  • LLM fine-tuning
  • ML governance
  • Agile environments
  • AWS ML certifications

Day to day

  • Build predictive models and deploy Generative AI and Agentic AI features for a document-based compliance platform.
  • Architect data-driven solutions that automate compliance workflows document review and regulatory mapping with robust production systems.
  • Develop end-to-end MLOps pipelines and production Python services while collaborating with cross-functional teams to translate business needs into ML solutions.