Design and build scalable data lakes, warehouses, and lakehouse architectures for a thematic research platform processing large volumes of financial data daily. Implement Python-based ETL/ELT pipelines, orchestrate workflows with Airflow, and develop ingestion workflows from third-party APIs. Must have 5+ years Python experience, 5+ years data processing with Pandas/Polars/PySpark/DuckDB, and 2+ years Big Data technologies (Spark, Snowflake).
Key Highlights
Design and build scalable data lakes, warehouses, and lakehouse architectures
Implement Python-based ETL/ELT pipelines with Airflow orchestration
Work with Snowflake, Spark, and AWS for high-performance data infrastructure
Key Responsibilities
Design and implement Python Data Engineering solutions
Design and build scalable Data Lakes, Data Warehouses, and Data Lakehouses
Design and implement robust ETL/ELT processes at scale using Python with Airflow orchestration
Develop sophisticated ingestion workflows from diverse third-party APIs and data sources
Manage and optimize various file formats (Parquet, Avro, ORC) and columnar storage for high-performance data retrieval
Work with AI development tools to support machine learning initiatives and advanced analytics
Act as a technical consultant for stakeholders and leadership to gather requirements and translate business goals into technical roadmaps
Work with Terraform and other tools to build AWS and on-prem infrastructure
Interested in remote work opportunities in Data Science? Discover Data Science Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
Benefits & Perks
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Flextime: Flexible schedule with remote and office options
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
Nice to Have
Familiarity with the fintech industry
Documentation skills for data pipelines and architecture designs
OpenSearch, Elasticsearch
AWS Sagemaker Studio, Jupyter for data analysis
Terraform
Scala