C

Senior Data Engineer

Conexess Group • United State
Remote
Apply
AI Summary

Join our enterprise data platform team as a Senior Data Engineer to build and operate a multi-account AWS environment. Design, build, and optimize data pipelines using AWS Glue, PySpark, and Apache Iceberg. Partner with data product teams to onboard new data sources and establish data contracts.

Key Highlights
Design, build, and optimize data pipelines using AWS Glue, PySpark, and Apache Iceberg
Partner with data product teams to onboard new data sources and establish data contracts
Troubleshoot AWS Glue jobs, DMS replication issues, and data-freshness problems
Key Responsibilities
Design, build, and optimize data pipelines using AWS Glue, PySpark, and Apache Iceberg
Own data product delivery from source-system ingestion through transformation, quality validation, and governed consumption
Partner with data product teams to onboard new data sources, define schemas, and establish data contracts
Technical Skills Required
Amazon Web Services SQL Python
Benefits & Perks
Fully remote opportunity
Long-term contract
Nice to Have
Data Mesh architecture and federated governance
Lakehouse architecture patterns
AWS DMS and change-data-capture ingestion

Job Description



Senior Data Engineer - Remote Contract

Our client is seeking a Senior Data Engineer to join its enterprise data platform team. This team builds and operates a multi-account AWS environment supporting more than a dozen data product teams across the organization. The platform spans data ingestion, transformation, governance and analytics using AWS Glue, Athena, Lake Formation, Redshift, S3 and Apache Iceberg.

This is a fully remote, long-term opportunity for an experienced engineer who can quickly take ownership of production data pipelines and analytics datasets.

What You’ll Do
  • Design, build and optimize data pipelines using AWS Glue, PySpark and Apache Iceberg
  • Own data product delivery from source-system ingestion through transformation, quality validation and governed consumption
  • Build and maintain Amazon Redshift integrations, including schemas, stored procedures, materialized views and Liquibase-managed DDL migrations
  • Write and optimize complex SQL for analytics transformations, reporting views and data-quality checks
  • Translate operational data models into star schemas and other dimensional models for analytics
  • Manage Apache Iceberg tables, including partition and schema evolution, compaction and orphan-file cleanup
  • Implement table- and column-level governance using Lake Formation permissions, tag-based access control and PII classification
  • Develop data-quality frameworks covering referential integrity, row-count reconciliation, completeness and anomaly detection
  • Partner with data product teams to onboard new data sources, define schemas and establish data contracts
  • Troubleshoot AWS Glue jobs, DMS replication issues and data-freshness problems
  • Own the complete lifecycle of changes, including development, deployment, validation and communication
  • Ensure pipelines and datasets remain accurate, reliable and healthy across environments
Required Qualifications
  • At least five years of hands-on data engineering experience building ETL/ELT pipelines and analytics platforms
  • Expert-level SQL skills, including complex joins, CTEs, window functions, query optimization and performance tuning
  • Strong Python and PySpark development experience
  • Hands-on AWS Glue experience, including job development, bookmarks, crawlers and performance optimization
  • Strong Amazon Redshift experience, including schema design, stored procedures, Spectrum, external schemas, materialized views and performance tuning
  • Experience with Apache Iceberg or another open table format, including ACID transactions, schema evolution, partition evolution and table maintenance
  • Strong understanding of dimensional data modeling, including star schemas, snowflake schemas and slowly changing dimensions
  • Experience with AWS data services such as Athena, Lake Formation, S3, DMS, Lambda and EventBridge
  • Understanding of data governance, tag-based access control, data classification and PII handling
  • Experience with Liquibase or a similar database migration tool
  • Familiarity with GitHub and CI/CD pipelines
  • Strong attention to data correctness, completeness and freshness
  • Ability to independently deliver production-ready data pipelines with limited oversight
Preferred Experience
  • Data Mesh architecture and federated governance
  • Lakehouse architecture patterns
  • AWS DMS and change-data-capture ingestion
  • Terraform or other infrastructure-as-code tools
  • Data catalog and lineage platforms such as AWS DataZone or IDERA ER/Studio
  • AI-assisted development tools such as Amazon Kiro or GitHub Copilot
  • Amazon SageMaker or other machine-learning platforms
  • Docker and container-based development
  • Agriculture, retail or other large-scale operational data environments

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

silicon data

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Russell Tobin

United State

Data Science Lead - AI and Machine Learning

Data Science
•
16h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

BNSF Railway

United State

Subscribe our newsletter

New Things Will Always Update Regularly