K

Senior Data Infrastructure Engineer - Kafka, Kubernetes, and Data Pipelines

Visa Sponsorship Relocation
Apply
AI Summary

Own and operate Kaufland’s self-hosted Apache Kafka backbone, Debezium CDC pipelines, and Kubernetes-based data infrastructure. Drive automation, observability, and cost management while supporting triage and senior-level design reviews. Lead multi-quarter initiatives with a focus on scalability and reliability.

Key Highlights
End-to-end ownership of Apache Kafka, Schema Registry, and Debezium-based CDC pipelines
Hands-on Kubernetes operations for Navarch workflow orchestrator (~7,500 tasks/day)
Automate infrastructure, access management, and observability with a cost-conscious mindset
Key Responsibilities
Run and own self-hosted Apache Kafka backbone and Schema Registry, including ACLs, client quotas, certificates, upgrades, and disaster recovery
Operate Debezium-based change data capture pipelines, managing connectors, snapshots, offsets, replica lag, and BigQuery sink loaders
Run workloads on Kubernetes, including Navarch workflow orchestrator, owning resource management, autoscaling, and debugging
Automate access management and infrastructure as code across self-hosted GitLab CI/CD pipelines
Own observability and infrastructure costs using Datadog, Slack-native alerting, and cost tracking for K8s, Kafka retention, and log volumes
Carry triage duties one week in three, handling ingestion requests, access requests, schema alerts, and supporting analysts
Act as a senior sparring partner, driving design reviews and multi-quarter initiatives alongside two senior data engineers
Participate in 24/7 on-call rotation (optional)
Technical Skills Required
Apache Kafka Kubernetes Python
Benefits & Perks
Flexible remote work within Germany or on-site in Cologne, Darmstadt, Düsseldorf, or Berlin
Relocation package for international hires
30 days of vacation per year and sabbatical options
Nice to Have
Object-oriented design and data modeling/dbt
BigQuery architecture expertise (access frameworks, partitioning, slot strategy)
Orchestration frameworks (Airflow, Dagster) or stream processing (Beam, Flink, Spark)
Internal automation tools (n8n, AI triage bots)
Terraform, Ansible, Pulumi, Helm, or other configuration management tools
Google Cloud Platform (GCP) experience

Job Description


Your tasks – this is what awaits you in detail

  • You run and own our self-hosted Apache Kafka backbone and Schema Registry end-to-end, managing ACLs, client quotas, certificates, upgrades, and disaster recovery
  • You operate our Debezium-based change data capture pipelines, overseeing connectors, snapshots, offsets, replica lag, and BigQuery sink loaders
  • You run our workloads on Kubernetes, including Navarch – our self-built workflow orchestrator executing ~7,500 SQL and Python tasks daily – owning resource management, autoscaling, and debugging
  • You automate and own access management and infrastructure as code across our self-hosted GitLab CI/CD pipelines, replacing manual tickets with code-driven policies
  • You own observability and infra costs – Datadog across Kafka and sink loaders, Slack-native alerting, and maintaining an honest view of K8s, Kafka retention, and log volume expenses
  • You carry your share of triage, one week in three – handling ingestion and access requests, schema alerts, and supporting analysts, while helping automate this layer via AI-assisted review
  • You act as a senior sparring partner alongside two senior data engineers (reporting to the Head of Data & Analytics) – driving design reviews and multi-quarter initiatives, with room to grow into BigQuery architecture over time: access frameworks, partitioning and slot strategy, the pipelines and models on top
  • Optional: You take part in our 24/7 on-call rotation

Your profile – this is what we expect from you

  • You have several years of experience operating production infrastructure you were personally responsible for, including on-call rotation. You have run systems that broke, can explain what went wrong and how you permanently fixed it, and can detail a specific trade-off you navigated backed by concrete numbers
  • You have deep hands-on expertise in Apache Kafka broker operations, including cluster sizing, partition rebalancing, replication, ISR behavior, and version upgrades
  • You have operated cross-system data replication in production – ideally Debezium on Kafka Connect, or database replication via binlog, GTID, WAL, or agent-based pipelines
  • You demonstrate strong production experience with Kubernetes, master resource management and autoscaling, and possess sharp debugging reflexes for complex failure modes
  • You manage infrastructure and access as code across multiple environments (ideally Terraform and Ansible; Pulumi, Helm or serious configuration management also counts) following a strict least-privilege IAM approach across multiple environments, ideally with GCP experience or transferable AWS/Azure knowledge
  • You use production-grade Python and SQL as practical tools, applying a cost-conscious mindset to query performance, scanning volume, and data pruning
  • You thrive in a low-process, Kanban-driven environment where seniors self-organize, rotate weekly triage duties without micro-management, and prefer automating recurring requests over manually repeating them
  • Optional: You bring skills or interest in object-oriented design, data modeling/dbt, BigQuery architecture, orchestration frameworks (Airflow, Dagster), stream processing (Beam, Flink, Spark), or internal automation (n8n, AI triage bots).
  • You are fluent in English (C1 level) and enjoy working in an international team environment

What we offer

  • Create your own work-life balance: You have the flexibility to choose between working remotely (within Germany) or from one of our locations in Cologne, Darmstadt, Düsseldorf, or in Berlin!
  • Do you want to move to Germany? No problem – we offer you an attractive relocation package to give you a smooth start.
  • Urban Sports Club and RSG Group Fitness Studios: Get top deals for fitness, swimming, yoga and more
  • Mental Well-Being: We support you on your well-being journey with special offerings such as Instahelp, the digital platform for online psychological counseling
  • Vacation & Sabbatical: Enjoy 30 days of vacation per year and the opportunity to take a sabbatical once you have been part of the team for a certain period!
  • Option for Pluxee restaurant vouchers: Buy Pluxee vouchers through us and benefit from tax-advantaged meal allowances!
  • ‘Deutschlandticket’: We subsidize your train season ticket for more mobility
  • Employee Discount: You will receive a monthly coupon for Kaufland.de
  • Free choice of operating system: MacOS or Ubuntu Linux, it’s up to you
  • Boost your growth: Benefit from our online language learning programs, diverse in-house training, and our automated 360-degree feedback. We cover the costs for relevant conferences, training opportunities, and approved team workshops to strengthen personal interactions
  • This is who we are: Our dynamic culture combines flat hierarchies, a start-up mentality, an international team of over 65 nationalities, and the strength of the Schwarz Group to provide you an agile and secure working environment.
  • We are team players: From day one, you will connect not only with your team but also with others through our digital onboarding journey, the buddy program, and regular team and company events, all-hands meetings, powerful mornings, and much more!

Check out our Principles & our blog for even more insights into our company culture!

Who we are

We are the tech powerhouse behind Kaufland’s international online marketplaces and a company within the Schwarz Group. Driven by a start-up mentality and working on an equal footing within flat hierarchies, we develop the best marketplace solutions for our customers, sellers and partners. #FromEuropeForEurope

Day by day, our Tech & Product Team of about 400 experts pursues the goal of creating the best possible customer shopping experience for our online marketplace. Through innovative technical developments, they not only create an outstanding shopping experience but also lay the foundation for an optimal selling experience for our sellers. Learn more about our division, areas, and cross-functional teams here!

Similar Jobs

Explore other opportunities that match your interests

Senior Systems Engineer – AWS European Sovereign Cloud (ESC) Managed Operations

Devops
6d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Amazon Web Services (AWS)

Germany

AWS European Sovereign Cloud Support Engineer

Devops
1w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Amazon Web Services (AWS)

Germany
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

nicsell

Germany

Subscribe our newsletter

New Things Will Always Update Regularly