P

Senior Data Engineer

pubx • United Kingdom
Remote
Apply
AI Summary

Design and maintain high-volume data pipelines, build event-driven components, and develop data models to support agentic AI features and core product workflows.

Key Highlights
Design and maintain high-volume data pipelines
Build event-driven components using Kafka and message queues
Develop data models and transformation layers to support analytics and ML/AI consumption
Key Responsibilities
Design and maintain high-volume data pipelines (batch + streaming) powering agentic AI features and core product workflows
Build event-driven components using Kafka and message queues, including idempotency patterns, replay strategies, and backfill mechanisms
Develop data models and transformation layers (lakehouse patterns, dbt-style modeling) supporting both analytics and ML/AI consumption
Own data quality and reliability: schema management, validation, lineage, SLAs, and incident response
Enable AI/ML workflows with robust datasets for training, evaluation, feature generation, and feedback loops from production agents
Deploy and operate data infrastructure on AWS using infrastructure-as-code
Technical Skills Required
Python Kafka SQL
Benefits & Perks
Remote work
Salary

Job Description


Why This Role Exists

PubX builds next-generation publisher-first agentic advertising marketplace infrastructure. Our AI makes real-time, revenue-critical decisions for digital publishers and Advertisers. Our Bid Intelligence uses machine learning to optimize every programmatic ad auction individually, generating measurable revenue uplift for publishers. We've priced over 1 trillion programmatic auctions, we’re currently ranked #5 globally in Prebid Analytics Adapter Rankings, and growing.


The problem we’re solving

Digital publishers leave significant revenue on the table because ad pricing is still largely manual, static or simple rule-based. Every ad impression is unique, but most pricing systems treat them the same. PubX's AI analyzes bid-stream data and historical patterns to arrange optimal deals, in near real-time.


As a founding member of AgenticAdvertising.org, we're building the next generation of autonomous advertising infrastructure.


What You’ll Work On

  • Design and maintain high-volume data pipelines (batch + streaming) powering agentic AI features and core product workflows
  • Build event-driven components using Kafka and message queues, including idempotency patterns, replay strategies, and backfill mechanisms
  • Develop data models and transformation layers (lakehouse patterns, dbt-style modeling) supporting both analytics and ML/AI consumption
  • Own data quality and reliability: schema management, validation, lineage, SLAs, and incident response
  • Enable AI/ML workflows with robust datasets for training, evaluation, feature generation, and feedback loops from production agents
  • Deploy and operate data infrastructure on AWS using infrastructure-as-code


What We’re Looking For

We’re looking for an experienced engineer who has worked on production systems and enjoys solving practical problems with AI.


You’ve likely have:

  • Strong data engineering fundamentals: data modeling, partitioning, performance tuning, and cost-aware design for high-volume workloads
  • Experience building streaming and event-driven systems (Kafka/queues), including handling late/out-of-order events, backfills, and real-world data edge cases
  • Strong SQL + Python skills, and comfort with modern data stack tooling (e.g., Spark, Airflow/Dagster, dbt, warehouse/lakehouse patterns)
  • Hands-on AWS experience with production operations for data systems: monitoring, incident response, and security considerations (PII, access control, encryption, auditability)
  • Familiarity integrating data with AI/ML and agentic systems: feature pipelines, evaluation datasets, grounding/citations inputs, and feedback capture from agent outcomes


You tend to:

  • Make pragmatic decisions balancing speed, quality, cost, and risk trade-offs.
  • Communicate technical ideas well in writing and conversation to both technical and non-technical audiences.
  • Write clean, well-tested code with thoughtful abstractions that’s easy to extend and operate.
  • Learn quickly when things are unfamiliar by prototyping, then hardening and documenting what you ship.


Bonus (not required):

  • Experience with AdTech or other high volume real-time systems


Who This Role Will Suit

This role suits engineers who like a mix of autonomy and collaboration, and who are comfortable working in an environment that’s still evolving.


We’re a distributed team with a growing engineering presence in India, so comfort with async collaboration and clear written communication is important.


We use agentic coding tools heavily (e.g. Cursor and Claude Code) to plan, scaffold, refactor, and debug production code, while maintaining strong engineering judgment and ownership of outcomes.


If you’re interested in building and shaping real systems in a growing product company, at the forefront of AdTech innovation, we’d love to hear from you.


We will process your personal data in accordance with our Recruitment Privacy Notice: https://pubx.ai/privacy/recruitment/


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Haystack

United Kingdom
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

link xperts

United Kingdom

Contract Business Analyst

Data Science
•
2d ago
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Associate

Jobgether

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly