N

Senior Site Reliability Engineer - Europe's Sovereign AI Cloud

nebul Netherlands
Visa Sponsorship
Apply
AI Summary

Ensure Europe's sovereign AI cloud's reliability and scalability. Troubleshoot complex issues across Kubernetes, NVIDIA GPU infrastructure, and cloud-native services. Automate operational tasks and improve platform stability.

Key Highlights
70% site reliability and operational troubleshooting
30% broader DevOps and platform engineering
Reduce dependency on senior engineers
Key Responsibilities
Monitor and improve platform reliability, availability, and performance
Troubleshoot incidents across Kubernetes, NVIDIA GPU infrastructure, Linux, networking, and platform services
Investigate issues affecting services written in Go and Python
Perform root-cause analysis and translate findings into platform improvements
Build and improve monitoring, metrics, logging, tracing, and alerting
Automate repetitive operational tasks using Go, Python, or scripting
Support Kubernetes cluster operations, deployments, upgrades, and platform changes
Improve the production readiness of new services and infrastructure components
Identify reliability risks and reduce operational dependency on senior engineers
Technical Skills Required
Kubernetes Linux Python Go
Benefits & Perks
Visa sponsorship available
Nice to Have
Experience with NVIDIA GPU infrastructure
Experience supporting AI, machine-learning, or high-performance computing workloads
Knowledge of Kubernetes GPU scheduling and resource management

Job Description


About Nebul

At Nebul, we’re building Europe’s sovereign AI cloud — trusted, secure, and purpose-built for the next generation of intelligent infrastructure.

Our platform combines Kubernetes, NVIDIA GPU infrastructure, cloud-native services and software written in Go and Python. As the platform grows, we need to increase our internal reliability capability and reduce the number of operational issues that depend on a small group of senior engineers.


What You’ll Be Doing

As a Site Reliability Engineer, you’ll work approximately 70% on site reliability and operational troubleshooting and 30% on broader DevOps and platform engineering.

Your primary responsibility will be to keep Nebul’s AI cloud platform stable, observable and operationally scalable. You’ll take ownership of incidents, investigate complex issues and ensure that senior engineers are not pulled into every troubleshooting session.

You’ll work across Kubernetes, NVIDIA GPU infrastructure, services written in Go and Python, networking and cloud-native platform components.

The role is not only about responding to incidents. You’ll identify recurring problems, automate operational work and improve the platform so that the same issues do not continue to return.


Key Responsibilities

  • Monitor and improve the reliability, availability and performance of Nebul’s AI cloud platform.
  • Troubleshoot incidents across Kubernetes, NVIDIA GPU infrastructure, Linux, networking and platform services.
  • Take ownership of technical troubleshooting sessions and coordinate issues through to resolution.
  • Investigate issues affecting services written in Go and Python.
  • Perform root-cause analysis and translate findings into structural platform improvements.
  • Build and improve monitoring, metrics, logging, tracing and alerting.
  • Ensure alerts are relevant, actionable and connected to clear operational procedures.
  • Create runbooks, escalation procedures and troubleshooting documentation.
  • Automate repetitive operational tasks using Go, Python or scripting.
  • Support Kubernetes cluster operations, deployments, upgrades and platform changes.
  • Improve the production readiness of new services and infrastructure components.
  • Work with engineering teams to improve resilience, observability and failure handling.
  • Identify reliability risks before they result in platform or customer impact.
  • Support operational improvements across GPU workloads and NVIDIA-based infrastructure.
  • Reduce the operational dependency on senior platform engineers.
  • Contribute to broader DevOps work when additional capacity is needed within the team.


What Your Day Will Not Look Like

  • Acting as a first-line support engineer who only closes tickets.
  • Escalating every complex issue to Diego or another senior engineer.
  • Spending all your time manually operating Kubernetes.
  • Resolving incidents without addressing their underlying causes.
  • Building isolated automation that is not integrated into the platform.


What You Bring

  • Strong experience as a Site Reliability Engineer, DevOps Engineer or Platform Engineer.
  • Hands-on production experience with Kubernetes.
  • Strong Linux and infrastructure troubleshooting skills.
  • Experience investigating issues across applications, infrastructure, networking and cloud platforms.
  • Experience with services written in Go or Python.
  • Practical experience with monitoring, logging, metrics and alerting.
  • Experience responding to production incidents and performing root-cause analysis.
  • The ability to independently lead complex troubleshooting sessions.
  • Experience automating operational work using Python, Go or scripting.
  • A solid understanding of cloud-native and distributed systems.
  • A calm, analytical and structured approach to incidents.
  • An ownership mindset and the ability to move from reactive troubleshooting to lasting improvements.


Bonus Points If You Have

  • Experience with NVIDIA GPU infrastructure.
  • Experience supporting AI, machine-learning or high-performance computing workloads.
  • Knowledge of Kubernetes GPU scheduling and resource management.
  • Experience with multi-tenant cloud environments.
  • Familiarity with Go-based cloud or platform services.
  • Experience with Infrastructure as Code and automated platform deployment.
  • Knowledge of distributed storage, networking or database troubleshooting.
  • Experience defining service-level indicators, objectives and operational reliability targets.
  • Experience working in sovereign, regulated or security-sensitive cloud environments.


Eligibility & Application Information

We welcome non-native Dutch speakers to apply. However, to be eligible, you must:

  • Have a valid work permit in the Netherlands. ( wo do offer sponsorship if needeed)
  • Reside in the Netherlands and be able to travel to the office in Leiden (near The Hague).
  • Be fluent in English. Dutch is not required.


Ready to make Europe’s sovereign AI cloud more reliable and operationally scalable?

Apply now through Frank Poll and help Nebul build a cloud platform that engineering teams and customers can depend on.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

nebul

Netherlands

Azure Cloud Engineer

Devops
1d ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

GreenFlux

Netherlands
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

A2G Consulting BV (A2G Technol...

Netherlands

Subscribe our newsletter

New Things Will Always Update Regularly