Ensure Europe's sovereign AI cloud's reliability and scalability. Troubleshoot complex issues across Kubernetes, NVIDIA GPU infrastructure, and cloud-native services. Automate operational tasks and improve platform stability.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
About Nebul
At Nebul, we’re building Europe’s sovereign AI cloud — trusted, secure, and purpose-built for the next generation of intelligent infrastructure.
Our platform combines Kubernetes, NVIDIA GPU infrastructure, cloud-native services and software written in Go and Python. As the platform grows, we need to increase our internal reliability capability and reduce the number of operational issues that depend on a small group of senior engineers.
What You’ll Be Doing
As a Site Reliability Engineer, you’ll work approximately 70% on site reliability and operational troubleshooting and 30% on broader DevOps and platform engineering.
Your primary responsibility will be to keep Nebul’s AI cloud platform stable, observable and operationally scalable. You’ll take ownership of incidents, investigate complex issues and ensure that senior engineers are not pulled into every troubleshooting session.
You’ll work across Kubernetes, NVIDIA GPU infrastructure, services written in Go and Python, networking and cloud-native platform components.
The role is not only about responding to incidents. You’ll identify recurring problems, automate operational work and improve the platform so that the same issues do not continue to return.
Key Responsibilities
- Monitor and improve the reliability, availability and performance of Nebul’s AI cloud platform.
- Troubleshoot incidents across Kubernetes, NVIDIA GPU infrastructure, Linux, networking and platform services.
- Take ownership of technical troubleshooting sessions and coordinate issues through to resolution.
- Investigate issues affecting services written in Go and Python.
- Perform root-cause analysis and translate findings into structural platform improvements.
- Build and improve monitoring, metrics, logging, tracing and alerting.
- Ensure alerts are relevant, actionable and connected to clear operational procedures.
- Create runbooks, escalation procedures and troubleshooting documentation.
- Automate repetitive operational tasks using Go, Python or scripting.
- Support Kubernetes cluster operations, deployments, upgrades and platform changes.
- Improve the production readiness of new services and infrastructure components.
- Work with engineering teams to improve resilience, observability and failure handling.
- Identify reliability risks before they result in platform or customer impact.
- Support operational improvements across GPU workloads and NVIDIA-based infrastructure.
- Reduce the operational dependency on senior platform engineers.
- Contribute to broader DevOps work when additional capacity is needed within the team.
Searching for Devops roles that provide visa sponsorship? Connect with international employers through Devops Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
What Your Day Will Not Look Like
- Acting as a first-line support engineer who only closes tickets.
- Escalating every complex issue to Diego or another senior engineer.
- Spending all your time manually operating Kubernetes.
- Resolving incidents without addressing their underlying causes.
- Building isolated automation that is not integrated into the platform.
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
What You Bring
- Strong experience as a Site Reliability Engineer, DevOps Engineer or Platform Engineer.
- Hands-on production experience with Kubernetes.
- Strong Linux and infrastructure troubleshooting skills.
- Experience investigating issues across applications, infrastructure, networking and cloud platforms.
- Experience with services written in Go or Python.
- Practical experience with monitoring, logging, metrics and alerting.
- Experience responding to production incidents and performing root-cause analysis.
- The ability to independently lead complex troubleshooting sessions.
- Experience automating operational work using Python, Go or scripting.
- A solid understanding of cloud-native and distributed systems.
- A calm, analytical and structured approach to incidents.
- An ownership mindset and the ability to move from reactive troubleshooting to lasting improvements.
Bonus Points If You Have
- Experience with NVIDIA GPU infrastructure.
- Experience supporting AI, machine-learning or high-performance computing workloads.
- Knowledge of Kubernetes GPU scheduling and resource management.
- Experience with multi-tenant cloud environments.
- Familiarity with Go-based cloud or platform services.
- Experience with Infrastructure as Code and automated platform deployment.
- Knowledge of distributed storage, networking or database troubleshooting.
- Experience defining service-level indicators, objectives and operational reliability targets.
- Experience working in sovereign, regulated or security-sensitive cloud environments.
Interested in opportunities specifically in Netherlands? Discover our dedicated Visa Sponsorship Jobs in Netherlands page featuring roles from top employers in this location.
Eligibility & Application Information
We welcome non-native Dutch speakers to apply. However, to be eligible, you must:
- Have a valid work permit in the Netherlands. ( wo do offer sponsorship if needeed)
- Reside in the Netherlands and be able to travel to the office in Leiden (near The Hague).
- Be fluent in English. Dutch is not required.
Ready to make Europe’s sovereign AI cloud more reliable and operationally scalable?
Apply now through Frank Poll and help Nebul build a cloud platform that engineering teams and customers can depend on.
Similar Jobs
Explore other opportunities that match your interests