All open positions
Senior Site Reliability Engineer, GPU Fleet
About Sesterce
Sesterce builds and runs AI campuses in Europe: dedicated supercomputers on our own sites, powered by low-carbon energy. Every Sesterce Campus uses the same factory-built power and cooling modules and is run by Sesterce OS, from the grid to the chip. We are a small team moving fast, on campuses in France and Finland and on clusters in Paris, Madrid and Munich.
Keep our GPU clusters healthy and fast.
You will
- build monitoring, alerting and automated repair for the fleet
- lead incident response and post-mortems
- automate routine operations on Sesterce OS.
You bring
- 5+ years as an SRE or production engineer on large fleets
- strong Linux, Kubernetes and Python or Go
- experience with GPU or HPC systems.