← Back to roles
Platform
Site Reliability Engineer
Ensure the availability, performance, and scalability of Arlequin's sovereign, cloud-native platform
5+ years, Senior
Full-time
Paris (hybrid / full-remote / on-site flexible)
About Arlequin
Arlequin AI is both a topological deep learning research lab and an AI platform. Our first product, HuDex, converts massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes instead of days, providing an operational advantage for government agencies and businesses.
~30 people, post-seed, Series A in progress. On-site in Paris.
The role
We have chosen a cloud-native, sovereign, end-to-end automated infrastructure, hosted on Scaleway. Joining us means taking part in building and operating a reliable, secure, and observable platform serving our research, product, and development teams.
The Platform team owns the technical foundations Arlequin AI runs on: DevOps, DevSecOps, MLOps, DataOps, FinOps, security and compliance, compute, and IT. Our role is not to build infrastructure for its own sake — we build self-service tools so engineers can ship without waiting on us, a Forward Deployed Engineer can deploy a client without our help, C-levels understand the real cost of what they sell, and researchers don’t have to deal with engineering questions to run their experiments. Today we are a team of 3, and we are looking for people to help structure the team.
As a Senior SRE, you join the Platform team to ensure the availability, performance, and scalability of our platforms, while giving development teams the tooling to be autonomous. You work on both the run and the build side, with a genuine culture of toil reduction.
Your responsibilities
Reliability & production: define, instrument, and track SLIs / SLOs / SLAs; continuously improve resilience (capacity planning, load testing, chaos engineering, disaster recovery, eliminating SPOFs); reduce toil through automation and self-service; handle incidents and run blameless post-mortems; support the scaling of training, inference, and scientific computing workloads
Infrastructure & IaC: design, deploy, and maintain infrastructure on Scaleway; industrialize IaC with OpenTofu and Terragrunt; administer and evolve production Kubernetes clusters
GitOps & delivery: operate and evolve continuous deployment following GitOps principles with ArgoCD; standardize CI/CD pipelines and release workflows; manage configuration and secrets securely
Observability: improve and maintain our LGTM stack (Loki, Grafana, Tempo, Mimir); design dashboards and a relevant alerting policy with Alertmanager; promote OpenTelemetry observability and train engineering teams on it
Culture & collaboration: spread SRE / DevOps best practices across teams; document the infrastructure and runbooks; contribute to architecture decisions and the platform’s technical roadmap
On-call: no rotation today, incidents are handled during business hours. When a rotation becomes necessary, it will never exceed one week on-call out of five, it will be paid, and you’ll take part in designing it
Stack
Cloud & orchestration: Scaleway, Kubernetes
IaC & GitOps: OpenTofu, Terragrunt, ArgoCD
Networking: Cilium
Observability: Loki, Grafana, Tempo, Mimir (LGTM), Prometheus, Alertmanager
Languages: Bash, Python
What we’re looking for
5+ years of experience in SRE / DevOps / Infrastructure, including significant experience with mission-critical production
Advanced command of Kubernetes in production environments
Solid experience with IaC (Terraform / OpenTofu, ideally Terragrunt)
Hands-on GitOps practice (ArgoCD or Flux)
Strong observability culture (ideally the LGTM stack)
Comfortable with scripting / automation (Bash, Python) and solid Linux fundamentals
Autonomy and a strong sense of ownership over the production scope, excellent communication, a feedback culture, technical curiosity, and pragmatism
Bonus
Hands-on experience with Scaleway
Knowledge of Cilium & service mesh
FinOps awareness
Process
First-fit interview (30 min)
Technical interview with the Platform team (coding + design)
Meeting with the hiring manager
Offer
Package
Competitive compensation including equity stake
Flexible remote policy: hybrid / full-remote / full on-site
Training, conference, and certification budget, with time dedicated to CNCF/LF open source contributions
Interested in this role?
Send us your profile and a few lines about what you would like to build with us.
Apply