← Back to roles
Platform
HPC Engineer
Design and operate Arlequin's cloud-native, sovereign HPC platform powering our AI products and research
5+ years, Senior
Full-time
Paris (hybrid / full-remote / on-site flexible)
About Arlequin
Arlequin AI is both a research lab in topological deep learning and an AI platform. Our first product, HuDex, converts massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes instead of days, providing an operational advantage for governments and businesses.
~30 people, post-seed, Series A in progress. On-site in Paris.
The role
We are building a cloud-native, sovereign, and end-to-end automated high-performance computing platform to support the scientific computing workloads of our products and research. Our products are also deployed on-premise, in air-gapped environments, for clients with the strongest security requirements.
The Platform team owns the technical foundations Arlequin AI runs on: DevOps, MLOps, FinOps, FieldOps, Security/Compliance, compute and IT. Our role is not to build infrastructure for its own sake — we build self-service tools so engineers can ship without waiting on us, Forward Deployed Engineers can deploy clients without our help, C-levels understand the real cost of what they sell, and researchers don’t have to deal with engineering questions to run their experiments. Today we are a team of 3, and we aim to grow the team to 8 people by December.
As an HPC Engineer, you join the Platform team to design and operate our compute platform covering all of Arlequin’s compute workloads, ensuring a performant, scalable, and reliable platform — 100% cloud-native and orchestrated by Kubernetes. Concretely, you operate GPU and CPU clusters on Kubernetes and optimize end-to-end performance (GPU, high-throughput network interconnect, high-performance storage), while building the tools and abstractions that make teams autonomous in using, launching, and monitoring their workloads.
Your responsibilities
Infrastructure: design, deploy, and maintain infrastructure on Scaleway; industrialize IaC with OpenTofu and Terragrunt
Data storage: design and operate high-performance storage (parallel/distributed file systems, data caches, object storage) and optimize end-to-end I/O for training and inference
Compute platform: design, deploy, and operate a Kubernetes compute cluster sized for heterogeneous workloads; set up scheduling (queues, priorities, gang scheduling, fair-sharing); optimize GPU sharing and utilization; deploy and maintain device plugins
Workloads: identify and eliminate bottlenecks across compute, network, and I/O; design and operate model serving (real-time, streaming, batch, progressive rollouts); operate large-scale batch pipelines and training pipelines with the research and MLOps teams
Instrumentation & optimization: monitor compute and GPUs; instrument inference services (latency, throughput, error rate, SLOs); optimize compute costs with unit metrics; configure alerting; improve utilization and right-sizing
Platform engineering: collaborate with research/product on the platform roadmap; support capacity planning; build self-service abstractions; document the platform and runbooks
On-call: no rotation today, incidents are handled during business hours. When a rotation becomes necessary, it will never exceed one week on-call out of five, and will be compensated
Stack
Cloud & orchestration: Scaleway, Kubernetes
GPU: NVIDIA GPU Operator, CUDA, DCGM
Serving & batch: SGLang, Kueue
Autoscaling: HPA / VPA, KEDA, Cluster Autoscaler
Distributed computing: NCCL, MPI, RDMA / RoCE
IaC & GitOps: OpenTofu, Terragrunt, ArgoCD
Storage: Blob Storage
Observability: Prometheus, DCGM Exporter, Grafana, Loki, Tempo, Mimir (LGTM), Alertmanager
Languages: Python, Bash
What we’re looking for
5+ years of experience in infrastructure / HPC / compute platforms, including production experience
Advanced proficiency with Kubernetes in a production environment
Hands-on experience with GPU workloads and batch schedulers: scheduling, resource management, and sharing
Experience operating a production service under latency and availability constraints: autoscaling, load management, SLOs, incident management
Strong IaC experience (Terraform / OpenTofu, ideally Terragrunt)
Good understanding of distributed computing (multi-node / multi-GPU) and large-scale batch processing
Autonomy and a strong sense of ownership over the production scope, excellent communication, a feedback culture, technical curiosity, and pragmatism
Bonus
Experience with Scaleway
In-depth knowledge of the NVIDIA / ROCm / TPU ecosystems
Hands-on experience with inference optimization: quantization, continuous batching, KV-cache / prefix caching, speculative decoding, prefill/decode disaggregation
RDMA / InfiniBand / RoCE and low-latency networking
Parallel file systems: Lustre, BeeGFS, GPFS
FinOps awareness, high-performance networking and storage knowledge, background in scientific computing or research, experience optimizing compute code
Process
First-fit interview (30 min)
Technical interview with the Platform team (coding + design)
Meeting with the hiring manager and the founders
Offer
Package
Competitive compensation including equity stake
Flexible remote policy: hybrid / full-remote / full on-site
Training, conference, and certification budget, with time dedicated to CNCF/LF open source contributions
Interested in this role?
Send us your profile and a few lines about what you would like to build with us.
Apply