← Back to roles

Platform

HPC Engineer

Design and operate Arlequin's cloud-native, sovereign HPC platform powering our AI products and research

5+ years, Senior

Full-time

Paris (hybrid / full-remote / on-site flexible)

About Arlequin

Arlequin AI is both a research lab in topological deep learning and an AI platform. Our first product, HuDex, converts massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes instead of days, providing an operational advantage for governments and businesses.
~30 people, post-seed, Series A in progress. On-site in Paris.

The role

We are building a cloud-native, sovereign, and end-to-end automated high-performance computing platform to support the scientific computing workloads of our products and research. Our products are also deployed on-premise, in air-gapped environments, for clients with the strongest security requirements.

The Platform team owns the technical foundations Arlequin AI runs on: DevOps, MLOps, FinOps, FieldOps, Security/Compliance, compute and IT. Our role is not to build infrastructure for its own sake — we build self-service tools so engineers can ship without waiting on us, Forward Deployed Engineers can deploy clients without our help, C-levels understand the real cost of what they sell, and researchers don’t have to deal with engineering questions to run their experiments. Today we are a team of 3, and we aim to grow the team to 8 people by December.

As an HPC Engineer, you join the Platform team to design and operate our compute platform covering all of Arlequin’s compute workloads, ensuring a performant, scalable, and reliable platform — 100% cloud-native and orchestrated by Kubernetes. Concretely, you operate GPU and CPU clusters on Kubernetes and optimize end-to-end performance (GPU, high-throughput network interconnect, high-performance storage), while building the tools and abstractions that make teams autonomous in using, launching, and monitoring their workloads.

Your responsibilities

  • Infrastructure: design, deploy, and maintain infrastructure on Scaleway; industrialize IaC with OpenTofu and Terragrunt

  • Data storage: design and operate high-performance storage (parallel/distributed file systems, data caches, object storage) and optimize end-to-end I/O for training and inference

  • Compute platform: design, deploy, and operate a Kubernetes compute cluster sized for heterogeneous workloads; set up scheduling (queues, priorities, gang scheduling, fair-sharing); optimize GPU sharing and utilization; deploy and maintain device plugins

  • Workloads: identify and eliminate bottlenecks across compute, network, and I/O; design and operate model serving (real-time, streaming, batch, progressive rollouts); operate large-scale batch pipelines and training pipelines with the research and MLOps teams

  • Instrumentation & optimization: monitor compute and GPUs; instrument inference services (latency, throughput, error rate, SLOs); optimize compute costs with unit metrics; configure alerting; improve utilization and right-sizing

  • Platform engineering: collaborate with research/product on the platform roadmap; support capacity planning; build self-service abstractions; document the platform and runbooks

  • On-call: no rotation today, incidents are handled during business hours. When a rotation becomes necessary, it will never exceed one week on-call out of five, and will be compensated

Stack

  • Cloud & orchestration: Scaleway, Kubernetes

  • GPU: NVIDIA GPU Operator, CUDA, DCGM

  • Serving & batch: SGLang, Kueue

  • Autoscaling: HPA / VPA, KEDA, Cluster Autoscaler

  • Distributed computing: NCCL, MPI, RDMA / RoCE

  • IaC & GitOps: OpenTofu, Terragrunt, ArgoCD

  • Storage: Blob Storage

  • Observability: Prometheus, DCGM Exporter, Grafana, Loki, Tempo, Mimir (LGTM), Alertmanager

  • Languages: Python, Bash

What we’re looking for

  • 5+ years of experience in infrastructure / HPC / compute platforms, including production experience

  • Advanced proficiency with Kubernetes in a production environment

  • Hands-on experience with GPU workloads and batch schedulers: scheduling, resource management, and sharing

  • Experience operating a production service under latency and availability constraints: autoscaling, load management, SLOs, incident management

  • Strong IaC experience (Terraform / OpenTofu, ideally Terragrunt)

  • Good understanding of distributed computing (multi-node / multi-GPU) and large-scale batch processing

  • Autonomy and a strong sense of ownership over the production scope, excellent communication, a feedback culture, technical curiosity, and pragmatism

Bonus

  • Experience with Scaleway

  • In-depth knowledge of the NVIDIA / ROCm / TPU ecosystems

  • Hands-on experience with inference optimization: quantization, continuous batching, KV-cache / prefix caching, speculative decoding, prefill/decode disaggregation

  • RDMA / InfiniBand / RoCE and low-latency networking

  • Parallel file systems: Lustre, BeeGFS, GPFS

  • FinOps awareness, high-performance networking and storage knowledge, background in scientific computing or research, experience optimizing compute code

Process

  1. First-fit interview (30 min)

  2. Technical interview with the Platform team (coding + design)

  3. Meeting with the hiring manager and the founders

  4. Offer

Package

  • Competitive compensation including equity stake

  • Flexible remote policy: hybrid / full-remote / full on-site

  • Training, conference, and certification budget, with time dedicated to CNCF/LF open source contributions

Interested in this role?

Send us your profile and a few lines about what you would like to build with us.

Apply