We are currently looking on behalf of one of our clients based in Basel for an AI Platform System Engineer. This is a permanent on-site position.
Your Responsibilities
Design and evolve scalable AI
infrastructure services optimized for GPU workloads
Implement and manage
enterprise-grade AI platforms (e.g., Kubernetes, OpenShift) for model
inference
Operate and maintain
high-performance computing (HPC) clusters, including bare-metal and
virtual GPU servers
Troubleshoot and resolve
complex infrastructure incidents across hardware and software stacks
Support data science teams with
infrastructure provisioning
Perform critical infrastructure
tasks and participate in on-call rotations
Your Profile
University degree (university /
FH) in Computer Science
5 years of IT system
administration experience, including 3+ years specializing in Linux (RHEL
preferred) and strong scripting skills in Shell, Python, and Ansible
Deep expertise in GPU
topologies and architectures (NVLink, PCIe switching) and container
orchestration (Kubernetes, Terraform, ArgoCD). Initial hands-on experience
in Prometheus, Grafana, and DCGM Exporter is required
Familiarity with ITIL
methodologies, combined with a proactive, accountable work ethic. Strong
analytical and problem-solving skills, effective prioritization, and the
ability to deliver results under pressure
Fluent in English; proficiency
in German is a plus