All jobs

uvation

Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

Serbia · Remote

About this role

Job Overview We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments. This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms . The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical. Key Responsibilities & Required Skills Linux & Bare Metal Infrastructure Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred) Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments Strong understanding of server hardware, including: BIOS/UEFI RAID controllers Firmware management iLO/iDRAC/IPMI NICs and SmartNICs HBA cards Hardware diagnostics and troubleshooting Experience designing, implementing, and supporting enterprise Linux infrastructure at scale AI Factory & GPU Infrastructure Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads Understanding of NVIDIA GPU technologies including: A100, H100, H200, B200, or equivalent GPU platforms NVIDIA DGX and OEM GPU servers GPU provisioning and lifecycle management GPU monitoring and performance optimization Knowledge of AI Factory architecture and infrastructure requirements Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads Understanding of: GPU resource allocation and scheduling Multi-GPU systems GPU networking requirements High-bandwidth, low-latency infrastructure design Familiarity with NVIDIA ecosystem technologies such as: CUDA NCCL GPUDirect Storage NVIDIA Fabric Manager NVIDIA Base Command (preferred) Enterprise Storage & Data Platforms Advanced Linux storage administration: LVM XFS, EXT4 NFS iSCSI Fibre Channel SAN Multipath I/O Strong hands-on experience with Ceph , including: Cluster architecture MON, OSD, MDS RBD, CephFS, RGW Capacity planning Performance tuning Failure recovery Experience with high-performance AI storage platforms such as: WEKA VAST Data Dell PowerScale Pure Storage FlashBlade NetApp Understanding of: NVMe-over-Fabrics (NVMe-oF) RDMA GPUDirect Storage Parallel file systems AI data pipelines Networking & Infrastructure Strong networking knowledge: Bonding VLANs Routing MTU optimization DNS DHCP Experience with high-performance data center networking: 100G/200G/400G Ethernet RoCE RDMA Spine-Leaf architectures Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting Operations & Reliability Experience with high availability, clustering, and disaster recovery Strong troubleshooting skills across: Linux operating systems Hardware platforms GPU infrastructure Networking Enterprise storage Experience supporting mission-critical production environments Bash and Python scripting for automation and operational efficiency Experience creating operational documentation, runbooks, and infrastructure standards Nice to Have Kubernetes infrastructure (especially AI/ML and GPU integration)

Skills and categories

Linux-Infrastructure-EngineeringBare-Metal-EngineeringStorage-EngineeringAI-Infrastructure-EngineeringGPU-Infrastructure-EngineeringLinux-Infrastructure-EngineerLinux-Infrastructure-SpecialistInfrastructure-Linux-Unix-AnalystSenior-AI-Infrastructure-EngineerAI-Infrastructure-EngineerAI-ML-Infrastructure-EngineerMachine-Learning-Infrastructure-EngineerInfrastructure-Platform-EngineerInfrastructure-EngineerDeveloper

Listing provided by Himalayas. Verify availability on the source board before applying — Kartavyam aggregates public listings as they appear and does not manage this application.

Similar roles

Browse all jobs