Loading jobs…
Loading jobs…
Uber — Europe Mont-Blanc, Rhône-Alpes
Location work modality: EMEA (remote) Start: ASAP Type of Contract: Permanent, full-time About Radian Arc Radian Arc, now part of InferX, Submer's AI cloud and GPU infrastructure platform, provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments. What impact you will have Mission: Design and build the observability platform that powers visibility, reliability, and performance insights for large-scale GPU cloud infrastructure as well as smaller edge deployments.
This role is responsible for designing and implementing key parts of the observability architecture across the platform, enabling engineering, operations, and customers to understand system behavior in real time across distributed AI workloads, GPU clusters, networking fabrics, storage systems, and edge inference environments. You will design and operate low-latency, high-scale telemetry pipelines that collect, process, and analyze metrics, logs, and traces from infrastructure running across core datacenter clusters and smaller edge deployments. The platform you build will support internal operations, automated reliability mechanisms, and customer-facing observability experiences.
As a senior engineer, you will lead delivery of major observability initiatives, contribute to the evolution of telemetry standards and SLO implementation, and work with other teams to ensure observability is effectively integrated into the platform architecture from infrastructure to application layers. You will collaborate closely with infrastructure, networking, storage, and platform engineering teams to provide clear visibility into performance bottlenecks, infrastructure degradation, and distributed workload behavior across both hyperscale GPU environments and smaller edge installations. This role contributes directly to improving platform reliability by analyzing production telemetry, identifying systemic issues, and driving improvements in performance, efficiency, and operational stability across the stack.
What You’Ll Do
Observability Platform Architecture Design and implement scalable telemetry pipelines for metrics, logs, and traces across distributed GPU infrastructure. Architect observability systems capable of ingesting high-cardinality telemetry from thousands of nodes and services. Build and operate telemetry storage systems optimized for large-scale time-series and event data.
Contribute to observability standards across services, including metrics, tracing instrumentation, logging, and SLO implementation. Infrastructure and Platform Observability Build visibility across compute, storage, and networking layers of the platform. Instrument GPU clusters, inference workloads, and distributed training environments.
Detect infrastructure degradation such as: GPU throttling, Network congestion, Storage latency, Hardware degradation. Implement telemetry pipelines for GPU, CPU, network, and storage performance metrics. Customer-Facing Observability Build dashboards and monitoring tools that expose system health and performance to both internal teams and customers.
Provide insights into workload performance including: GPU utilization, Storage throughput, Network latency, Distributed inference performance. Develop performance analysis tools that help customers understand system bottlenecks. Network and Infrastructure Telemetry Develop and maintain network observability platforms.
Build telemetry collectors and exporters using Python or Go. Ingest telemetry from infrastructure components including: NVIDIA Cumulus Linux, VyOS routers, Citrix NetScaler / WAF.