Loading jobs…
Loading jobs…
Coupang — Mountain View, California
Company Introduction We exist to wow our customers. ” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.
We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations.
At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real.
We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Role Overview The Resource Fabric Engineering team builds and operates the foundational infrastructure and developer tools that power the AI lifecycle across Coupang.
Our mission is to provide scalable, reliable, and intelligent platforms that enable teams to efficiently build, deploy, and manage AI workloads. We develop core services including resource management, workflow orchestration, model lifecycle tooling, and observability solutions that support machine learning innovation at scale. We are looking for a Staff Software Engineer to lead the design and evolution of large-scale distributed systems and drive technical excellence across the platform.
What You Will Do
Design and implement the development of a next-generation Resource Manager and AI developer tools, enabling efficient, secure, and policy-compliant data access across hybrid environments. Lead the design of scalable, event-driven microservices for optimized resource management. Build and enhance developer tools that improve workflow orchestration, model management, observability, and operational efficiency.
Mentor junior engineers and contribute to architectural decisions that shape the future of AI infrastructure. Collaborate with Product owner, TPMs, and Compliance teams to define end-to-end ML model development lifecycle experiences. Drive operational excellence by improving system reliability, performance, monitoring, and incident management practices.
Take end-to-end ownership of critical platform initiatives from design through operation and continuous improvement.
Qualifications
Bachelor’s degree in Computer Science, Engineering, or a related technical field. 8 years of professional software development experience, or 6 years with an advanced degree. Experience designing, building, and operating large-scale distributed systems in cloud environments.
Proficiency in at least one modern programming language such as Java, Go, or Python. Experience with AWS, Azure, or GCP, and cloud-native development practices.
Preferred Qualifications
Strong expertise in distributed systems technologies including Kafka, Cassandra, ClickHouse, MongoDB, or similar platforms. Experience with Kubernetes, Docker, gRPC, Spring, and microservices architecture. Experience building infrastructure, tooling, or platforms supporting machine learning and AI workloads.
, TensorFlow, PyTorch, ONNX). , Prometheus, Grafana, ELK stack). Deep expertise in Java Spring Boot, Hibernate, and RESTful API design.
Strong understanding data governance, compliance, and access control in distributed systems.