Job Description
The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.
Company Overview
A leading global technology and logistics platform is building advanced Generative AI infrastructure that enables teams across its organization to safely bring AI-powered products, agents, automation, and personalization into production. The organization’s GenAI Platform team develops shared machine learning infrastructure, including open-weight large language models and vision-language models, real-time GPU serving, high-throughput batch inference, fine-tuning systems, gateways, evaluation infrastructure, guardrails, and cost attribution capabilities.
Key Responsibilities
- Lead the design of infrastructure that enables teams to move Generative AI concepts from prototype to production.
- Own and evolve open-weight model serving infrastructure, including real-time GPU endpoints, high-throughput batch inference, and fine-tuning using SFT, DPO, and LoRA.
- Develop and architect scalable systems for model serving, batch inference, GPU autoscaling, GPU utilization, and fine-tuning.
- Optimize GPU inference costs and latency while maintaining reliability, scalability, and performance.
- Build infrastructure that enables product teams to select between open-weight and closed-source models with appropriate reliability, fallback, observability, and cost controls.
- Develop platforms that support rapid experimentation while maintaining production standards for latency, monitoring, SLOs, operational readiness, and reliability.
- Collaborate with ML engineers, product engineers, data scientists, and platform teams to transform emerging Generative AI capabilities into reusable platform infrastructure.
- Provide technical leadership for the organization’s centralized Generative AI platform.
- Explore emerging technologies and approaches including reinforcement learning, agent optimization, post-training techniques, and agentic systems.
- Mentor engineers and establish strong technical practices across the team.
- Design systems capable of supporting AI-powered products, agents, automation, and personalization at scale.
Required Qualifications
- Bachelor’s, Master’s, or PhD in Computer Science or an equivalent technical discipline.
- 6+ years of professional software engineering experience.
- Strong backend engineering fundamentals, particularly in Python and distributed systems.
- Proven experience designing and owning production services, APIs, data pipelines, or machine learning infrastructure at scale.
- Experience operating production systems, including observability, debugging, reliability, incident response, and performance and cost optimization.
- Hands-on experience with LLM inference and/or fine-tuning open-weight models in production.
- Experience with model serving, including latency, throughput, batching, autoscaling, and GPU utilization, and/or fine-tuning techniques such as SFT, DPO, and LoRA.
- Demonstrated technical leadership in ambiguous and rapidly evolving technical environments.
- Experience mentoring engineers and transforming customer requirements into reusable platform capabilities.
- Proficiency using AI coding tools such as Claude Code, Codex, or Cursor throughout the software development lifecycle, including software design, code generation, testing, monitoring, and release.
Preferred Qualifications
- Experience with LLM inference engines and serving frameworks such as vLLM, SGLang, or TensorRT-LLM in production.
- Experience with distributed or multi-node fine-tuning and training pipelines, including SFT, DPO, RLHF, LoRA, data preparation, and evaluation.
- Experience optimizing GPU performance, including distributed inference, KV-cache and memory optimization, quantization, and throughput tuning.
- Knowledge of quantization technologies such as FP8, INT8, AWQ, and GPTQ.
- Experience with Kubernetes and cloud infrastructure such as AWS or GCP.
- Experience working with GPUs, serverless or elastic GPU platforms, and high-throughput batch processing systems.
- Experience building LLM gateways, model routing systems, vendor abstraction layers, or cost attribution platforms.
- Experience developing internal developer platforms, self-service infrastructure, or machine learning platforms.
- Experience building and deploying AI agents or MCP servers in production.
- Experience with evaluation systems, LLM observability, tracing, Retrieval-Augmented Generation (RAG), search systems, or vector databases.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
- No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
- No Client Relationship: This company is not a client of SpringCube unless stated.
- To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
- No Liability: SpringCube is not liable for inaccuracies.