Back to Job Listings

Machine Learning Infrastructure – Generative AI

San Francisco

SpringCube

Full-time - Senior Engineer

Social Networking & Media

Posted 3 weeks ago

Disclosed upon interview

Contact Employer
  • Share:
Send Feedback
Report This Job

Job Description

The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.

Company Overview

A leading global technology and logistics platform is building advanced Generative AI infrastructure that enables teams across its organization to safely bring AI-powered products, agents, automation, and personalization into production. The organization’s GenAI Platform team develops shared machine learning infrastructure, including open-weight large language models and vision-language models, real-time GPU serving, high-throughput batch inference, fine-tuning systems, gateways, evaluation infrastructure, guardrails, and cost attribution capabilities.

Key Responsibilities

  • Lead the design of infrastructure that enables teams to move Generative AI concepts from prototype to production.
  • Own and evolve open-weight model serving infrastructure, including real-time GPU endpoints, high-throughput batch inference, and fine-tuning using SFT, DPO, and LoRA.
  • Develop and architect scalable systems for model serving, batch inference, GPU autoscaling, GPU utilization, and fine-tuning.
  • Optimize GPU inference costs and latency while maintaining reliability, scalability, and performance.
  • Build infrastructure that enables product teams to select between open-weight and closed-source models with appropriate reliability, fallback, observability, and cost controls.
  • Develop platforms that support rapid experimentation while maintaining production standards for latency, monitoring, SLOs, operational readiness, and reliability.
  • Collaborate with ML engineers, product engineers, data scientists, and platform teams to transform emerging Generative AI capabilities into reusable platform infrastructure.
  • Provide technical leadership for the organization’s centralized Generative AI platform.
  • Explore emerging technologies and approaches including reinforcement learning, agent optimization, post-training techniques, and agentic systems.
  • Mentor engineers and establish strong technical practices across the team.
  • Design systems capable of supporting AI-powered products, agents, automation, and personalization at scale.

Required Qualifications

  • Bachelor’s, Master’s, or PhD in Computer Science or an equivalent technical discipline.
  • 6+ years of professional software engineering experience.
  • Strong backend engineering fundamentals, particularly in Python and distributed systems.
  • Proven experience designing and owning production services, APIs, data pipelines, or machine learning infrastructure at scale.
  • Experience operating production systems, including observability, debugging, reliability, incident response, and performance and cost optimization.
  • Hands-on experience with LLM inference and/or fine-tuning open-weight models in production.
  • Experience with model serving, including latency, throughput, batching, autoscaling, and GPU utilization, and/or fine-tuning techniques such as SFT, DPO, and LoRA.
  • Demonstrated technical leadership in ambiguous and rapidly evolving technical environments.
  • Experience mentoring engineers and transforming customer requirements into reusable platform capabilities.
  • Proficiency using AI coding tools such as Claude Code, Codex, or Cursor throughout the software development lifecycle, including software design, code generation, testing, monitoring, and release.

Preferred Qualifications

  • Experience with LLM inference engines and serving frameworks such as vLLM, SGLang, or TensorRT-LLM in production.
  • Experience with distributed or multi-node fine-tuning and training pipelines, including SFT, DPO, RLHF, LoRA, data preparation, and evaluation.
  • Experience optimizing GPU performance, including distributed inference, KV-cache and memory optimization, quantization, and throughput tuning.
  • Knowledge of quantization technologies such as FP8, INT8, AWQ, and GPTQ.
  • Experience with Kubernetes and cloud infrastructure such as AWS or GCP.
  • Experience working with GPUs, serverless or elastic GPU platforms, and high-throughput batch processing systems.
  • Experience building LLM gateways, model routing systems, vendor abstraction layers, or cost attribution platforms.
  • Experience developing internal developer platforms, self-service infrastructure, or machine learning platforms.
  • Experience building and deploying AI agents or MCP servers in production.
  • Experience with evaluation systems, LLM observability, tracing, Retrieval-Augmented Generation (RAG), search systems, or vector databases.

Disclaimer

SpringCube curates tech job listings from various company websites to support tech professionals globally.

  1. No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
  2. No Client Relationship: This company is not a client of SpringCube unless stated.
  3. To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
  4. No Liability: SpringCube is not liable for inaccuracies.
‹