Back to Job Listings

Senior Machine Learning Engineer, Services/MLOps

San Francisco

SpringCube

Full-time - Senior Engineer

IT Cloud Computing, Software & SaaS

Posted 4 weeks ago

Disclosed upon interview

Contact Employer
  • Share:
Send Feedback
Report This Job

Job Description

The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.

Company Overview

A leading global technology organization is expanding its enterprise managed-service capabilities for custom multimedia Generative AI. Its platform enables customers to create and deploy deep-tuned image, video, and 3D models using proprietary customer intellectual property, combined with creative production workflows and media-intelligence capabilities across a range of digital surfaces.

Key Responsibilities

  • Own the complete serving lifecycle for heterogeneous machine learning pipelines, including packaging, versioned releases, canary deployments, rollbacks, and autoscaling.
  • Deploy model pipelines as production services and scale them to enterprise traffic while meeting latency and throughput targets.
  • Ensure production model quality remains consistent with training and reference environments by addressing differences in precision, preprocessing, and model versions.
  • Design enterprise-ready systems with strong tenancy boundaries, data isolation, and controls for protecting customer intellectual property.
  • Build and maintain platform capabilities for rapid pipeline deployment, observability, monitoring, and alerting.
  • Define and enforce quality gates within deployment pipelines through automated evaluations, regression detection, and model-drift monitoring.
  • Manage GPU capacity and infrastructure costs through utilization optimization, efficient batching, and appropriate sizing of accelerator fleets.
  • Participate in production ML operations, including on-call responsibilities, incident response, and postmortem activities for availability and latency issues.
  • Build externalizable data pipelines that support self-service fine-tuning workflows for enterprise customers where applicable.
  • Develop optimized VLM deployments for media intelligence and content-querying applications where applicable.

Cross-Functional Collaboration

  • Partner with Applied Science teams to transition research models into reliable, high-throughput production services while maintaining model quality.
  • Collaborate with ML Engineering leadership and AI Platform teams on shared infrastructure, accelerator capacity, and scalable model-serving capabilities.
  • Work with creative production and product teams to translate workflows into dependable and performant machine learning services.

Required Qualifications

  • 5+ years of experience in machine learning engineering, with significant ownership of production ML or inference services at scale.
  • Strong Python and deep-learning engineering skills, including hands-on experience with PyTorch.
  • Experience deploying and scaling model-backed services in production environments.
  • Experience composing multi-model pipelines and serving them through APIs, including orchestration, batching, autoscaling, and version management.
  • Proven experience building observability, monitoring, and alerting systems for production services.
  • Experience working with multiple Generative AI architectures, including LLMs, VLMs, diffusion models, transformer models, and 3D or mesh technologies.
  • Experience integrating, optimizing, and evaluating machine learning models in collaboration with research or applied science teams.
  • Experience designing multi-tenant systems and implementing data isolation within enterprise or regulated environments.
  • Proficiency with containers and orchestration technologies such as Docker and Kubernetes.
  • Experience with CI/CD practices for machine learning systems.
  • Experience working with at least one major cloud platform, such as AWS or Azure.
  • Experience optimizing GPU inference for latency and cost through techniques such as quantization, batching, and specialized serving runtimes.
  • Strong data-driven problem-solving skills and excellent communication abilities for cross-functional collaboration.

Preferred Qualifications

  • Experience with custom CUDA development.
  • Experience operating large-scale Generative AI inference systems.
  • Experience with enterprise media intelligence and multimedia AI applications.
  • Experience building self-service fine-tuning platforms for enterprise customers.
  • Experience optimizing VLM deployments for media intelligence and content querying.

Education

  • Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent practical experience building and operating production machine learning systems.

Disclaimer

SpringCube curates tech job listings from various company websites to support tech professionals globally.

  1. No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
  2. No Client Relationship: This company is not a client of SpringCube unless stated.
  3. To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
  4. No Liability: SpringCube is not liable for inaccuracies.
‹