Job Description
The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.
Company Overview
A leading global technology organization is expanding its enterprise managed-service capabilities for custom multimedia Generative AI. Its platform enables customers to create and deploy deep-tuned image, video, and 3D models using proprietary customer intellectual property, combined with creative production workflows and media-intelligence capabilities across a range of digital surfaces.
Key Responsibilities
- Own the complete serving lifecycle for heterogeneous machine learning pipelines, including packaging, versioned releases, canary deployments, rollbacks, and autoscaling.
- Deploy model pipelines as production services and scale them to enterprise traffic while meeting latency and throughput targets.
- Ensure production model quality remains consistent with training and reference environments by addressing differences in precision, preprocessing, and model versions.
- Design enterprise-ready systems with strong tenancy boundaries, data isolation, and controls for protecting customer intellectual property.
- Build and maintain platform capabilities for rapid pipeline deployment, observability, monitoring, and alerting.
- Define and enforce quality gates within deployment pipelines through automated evaluations, regression detection, and model-drift monitoring.
- Manage GPU capacity and infrastructure costs through utilization optimization, efficient batching, and appropriate sizing of accelerator fleets.
- Participate in production ML operations, including on-call responsibilities, incident response, and postmortem activities for availability and latency issues.
- Build externalizable data pipelines that support self-service fine-tuning workflows for enterprise customers where applicable.
- Develop optimized VLM deployments for media intelligence and content-querying applications where applicable.
Cross-Functional Collaboration
- Partner with Applied Science teams to transition research models into reliable, high-throughput production services while maintaining model quality.
- Collaborate with ML Engineering leadership and AI Platform teams on shared infrastructure, accelerator capacity, and scalable model-serving capabilities.
- Work with creative production and product teams to translate workflows into dependable and performant machine learning services.
Required Qualifications
- 5+ years of experience in machine learning engineering, with significant ownership of production ML or inference services at scale.
- Strong Python and deep-learning engineering skills, including hands-on experience with PyTorch.
- Experience deploying and scaling model-backed services in production environments.
- Experience composing multi-model pipelines and serving them through APIs, including orchestration, batching, autoscaling, and version management.
- Proven experience building observability, monitoring, and alerting systems for production services.
- Experience working with multiple Generative AI architectures, including LLMs, VLMs, diffusion models, transformer models, and 3D or mesh technologies.
- Experience integrating, optimizing, and evaluating machine learning models in collaboration with research or applied science teams.
- Experience designing multi-tenant systems and implementing data isolation within enterprise or regulated environments.
- Proficiency with containers and orchestration technologies such as Docker and Kubernetes.
- Experience with CI/CD practices for machine learning systems.
- Experience working with at least one major cloud platform, such as AWS or Azure.
- Experience optimizing GPU inference for latency and cost through techniques such as quantization, batching, and specialized serving runtimes.
- Strong data-driven problem-solving skills and excellent communication abilities for cross-functional collaboration.
Preferred Qualifications
- Experience with custom CUDA development.
- Experience operating large-scale Generative AI inference systems.
- Experience with enterprise media intelligence and multimedia AI applications.
- Experience building self-service fine-tuning platforms for enterprise customers.
- Experience optimizing VLM deployments for media intelligence and content querying.
Education
- Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent practical experience building and operating production machine learning systems.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
- No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
- No Client Relationship: This company is not a client of SpringCube unless stated.
- To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
- No Liability: SpringCube is not liable for inaccuracies.