Back to Job Listings

Principal Scientist – Data Pipeline Engineer

San Francisco

SpringCube

Full-time - Principal Engineer

Software, SaaS, Cloud & Infrastructure

Posted 3 weeks ago

$200,000 - $250,000

Contact Employer
  • Share:
Send Feedback
Report This Job

Job Description

The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.

Company Overview

A leading global technology organization is advancing multimodal generative AI through large-scale foundation models spanning image, video, and audio. Its engineering teams are building sophisticated data infrastructure capable of processing billions of assets and transforming raw multimodal data into high-quality, training-ready datasets. The organization focuses on combining advanced machine learning with highly scalable distributed systems to accelerate model development and improve training outcomes.

The organization is seeking a Principal ML Engineer to architect and scale the multimodal data processing pipelines and infrastructure supporting multimodal foundation models. This role sits at the intersection of data engineering and applied machine learning, with responsibility for building distributed, GPU-accelerated systems that transform billions of raw assets into training-ready data at scale.

The work will directly influence the speed, efficiency, reliability, and quality of model training by improving data pipeline throughput and ensuring high-quality data reaches training systems. This is a senior individual contributor position with broad technical influence across data engineering, infrastructure, and machine learning teams.

Key Responsibilities

Optimize Data Processing Pipelines at Scale

  • Architect and optimize large-scale distributed pipelines that process billions of image, video, and audio assets through machine learning workflows into training-ready data.
  • Increase inference throughput across pipelines through batching, parallelism, and improved hardware utilization to process raw data faster and more cost-effectively.
  • Identify and eliminate bottlenecks throughout ingestion, processing, and delivery, including storage, I/O, and compute scheduling.

Architect Scalable Data Infrastructure

  • Design systems capable of reliably storing, indexing, and serving billions of data points requiring substantial processing.
  • Work with large-scale databases, distributed storage systems, and high-throughput computing infrastructure.
  • Apply deep expertise in distributed systems and frameworks such as Ray, Spark, or equivalent technologies to orchestrate large-scale GPU- and CPU-intensive workloads.
  • Own key architectural decisions involving databases, storage, job scheduling, and GPU cluster utilization to support continued platform and model growth.

Drive Data Curation for Model Training

  • Apply strong machine learning expertise, particularly in inference optimization for VLMs and LLMs and data curation for model training.
  • Partner closely with modeling teams to understand which datasets and data characteristics improve training outcomes.
  • Translate modeling requirements into scalable data pipeline and curation strategies.
  • Operate as a hands-on technical leader bridging data engineering and applied machine learning.

Required Qualifications

  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems, including experience operating at very large scale involving billions of records or assets.
  • Strong software engineering background with hands-on expertise in distributed systems and frameworks such as Ray, Spark, or equivalent large-scale data processing technologies.
  • Proficiency in Python, along with strong experience in a systems-level programming language such as C++, Rust, Go, or Java.
  • Strong debugging skills across distributed systems and machine learning-focused runtime environments.
  • Deep knowledge of databases and storage systems at scale, including data lakes, indexing, and retrieval across billions of data points.
  • Strong machine learning background, particularly in optimizing GPU inference pipelines for VLMs, LLMs, or other large models.
  • Experience with batching, quantization, model serving, and throughput-versus-latency optimization.
  • Experience with data curation for model training, including understanding what makes data valuable for generative and multimodal models.
  • Ability to work across the full technology stack, from low-level systems and GPU optimization to higher-level data strategy and curation.
  • Strong communication and collaboration skills with the ability to work effectively across data, infrastructure, and modeling teams.
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field.

Disclaimer

SpringCube curates tech job listings from various company websites to support tech professionals globally.

  1. No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
  2. No Client Relationship: This company is not a client of SpringCube unless stated.
  3. To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
  4. No Liability: SpringCube is not liable for inaccuracies.
‹