Back to Job Listings

Principal Scientist – Data Pipeline Engineer

San Francisco

SpringCube

Full-time - Principal Engineer

IT Cloud Computing, Software & SaaS

Posted 3 weeks ago

Disclosed upon interview

Contact Employer
  • Share:
Send Feedback
Report This Job

Job Description

The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.

Company Overview

A leading global technology organization is developing advanced generative and multimodal AI technologies across image, video, and audio applications. Its AI teams build large-scale foundation models and the data infrastructure required to train and improve these models efficiently.

The organization is seeking a Principal Scientist – Data Pipeline Engineer to architect and scale multimodal data processing pipelines and infrastructure supporting foundation models. This senior individual contributor role sits at the intersection of data engineering and applied machine learning, with broad technical influence across data, infrastructure, and modeling teams.

The successful candidate will design distributed, GPU-accelerated systems capable of transforming billions of raw assets into high-quality, training-ready data. The role will directly influence pipeline throughput, reliability, cost efficiency, and the quality of data delivered to model training.

Key Responsibilities

Optimize Data Processing Pipelines at Scale

  • Architect and optimize large-scale distributed pipelines that process billions of images, video, and audio assets through machine learning workflows into training-ready data.
  • Scale inference throughput across pipelines through batching, parallelism, and improved hardware utilization.
  • Identify and eliminate bottlenecks across ingestion, processing, and delivery, including storage, I/O, and compute scheduling.
  • Improve the speed, reliability, and cost efficiency of large-scale data processing workflows.

Architect Scalable Data Infrastructure

  • Design systems that reliably store, index, and serve billions of data points requiring substantial processing.
  • Develop infrastructure using large-scale databases, distributed storage, and high-throughput computing systems.
  • Apply deep expertise in distributed systems and frameworks such as Ray or equivalent technologies to orchestrate large-scale GPU- and CPU-intensive workloads.
  • Own architecture decisions involving databases, storage systems, job scheduling, and GPU cluster utilization.
  • Develop infrastructure capable of scaling alongside increasing data volumes and model complexity.

Drive Data Curation for Model Training

  • Apply strong machine learning expertise, particularly in inference optimization for VLMs and LLMs and data curation for model training.
  • Partner closely with modeling teams to understand which data characteristics improve training outcomes.
  • Translate model training requirements into effective data pipeline and curation strategies.
  • Serve as a hands-on technical leader connecting data engineering and applied machine learning.

Required Qualifications

  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems, including experience working at large scale with billions of records or assets.
  • Strong software engineering background with hands-on expertise in distributed systems.
  • Experience with Ray, Spark, or equivalent large-scale data processing frameworks.
  • Proficiency in Python and strong experience with a systems-level programming language such as C++, Rust, Go, or Java.
  • Strong debugging skills across distributed and machine learning-focused runtime environments.
  • Deep knowledge of databases and storage systems at scale, including data lakes, indexing, and retrieval across billions of data points.
  • Strong machine learning background, particularly experience optimizing GPU inference pipelines for VLMs, LLMs, or other large models.
  • Experience with batching, quantization, model serving, and throughput-versus-latency optimization.
  • Experience with data curation for model training and an understanding of what makes data valuable for generative and multimodal models.
  • Ability to work across the full technology stack, from low-level systems and GPU optimization to higher-level data strategy and curation.
  • Strong communication skills and the ability to collaborate effectively across data, infrastructure, and modeling teams.
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field.

Disclaimer

SpringCube curates tech job listings from various company websites to support tech professionals globally.

  1. No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
  2. No Client Relationship: This company is not a client of SpringCube unless stated.
  3. To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
  4. No Liability: SpringCube is not liable for inaccuracies.
‹