Back to Job Listings

Staff ML Data Engineer (Datagrid)

San Francisco

SpringCube

Full-time - Senior Engineer

Social Networking & Media

Posted 3 weeks ago

Disclosed upon interview

Contact Employer
  • Share:
Send Feedback
Report This Job

Job Description

The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.

Company Overview

A leading construction technology and software organization is expanding its AI and Frontier Models capabilities and is seeking an experienced Staff ML Data Engineer to support advanced machine learning research and applied AI products. The organization develops data systems that enable teams to work with large-scale, multimodal datasets and transform research initiatives into reliable production applications.

The role focuses on building robust data infrastructure for frontier-scale machine learning, with particular emphasis on spatial intelligence and multimodal data. The successful candidate will help researchers and engineers discover, curate, transform, and operate on large datasets while ensuring that data systems remain scalable, observable, reproducible, and production-ready.

The Staff ML Data Engineer will work closely with ML researchers, applied ML engineers, and system architects to translate ambiguous research requirements into scalable data systems. The position remains deeply hands-on while providing technical leadership in data architecture, data quality, operational excellence, and engineering standards.

Key Responsibilities

  • Act as the technical lead for data engineering initiatives supporting frontier model research and applied machine learning systems.
  • Design, build, and maintain scalable batch and streaming pipelines for multimodal data, including documents, images, and spatial metadata.
  • Partner with researchers and architects to translate experimental workflows into reliable and repeatable data systems.
  • Lead the development of dataset curation, versioning, and lineage workflows that support rapid experimentation and reproducibility.
  • Establish and maintain standards for data quality, validation, observability, and cost efficiency across AI data pipelines.
  • Contribute to data architecture decisions spanning research environments and production systems.
  • Identify gaps and inefficiencies within existing data workflows and conduct proofs of concept to evaluate potential improvements.
  • Mentor other engineers through code reviews, architecture discussions, and hands-on collaboration.
  • Provide technical guidance on data infrastructure supporting machine learning training, evaluation, and inference workflows.
  • Help establish engineering practices that improve the reliability, scalability, and operational efficiency of AI data systems.

Required Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 8+ years of experience designing and operating complex data systems in production or research-adjacent environments.
  • Strong proficiency in SQL and Python.
  • Experience working with data-intensive or distributed systems.
  • Proven experience building scalable data pipelines supporting machine learning training, evaluation, or inference workflows.
  • Strong understanding of data modeling, dataset lifecycle management, and data quality best practices.
  • Ability to operate effectively in highly ambiguous problem spaces while collaborating with researchers and architects.
  • Demonstrated ability to provide technical leadership through direct contribution, mentorship, and establishment of engineering standards.
  • Strong communication skills, including the ability to explain technical tradeoffs to both research and engineering audiences.

Preferred Qualifications

  • Experience with large-scale dataset curation and annotation workflows.
  • Experience with experiment tracking and reproducibility tooling.
  • Familiarity with Databricks, Apache Spark, lakehouse architectures, and cloud data warehouses.
  • Experience with Kafka, Pub/Sub, or event-driven data architectures.
  • Familiarity with Airflow, Dagster, and data quality or lineage tools.
  • Experience with AWS or GCP cloud environments.
  • Experience working with containerized data workloads, CI/CD, and infrastructure-as-code.
  • Experience optimizing data pipelines for GPU-backed training and large-scale inference workloads.
  • Understanding of performance optimization and cost management for large-scale machine learning data infrastructure.

Disclaimer

SpringCube curates tech job listings from various company websites to support tech professionals globally.

  1. No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
  2. No Client Relationship: This company is not a client of SpringCube unless stated.
  3. To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
  4. No Liability: SpringCube is not liable for inaccuracies.
‹