Job Description
This job was selected by the SpringCube team to help AI, Data and Cloud Engineers discover relevant San Francisco Bay Area employers. Sign up to view the full employer details and apply directly with the hiring company.
Company Overview
A technology company is developing a new generation of computing experiences designed to make computers feel more natural and human. Its work focuses on voice agents and advanced AI technologies, bringing together experienced leaders and technical experts across hardware and software to develop systems capable of seeing, hearing, and collaborating with people in natural ways.
The Data & Analytics organization supports the development of AI systems through sophisticated data infrastructure and engineering capabilities. Its data includes conversations, voice recordings, sensor signals, product telemetry, and other complex multimodal sources.
The organization is seeking a Data Engineer, Machine Learning to build and maintain the data pipelines that support AI and machine learning models. This role will work closely with machine learning engineers and researchers to ensure they have the right data, in the right format, and at the right time to train, evaluate, and deploy models.
Key Responsibilities
- Design and build production data pipelines that prepare conversational, voice, and multimodal data for model training and evaluation.
- Partner directly with machine learning engineers to understand data requirements for new models and experiments.
- Deliver high-quality datasets that meet the requirements of machine learning workflows.
- Build and maintain infrastructure for dataset versioning, lineage tracking, and reproducibility.
- Ensure training runs can be traced back to their specific source datasets and data versions.
- Develop data quality frameworks that identify potential issues before they impact model quality.
- Implement schema validation, data drift detection, and coverage monitoring.
- Optimize large-scale data processing for cost and performance across cloud infrastructure.
- Build internal tools that enable ML engineers and researchers to independently discover, explore, and request data.
- Define and enforce data governance and privacy standards, particularly for sensitive conversational and voice data.
- Contribute to architectural decisions for the broader data platform as data volume and organizational requirements grow.
- Support the complete machine learning data lifecycle, from data collection and labeling through training and evaluation.
- Collaborate closely with ML engineers and researchers to improve data workflows and accelerate model development.
Required Qualifications
- 5+ years of experience in data engineering, including meaningful experience supporting machine learning or AI teams.
- Strong SQL and Python skills with the ability to use both regularly in production environments.
- Experience building and operating ETL/ELT pipelines at scale using modern data platforms and tooling.
- Experience with workflow orchestration systems such as Airflow, Dagster, or Prefect.
- Hands-on experience with machine learning data workflows, including training data pipelines, dataset versioning, data labeling pipelines, or model evaluation data.
- Strong understanding of machine learning workflows and the characteristics of high-quality training datasets.
- Understanding of how data quality can directly affect machine learning model performance.
- Experience working with unstructured and semi-structured data, including audio, text, and JSON logs.
- Strong communication skills and the ability to bridge data infrastructure requirements with machine learning needs.
- Ability to collaborate effectively with machine learning engineers, researchers, and other technical teams.
Preferred Qualifications
- Experience with vector databases, embedding storage, or feature stores.
- Experience working with hardware or embedded-system data, including telemetry, sensor data, or real-time streams.
- Experience with distributed computing frameworks such as Ray or Spark.
- Experience with Kubernetes and managed Kubernetes environments such as GKE or EKS.
- Knowledge of data privacy frameworks, particularly those applicable to voice and conversational data.
- Experience building internal developer tools or self-service data platforms.
- Experience designing infrastructure that supports scalable machine learning data workflows.
- Familiarity with data lineage, reproducibility, and governance practices for AI systems.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
1.No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
2.No Client Relationship: This company is not a client of SpringCube unless stated.
3.To Apply: Click the “Apply” button to be redirected to the hiring company’s application page for this job.
4.No Liability: SpringCube is not liable for inaccuracies.