Back to Job Listings

Machine Learning Infrastructure – Generative AI

San Francisco

SpringCube

Full-time - Senior Engineer

Social Networking & Media

Posted 4 weeks ago

Disclosed upon interview

Contact Employer
  • Share:
Send Feedback
Report This Job

Job Description

The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.

Company Overview

A leading global technology platform is building a centralized Generative AI infrastructure platform that enables engineering teams across its ecosystem to safely bring AI-powered products, agents, automation, and personalization into production. The platform focuses on increasing the speed and business impact of Generative AI through scalable infrastructure, evaluation systems, observability, model serving, guardrails, and cost management.

The organization is seeking a Software Engineer, Machine Learning Infrastructure – Generative AI to join a high-impact engineering team focused on building production infrastructure for Generative AI. The primary focus of this role will be the evaluation and LLM observability platform, enabling engineering teams to evaluate, trace, monitor, and continuously improve the quality of LLM and agent-based products.

Key Responsibilities

  • Build infrastructure that helps engineering teams move Generative AI concepts from prototypes into reliable production systems.
  • Develop and maintain a unified evaluation platform including evaluation SDKs, OpenTelemetry trace and score ingestion, LLM-as-judge workflows, offline and online evaluation pipelines, and agent simulations.
  • Design scalable systems supporting evaluation workflows, LLM observability, trace and score ingestion, and agent simulation.
  • Develop reliable measurement and quality systems that help teams identify regressions and compare the performance of open-weight and closed-source models.
  • Build platforms that support rapid experimentation while meeting production requirements for latency, scalability, monitoring, service-level objectives, operational playbooks, and reliability.
  • Collaborate with machine learning engineers, product engineers, data scientists, and platform teams across multiple business units.
  • Help develop centralized Generative AI platform capabilities that connect evaluation, observability, agent optimization, and post-training workflows.
  • Contribute to systems supporting automated evaluation, agent simulation, reward modeling, and RLHF/RLVR evaluation.
  • Improve the reliability, scalability, and operational excellence of Generative AI infrastructure.
  • Help establish durable platform primitives that support AI-powered products, agents, automation, and personalization.

Required Qualifications

  • Bachelor’s, Master’s, or PhD in Computer Science or an equivalent field.
  • 3+ years of professional experience in software engineering.
  • Strong backend engineering fundamentals, particularly in Python and distributed systems.
  • Experience building production services, APIs, data pipelines, or machine learning infrastructure at scale.
  • Experience operating production systems, including observability, debugging, reliability, incident response, and performance and cost optimization.
  • Hands-on experience with evaluation, LLM observability, or measurement systems for machine learning and LLM products in production.
  • Experience with evaluation pipelines, tracing and scoring, offline and online quality metrics, or experimentation.
  • Proficiency using AI coding tools such as Claude Code, Codex, or Cursor throughout the software development lifecycle, including software design, code generation, testing, monitoring, and release processes.

Preferred Qualifications

  • Experience with LLM evaluation methodologies, including LLM-as-judge design and calibration, evaluation drift detection, human-in-the-loop labeling, or evaluation harnesses for agents and multi-step systems.
  • Experience with LLM observability and tracing, including OpenTelemetry, trace and score ingestion, and instrumentation SDKs.
  • Experience building and deploying AI agents or MCP servers in production, including agent evaluation and simulation.
  • Experience developing data pipelines, streaming ingestion systems, and analytical data stores such as SQL and columnar or OLAP databases for high-volume telemetry.
  • Experience with LLM gateways, model routing, vendor abstraction, or cost attribution systems.
  • Experience building developer platforms, internal platforms, or self-service infrastructure.
  • Experience with Kubernetes and cloud infrastructure such as AWS or GCP.
  • Experience with high-throughput batch processing systems.
  • Knowledge of retrieval-augmented generation (RAG), search systems, vector databases, or open-weight LLM inference and fine-tuning.

Disclaimer

SpringCube curates tech job listings from various company websites to support tech professionals globally.

  1. No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
  2. No Client Relationship: This company is not a client of SpringCube unless stated.
  3. To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
  4. No Liability: SpringCube is not liable for inaccuracies.
‹