Back to Job Listings

Machine Learning Infrastructure Engineer

San Francisco

SpringCube

Full-time - Senior Engineer

Retail & Ecommerce

Posted 3 weeks ago

Disclosed upon interview

Contact Employer
  • Share:
Send Feedback
Report This Job

Job Description

The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.

The salary or hourly rate range may be inclusive of several levels that would be applicable to the position. Final salary or hourly rate will be based on a number of factors including level, relevant prior experience, skills, and expertise. This range is inclusive of base salary or hourly rate only and does not include benefits or equity.

Company Overview

A rapidly growing live commerce and marketplace platform is building the infrastructure needed to support a new generation of online commerce. The organization enables buyers and sellers to connect through live shopping experiences across categories including trading cards, fashion, electronics, live plants, and other products.

With a rapidly expanding global presence, the organization is focused on developing innovative technology, moving quickly, staying closely connected to users, and creating meaningful product experiences. Its engineering teams work across multiple global hubs while maintaining a strong emphasis on collaboration, innovation, and measurable impact.

Role Overview

The organization is seeking a Machine Learning Infrastructure Engineer to design and scale the core infrastructure that powers machine learning systems and self-hosted large language model applications across the business.

The successful candidate will work closely with machine learning scientists to bring advanced models into production and develop infrastructure that makes sophisticated ML systems dependable, efficient, and scalable. The role will involve low-latency large-model serving, distributed training, high-throughput GPU inference, and the development of infrastructure supporting critical product experiences.

Key Responsibilities

  • Own the infrastructure powering AI and ML models across critical business areas, including growth, recommendations, trust and safety, fraud, seller tooling, and other product surfaces.
  • Prototype, deploy, and productionize innovative machine learning architectures that directly influence user experiences and marketplace dynamics.
  • Design and scale inference infrastructure capable of serving large models with low latency and high throughput.
  • Build distributed training and inference pipelines using GPUs, model parallelism, and data parallelism.
  • Develop reliable infrastructure that supports the deployment and operation of advanced machine learning and large language models.
  • Collaborate closely with machine learning scientists and cross-functional teams to bring cutting-edge models into production.
  • Drive technical initiatives across multiple product areas and communicate findings and recommendations to leadership and product teams.
  • Take on complex technical challenges as AI and ML capabilities continue to scale across the organization.
  • Contribute to reliable, reproducible, and well-tested engineering practices within a distributed working environment.

Required Qualifications

  • 4+ years of professional experience developing machine learning systems and algorithms.
  • 3+ years of software engineering experience building and maintaining production systems serving consumer-scale workloads.
  • 1+ year of professional experience developing software using Python.
  • Bachelor’s degree in Computer Science, Statistics, Applied Mathematics, or a related technical field, or equivalent professional experience.
  • Ability to work autonomously and drive initiatives across multiple product areas.
  • Strong ability to communicate technical findings and recommendations with leadership and product teams.
  • Experience working with operational, search, and key-value databases such as PostgreSQL, DynamoDB, Elasticsearch, and Redis.
  • Strong understanding of visualization and monitoring tools such as DataDog and Grafana.
  • Familiarity with cloud computing platforms and managed services such as AWS SageMaker, Lambda, Kinesis, S3, EC2, EKS/ECS, Apache Kafka, and Flink.
  • Ability to collaborate effectively in a remote or distributed working environment.
  • Strong commitment to well-tested, reproducible engineering practices.
  • Exceptional documentation and communication skills.

Working Arrangement

The role offers flexibility to work from home or from one of the organization’s global office hubs. In-person collaboration is valued for planning, problem-solving, and team connection.

For US-based team members, individuals in this role are expected to live within commuting distance of the New York, Seattle, Los Angeles, or San Francisco hubs.

Disclaimer

SpringCube curates tech job listings from various company websites to support tech professionals globally.

  1. No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
  2. No Client Relationship: This company is not a client of SpringCube unless stated.
  3. To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
  4. No Liability: SpringCube is not liable for inaccuracies.
‹