Job Description
This job was selected by the SpringCube team to help AI, Data and Cloud Engineers discover relevant San Francisco Bay Area employers. Sign up to view the full employer details and apply directly with the hiring company.
Company Overview
A Silicon Valley-based technology company is developing cutting-edge AI and machine learning products designed to address healthcare staffing challenges and administrative burdens. Its solutions leverage conversational AI and patented voice biomarker technology to support better healthcare delivery and improve experiences for healthcare organizations and their users.
Key Responsibilities
Drive High-Impact Projects
- Own data science and AI evaluation projects from problem definition and methodology development through analysis, communication of results, and implementation.
- Lead projects that generate actionable insights and measurable improvements to AI systems.
Drive Continuous Improvement of LLM-as-a-Judge
- Build data and evaluation infrastructure for a closed-loop AI quality system.
- Develop and maintain golden datasets and ground-truth benchmarks.
- Calibrate judge models and conduct detailed error analysis.
- Incorporate production feedback into ongoing evaluation processes.
- Identify AI failure modes and translate findings into improvements for datasets, evaluation methodologies, LLM-as-a-Judge systems, and underlying AI products.
Develop Evaluation Frameworks
- Design and implement robust evaluation frameworks, methodologies, and metrics for LLM- and speech-based AI systems.
- Develop scalable approaches for measuring AI quality and reliability.
- Establish evaluation processes that support faster customer onboarding and continuous product improvement.
Required Qualifications
- 5+ years of industry experience in data science, machine learning, or a related quantitative field.
- Bachelor’s or Master’s degree in Computer Science, Data Science, Statistics, Mathematics, or a related quantitative discipline.
- Strong proficiency in Python and relevant libraries such as Pydantic, Pandas, NumPy, and Scikit-learn.
- Strong foundation in statistical analysis, experimental design, hypothesis testing, and A/B testing methodologies.
- Experience working with large-scale datasets, data pipelines, and analytical tools to generate actionable insights.
- Familiarity with evaluation metrics such as precision, recall, and agreement metrics including Cohen’s Kappa.
- Experience calibrating AI judge models against human ground truth.
- Demonstrated ability to independently own and drive projects from problem definition through execution and implementation.
- Excellent communication and interpersonal skills.
- Ability to collaborate effectively with AI/ML Engineers, Product Managers, and other cross-functional stakeholders in a fast-paced environment.
Preferred Qualifications
- Experience with Natural Language Processing (NLP) techniques.
- Understanding of LLM architectures and the challenges associated with evaluating large language models.
- Experience with speech technologies and speech system evaluation.
- Experience working in a regulated industry such as healthcare or finance.
- Familiarity with data governance principles, particularly those related to PII and PHI.
- Experience with MLOps practices and deploying machine learning models to production.
- Familiarity with Langfuse.
- Experience with cloud platforms such as GCP, AWS, or Azure.
- Experience with data platforms such as Databricks.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
1.No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
2.No Client Relationship: This company is not a client of SpringCube unless stated.
3.To Apply: Click the “Apply” button to be redirected to the hiring company’s application page for this job.
4.No Liability: SpringCube is not liable for inaccuracies.