Job Description
The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.
Company Overview
A fast-growing AI technology company is developing ambitious AI products focused on building reliable and verifiable artificial intelligence systems. The organization places strong emphasis on rigorous evaluation, verification, and measurable performance to ensure that AI systems deliver accurate and dependable results.
The organization is seeking an AI Engineer, Evaluation to build the testing and verification systems that determine whether AI products and agent systems actually work as intended. This role sits close to the core of the product and focuses on designing evaluation suites, verification harnesses, and quality gates that establish clear evidence of AI system performance.
The successful candidate will be skeptical by default, rigorous about measurement, and motivated by turning uncertain results into reliable evidence. The role combines software quality assurance, engineering discipline, analytical thinking, and domain judgment to establish a high standard for AI system reliability.
Key Responsibilities
- Design and build evaluation suites for AI products and agent systems.
- Build verification harnesses that confirm AI systems have produced the intended results.
- Define quality gates that determine what is ready to ship and what requires further development.
- Translate domain expertise and customer requirements into measurable and testable checks.
- Identify and investigate failure modes, including regressions, hallucinations, and silent errors.
- Make evaluation results clear and actionable for engineering, product, and customer-facing teams.
- Integrate evaluation systems into CI pipelines and the broader software development lifecycle.
- Establish and continuously raise standards for determining whether AI systems are functioning correctly.
- Develop automated testing and verification processes that improve the reliability of AI products.
- Collaborate with engineering and product teams to identify areas where evaluation can improve system quality.
Required Qualifications
- Experience testing, evaluating, or performing quality assurance for complex software systems.
- Strong understanding of how LLMs and AI agent systems can fail.
- Strong analytical skills with a rigorous and skeptical approach to evaluating system performance.
- Ability to write code for testing frameworks, verification harnesses, and automation.
- Exceptional attention to detail and a strong commitment to software quality.
- Excellent written and verbal communication skills.
- Ability to operate effectively in a fast-moving and highly demanding environment.
- Strong sense of ownership and accountability.
Preferred Qualifications
- Experience building LLM evaluations, benchmarks, or AI test infrastructure.
- Background in QA, SDET, software testing, or test automation.
- Domain expertise in an industry or vertical where correctness and reliability are critical.
- Experience with statistical evaluation and measurement methodologies.
- Experience developing automated verification systems for AI or complex software products.
Compensation
- Competitive salary.
- Meaningful equity.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
- No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
- No Client Relationship: This company is not a client of SpringCube unless stated.
- To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
- No Liability: SpringCube is not liable for inaccuracies.