Job Description
The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.
Company Overview
The organization is establishing a new state of trust for global commerce and capital markets through automating and streamlining the work of assurance and audit practitioners specifically within cybersecurity, privacy, and financial audit. Put simply, it builds software for the people who enable trust between businesses.
What You’ll Own
Measurable AI Agents
- Design and build a unified evaluation platform that serves as the single source of truth for all agentic systems and audit workflows.
- Build observability systems that surface agent behavior, trace execution, and failure modes in production, and feedback loops that turn production failures into first-class evaluation cases.
- Own the evaluation infrastructure stack including integration with LangSmith and LangGraph.
- Translate customer problems into concrete agent behaviors and workflows.
- Integrate and orchestrate LLMs, tools, retrieval systems, and logic into cohesive, reliable agent experiences.
Rapid Model Evaluation
- Build automated pipelines that evaluate new models against all critical workflows within hours of release.
- Design evaluation harnesses for the most complex agentic systems and workflows.
- Implement comparison frameworks that measure effectiveness, consistency, latency, and cost across model versions.
- Design guardrails and monitoring systems that catch quality regressions before they reach customers.
AI-Native Engineering Execution
- Use AI as core leverage in how systems are designed, built, tested, and iterated.
- Prototype quickly to resolve uncertainty, then harden systems for enterprise-grade reliability.
- Build evaluations, feedback mechanisms, and guardrails so agents improve over time.
- Work with SMEs and ML Engineers to create evaluation datasets by curating production traces.
- Design prompts, retrieval pipelines, and agent orchestration systems that perform reliably at scale.
Ownership of Quality and Large Product Areas
- Define and document evaluation standards, best practices, and processes for the engineering organization.
- Advocate for evaluation-driven development and make it easy for the team to write and run evaluations.
- Partner with product and ML engineers to integrate evaluation requirements into agent development from day one.
- Take full ownership of large product areas rather than executing on narrow tasks.
Who You Are
The ideal candidate is an engineer who believes that evaluations are foundational to building reliable AI systems, not a nice-to-have. The following operating principles should resonate with the candidate:
- Evaluation-first mindset: Understands that for an AI company, not being able to evaluate a new model quickly is unacceptable.
- AI-native instincts: Treats LLMs, agents, and automation as fundamental building blocks and parts of the craft of engineering.
- Data-driven rigor: Makes decisions based on metrics and is focused on measuring what matters.
- Production-oriented: Understands that evaluations must work on real production behavior, not just offline datasets.
- Strong product judgment: Can decide what matters and why without waiting for guidance, not just how to implement it.
- Bias to building: Moves quickly and builds working systems rather than relying on perfect specifications.
Experience
The organization cares more about capability and trajectory than years on a resume, but most strong candidates will have:
- Multiple years of experience shipping production software in complex, real-world systems.
- Experience with TypeScript, React, Python, and Postgres.
- Experience building and deploying LLM-powered features serving production traffic.
- Experience implementing evaluation frameworks for model outputs and agent behaviors.
- Experience designing observability or tracing infrastructure for AI/ML systems.
- Experience working with vector databases, embedding models, and RAG architectures.
- Experience with evaluation platforms such as LangSmith, Langfuse, or similar technologies.
- Comfort operating in ambiguity and taking responsibility for outcomes.
- Deep empathy for professional-grade, mission-critical software. Experience with audit and accounting workflows is not required.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
- No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
- No Client Relationship: This company is not a client of SpringCube unless stated.
- To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
- No Liability: SpringCube is not liable for inaccuracies.