Job Description
This job was selected by the SpringCube team to help AI, Data and Cloud Engineers discover relevant San Francisco Bay Area employers. Sign up to view the full employer details and apply directly with the hiring company.
Company Overview
A technology company is building an identity infrastructure platform designed to help institutions and businesses verify organizations, understand ownership relationships, and assess risk. The platform combines public records, IRS data, sanctions lists, web signals, and fraud telemetry from more than 2,200 financial institutions into a comprehensive business identity graph.
Key Responsibilities
- Build and maintain ETL/ELT pipelines that ingest and normalize public records, web signals, fraud telemetry, and other data sources.
- Develop data models and transformation layers using technologies such as Dataflow, Spark, and Airflow.
- Build data infrastructure that supports fraud detection, Know Your Business (KYB), and customer-facing APIs.
- Implement data quality checks, observability tooling, monitoring, and alerting to identify issues before they affect customers.
- Tune pipelines and queries for performance, data freshness, reliability, and cloud infrastructure costs.
- Collaborate with data scientists, ML engineers, and product teams to provide clean, well-modeled data for entity resolution and scoring systems.
- Help ensure data pipelines meet applicable security and regulatory requirements for sensitive information.
- Support standards and requirements related to SOC 2, GDPR, KYC, and KYB.
- Document data pipelines, models, and infrastructure to improve team knowledge and operational efficiency.
- Communicate effectively with both technical and non-technical stakeholders.
- Take ownership of data engineering projects and contribute production code early in the role.
- Help develop scalable infrastructure for real-time entity resolution, fraud detection, and identity-related applications.
Required Qualifications
- 1+ year of experience in data engineering.
- Experience working with Python and SQL.
- Experience with cloud-native data platforms.
- Hands-on experience building and maintaining production ETL/ELT pipelines.
- Working knowledge of modern data engineering tools such as Dataflow, Spark, Airflow, or equivalent technologies.
- Experience with cloud data warehouses or data lakes such as BigQuery, Snowflake, or equivalent platforms.
- Solid understanding of data modeling principles.
- Strong focus on data integrity, reliability, and quality.
- Experience working with both structured and unstructured data.
- Understanding of scalable data architecture and production data systems.
Preferred Qualifications
- Interest in AI/ML infrastructure and experience working closely with machine learning systems.
- Experience with streaming or real-time data technologies such as Kafka or Pub/Sub.
- Exposure to KYC, KYB, fraud, risk, or underwriting data.
- Experience working with sensitive data and understanding the importance of security, privacy, and responsible data handling.
- GCP experience, particularly with BigQuery, Cloud Run, Dataflow, and Pub/Sub.
- Experience building systems in environments with limited predefined processes.
- Strong communication skills and the ability to translate technical concepts for non-technical stakeholders.
- Ability to receive direct feedback and quickly incorporate it into work.
- Strong curiosity and willingness to learn new technologies and approaches.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
1.No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
2.No Client Relationship: This company is not a client of SpringCube unless stated.
3.To Apply: Click the “Apply” button to be redirected to the hiring company’s application page for this job.
4.No Liability: SpringCube is not liable for inaccuracies.**