Job Description
The SpringCube team curated the following job opportunity to help you in your job search. Explore the position below to find your next career move.
Company Overview
A rapidly growing technology company is focused on transforming the television advertising industry through the application of artificial intelligence, machine learning, sophisticated media buying, and proprietary analytics. Its technology platform enables businesses to reach customers through linear and streaming television advertising while providing an automated, digital-like advertising experience.
Key Responsibilities
- Own the reliability and availability of data platform infrastructure across development, staging, and production environments.
- Establish and improve disciplined environment promotion processes, ensuring appropriate controls between staging and production.
- Define and maintain standard operating procedures for deployments, maintenance windows, and change management.
- Instrument and monitor platform health using observability tools and develop meaningful alerting systems.
- Participate in architecture and deployment discussions and identify potential risks before systems reach production.
- Collaborate with data scientists, engineers, and product managers to address infrastructure requirements.
- Identify and remediate reliability risks before they develop into production incidents.
- Support customer-facing and internal systems with a strong focus on stability and operational excellence.
- Administer, scale, secure, and continuously improve data platform infrastructure.
- Promote methodical and controlled approaches to production changes and infrastructure management.
Required Qualifications
- 3+ years of experience in cloud infrastructure, Site Reliability Engineering (SRE), platform engineering, or a related field.
- Experience with AWS is preferred; experience with GCP or Azure is also applicable.
- Strong operational discipline and experience maintaining production systems.
- Experience with high-availability architectures, including blue/green deployments, data replication, and load balancing.
- Experience with workflow orchestration, such as Airflow, DAG-based schedulers, job scheduling, or large-scale cron systems.
- Strong Linux fundamentals and scripting experience using Bash, Python, or similar languages.
- Experience with distributed data processing technologies such as Spark, PySpark, or comparable frameworks, or experience managing clusters that operate them.
- Experience with containerization and orchestration technologies such as Kubernetes and Docker.
- Experience with data ingestion, ETL, or streaming systems such as Kafka and Flink, or experience operating message queues and data pipelines.
- Experience with infrastructure-as-code and provisioning technologies such as Terraform and Helm.
- Experience with OLAP and OLTP databases such as ClickHouse, PostgreSQL, Redshift, or comparable platforms.
- Understanding of query patterns, indexing, and database operational practices.
- Experience with monitoring, logging, and observability platforms such as Datadog, Prometheus, Kibana, or similar technologies.
- Experience administering and scaling managed data platforms such as Databricks.
- Understanding of network infrastructure fundamentals, including load balancers, DNS, auto-scaling, multi-region architectures, and proxies.
- Knowledge of security and access management principles, including least-privilege access, secrets management, and controls for data systems.
- Familiarity with MLOps concepts or tooling is a plus.
What the Organization Values
- Humility: Recognizes limitations, asks questions, and seeks guidance when working in unfamiliar areas.
- Methodical Execution: Minimizes unnecessary variables, avoids premature optimization, and completes work before taking on additional initiatives.
- Communication: Communicates planned changes clearly, particularly when working with shared or production environments.
- Ownership: Takes responsibility when issues occur and looks for solutions rather than shifting blame.
- Independence: Can independently drive projects from ambiguous requirements through high-quality delivery while recognizing when assistance is needed.
- Operational Discipline: Treats production environments with care and follows established maintenance, deployment, and change-management practices.
- Reliability Mindset: Prioritizes platform stability, availability, and long-term operational health.
Disclaimer
SpringCube curates tech job listings from various company websites to support tech professionals globally.
- No Endorsement: Job ads on SpringCube do not imply endorsement of their authenticity or quality.
- No Client Relationship: This company is not a client of SpringCube unless stated.
- To Apply: Click the Apply button to be redirected to the hiring company’s application page for this job.
- No Liability: SpringCube is not liable for inaccuracies.