California, Sunnyvale
08/12/2026
Contract
Active
Job Summary
We are seeking an experienced Senior Data Engineer with strong Spark and streaming expertise to build real-time, scalable data pipelines using technologies such as Spark, Kafka, and cloud services. The ideal candidate will have experience designing and optimizing data solutions on GCP to support analytics and machine learning workloads.
Key Responsibilities
• Design, develop, and maintain ETL/ELT data pipelines for batch and real-time data ingestion, transformation, and loading using Spark (PySpark/Scala) and streaming technologies such as Kafka and Flink.
• Build and optimize scalable data architectures, including data lakes, data warehouses such as BigQuery, and streaming platforms.
• Optimize Spark jobs, SQL queries, and data processing workflows for speed, efficiency, and cost-effectiveness.
• Implement data quality checks, monitoring, and alerting systems to ensure data accuracy and consistency.
• Develop and maintain scalable data processing solutions supporting analytics and machine learning workloads.
Required Qualifications
• Minimum 8 years of total IT experience.
• Minimum 4 years of recent experience working with Google Cloud Platform (GCP).
• Strong proficiency in Python and SQL, with experience in Scala and/or Java preferred.
• Strong expertise in Apache Spark, including Spark SQL, DataFrames, and Spark Streaming.
• Hands-on experience with streaming technologies such as Apache Kafka, Flink, or Pub/Sub.
• Experience designing and implementing batch and real-time data pipelines.
• Experience with GCP data services and cloud-based data engineering solutions.
• Knowledge of data warehousing technologies such as Snowflake and Redshift.
• Knowledge of NoSQL databases.
• Strong understanding of data quality, monitoring, and alerting practices.
Preferred Qualifications
• Experience with Azure data services.
• Experience with Apache Airflow.
• Experience with Databricks.
• Experience with Docker and Kubernetes.
• Experience optimizing data pipelines and distributed processing workloads for performance and cost efficiency.
.
.
.