Location:
CN-Shenzhen-HyQ
Shift:
Standard - 40 Hours (China)
Scheduled Weekly Hours:
40
Worker Type:
Permanent
Job Summary:
Build and maintain data pipelines and platform components across the full data stack. You will work within an established architecture, contributing reliable, well-tested code while growing your expertise across streaming, batch, storage, and analytics technologies.
Job Duties:
Responsibilities
- Data Pipelines: Develop, test, and deploy batch (Spark) and streaming (Flink/Spark Structured Streaming) jobs on OpenShift
- Message & Stream: Create and manage Kafka topics, configure connectors, monitor consumer lag, troubleshoot message delivery
- Lakehouse & Storage: Create and maintain Iceberg tables (schema changes, partitioning, maintenance); operate MinIO for data lake workloads
- Query & Serving: Write SQL transformations (dbt), build Trino queries for federated analytics, manage materialized views on StarRocks/ClickHouse
- Orchestration: Author and maintain Airflow/Dagster DAGs; troubleshoot pipeline failures and data quality issues
- Observability: Build Grafana dashboards and Prometheus alerts for pipeline health, data freshness, and quality metrics
- Kubernetes Operations: Write basic Helm charts, manage ConfigMaps/Secrets, troubleshoot pod failures under guidance
- CI/CD: Maintain CI pipelines for data jobs, contribute to ArgoCD application definitions
- Write clear documentation: runbooks, pipeline specs, onboarding guides
- Participate in code reviews and agile ceremonies
Required Skills & Experience
- 3 – 6 years in data engineering, software engineering, or DevOps
- Working knowledge of Kubernetes: Pods, Deployments, Services, ConfigMaps, basic Helm usage
- Practical Spark experience: PySpark or Spark SQL batch ETL, reading/writing to object storage and Iceberg tables
- Basic Kafka experience: producing and consuming messages, understanding partitions and offsets
- Understanding of data lake/lakehouse concepts: columnar formats (Parquet/ORC), open table formats (Iceberg/Delta Lake), partitioning, schema evolution
- Experience with an S3-compatible object store (MinIO, AWS S3, Ceph)
- Solid SQL skills; proficient in Python
- Comfortable with Git, Docker, CI/CD, and Linux shell scripting
Nice to Have
- Exposure to Trino or Presto for federated queries
- Familiarity with a real-time OLAP engine: StarRocks, ClickHouse, or Doris
- Experience with dbt for data transformation
- Workflow orchestration: Airflow or Dagster
- Interest in data observability, data mesh, and data quality frameworks
Company Introduction:
ITD SZ
港交所科技(深圳)有限公司,是2016年12月28日于深圳市前海自贸区成立的外商独资企业。
作为港交所的技术子公司,港交所科技(深圳)有限公司主要是为集团及其附属公司提供计算机软件、计算机硬件、信息系统、云存储、云计算、物联网和计算机网络的开发、技术服务、技术咨询、技术转让;经济信息咨询、企业管理咨询、商务信息咨询、商业信息咨询、信息系统设计、集成、运行维护;数据库管理、大数据分析;以承接服务外包方式提供系统应用管理和维护、信息技术支持管理、数据处理等信息技术和业务流程外包服务。