For our client, a global leader in its industry, we are looking for a Data Integration & Systems Engineer
Working mode: Hybrid
Type of contract: B2B
Primary purpose:
Design, build, and operate scalable data pipelines and data flows that enable Next Best Action (NBA) decisioning to run reliably across on-premises and cloud platforms. Ensure that data is correctly migrated from Hadoop/Cloudera to Google Cloud Platform (GCP) and integrated with Pega CDH/Infinity to support real-time and batch decisioning for the MVP and future scale.
Key responsibilities:
Design and implement end-to-end data pipelines across on-premises and Google Cloud Platform environments
Build and maintain batch and near-real-time data flows, ensuring availability, freshness, and performance
Drive the migration and modernisation of data from Hadoop/Cloudera to GCP (BigQuery and related services)
Ensure reliable, scalable, and high-performance data delivery for decisioning and analytics use cases
Implement and maintain data integration with Pega CDH/Infinity, ensuring that data is correctly structured and available for decision strategies and AI models
Ensure data quality, lineage, traceability, and compliance with governance and regulatory requirements
Implement monitoring, validation, and reconciliation processes for data pipelines
Collaborate with data scientists, decisioning subject matter experts (SMEs), architects, and product teams to align data flows with business needs
Apply modern development practices, including CI/CD, version control, and automated testing
Use AI-assisted development tools (e.g., GitHub Copilot, Codex, Claude, Google AI tools) to improve productivity and quality
Essential technical experience:
Google Cloud Platform: hands-on experience with BigQuery and cloud-based data pipelines (must have)
Hadoop/Cloudera ecosystem: strong experience with on-premises big data platforms and migration to the cloud (must have)
Data pipeline development: proven experience building scalable batch and real-time data flows
Spark and Scala: strong hands-on experience with distributed data processing
Python development: solid programming skills for data engineering use cases
SQL: advanced skills in data querying, transformation, and optimisation
Data modelling: strong experience designing robust and scalable data structures
Data integration: experience integrating data with enterprise platforms, including Pega CDH/Infinity
Software engineering practices: experience with CI/CD, version control (e.g., Git), and testing frameworks
Distributed systems: experience working in large-scale, high-volume data environments
Data quality and governance: experience with validation, monitoring, lineage, and compliance requirements
AI-assisted development: practical experience with tools such as GitHub Copilot or similar
Agile ways of working: experience delivering in cross-functional, iterative environments
