Data Engineer (Expert) 3352
via Mediro ICT · Pretoria, Gauteng
- Pay
- On application
- Type
- Contract
- Sector
- IT & Internet
- Closes
- in 18 days
- Recruiter
- Mediro ICT · All jobs at Mediro ICT
About the role
Our client in Pretoria is recruiting for a Data Engineer (Expert) to join their team.
- Design, build and operate operational data integration solutions focused on making enterprise data available, trusted and consumable.
- Ingest and stream data from various operational sources into real-time pipelines using AWS-native and opensource streaming technologies.
- Build and maintain Kafka topics, producers/consumers, connectors and streaming applications to enable realtime flows.
- Implement API-based integrations, webhook listeners and file ingestion solutions.
- Develop robust ETL/ELT and stream processing logic (Glue, Lambda, Spark on EMR) to normalise, correlate and enrich data across multiple sources.
- Ensure pipelines are idempotent and resilient with retries, replay, backpressure and compensation strategies.
- Apply data modelling and metadata practices to ensure consistent schemas and discoverability (AWS Glue Data Catalog, Lake Formation).
- Implement and monitor data quality checks, anomaly detection and validation rules; drive remediation where needed using AWS monitoring tools.
- Maintain data lineage, access control and security practices to meet governance and compliance requirements (IAM, KMS, Lake Formation).
- Collaborate with automation and orchestration teams to deploy integration components using AWS-focused IaC and CI/CD pipelines.
- Own pipeline alerting, runbooks, and participate in incident response/on-call rotations; troubleshoot production issues and perform root cause analysis using AWS observability tooling.
- Build testable pipelines with unit/integration tests, contract tests and CI/CD deployment pipelines for streaming workloads targeting AWS environments.
What they're looking for
- Qualifications/Experience:
- Extensive hands-on experience (typically 6+ years) in data engineering, integration or streaming roles with demonstrable production experience.
- Proven track record building and operating streaming platforms (Kafka/MSK) and API-based integrations, with strong Python and Java skills and experience with enterprise databases and query languages.
- Strong analytical thinking, curiosity about data, attention to detail, structured problem solving and ownership — able to drive topics to completion.
- Preferred certifications: Confluent Certified Developer for Apache Kafka, Microsoft streaming eventhubs, AWS Certified Data Engineer.
- Essential Skills Requirements:
- Hands-on experience with Kafka and event streaming platforms for real-time data movement.
- Proven experience with API integration patterns, webhooks and event/webhook ingestion.
- Strong proficiency in Python for data engineering, ingestion pipelines and automation.
- Strong proficiency in Java for stream processing or connector development.
- Solid competence with enterprise databases and query languages, including performance tuning and query optimisation for OLTP/operational workloads.
- Experience with NoSQL/document stores such as Amazon DynamoDB, MongoDB.
- Experience in data modelling to design schemas and standardised data representations.
- Experience with schema registries and contract-first designs (Avro, Protobuf) to manage producer/consumer compatibility.
- Strong understanding and practice of data quality techniques and tooling to ensure trusted data.
- Knowledge of metadata management and cataloguing to support discoverability and lineage.
- Familiarity with ETL/ELT patterns and best practices for performant, reliable data pipelines.
- Observability for streaming: experience with metrics, tracing and logging on AWS (CloudWatch, OpenTelemetry, Prometheus/Grafana).
- Advantageous Skills Requirements:
- Awareness of frontend frameworks (e.g., Angular) to better understand downstream consumers.
- Experience operating container platforms and orchestration (Kubernetes/EKS) for scalable stream processing on AWS.
- Familiarity with enterprise systems like SAP and working with their integration interfaces.
- Experience with big data ecosystems (e.g., EMR, S3, Hadoop) and distributed storage/processing on AWS.
- Working knowledge of AWS analytics/data platform services (Glue, Athena, Kinesis, Redshift, Lake Formation, MSK).
- Knowledge of message delivery semantics, partitioning strategies and capacity planning for high-throughput pipelines on AWS.
More IT and developer jobs · More IT and developer jobs in Pretoria · All jobs at Mediro ICT
