Loading open roles
Loading open roles
Loading role
HireBound · posted 9 months ago
Job Title: Kafka DevOps Engineer
Department: Data Platform & Integration
Job Purpose: We are building the Common Data Backbone (CDB)—Fugro’s strategic data platform, which enables discovery, governance, and integration across our global geospatial data ecosystem. The CDB connects multiple cloud services and end-user applications through Apache Kafka, serving as the integration solution within the CDB for event orchestration and integration services.
To further develop and deploy our CDB, we want to strengthen the team with an experienced Kafka DevOps Engineer who will expand and mature the Kafka infrastructure on AWS. This role focuses on secure cluster setup, lifecycle management, performance tuning, Dev-Ops and reliability engineering to ensure Kafka runs at enterprise-grade standards.
Key Responsibilities
• Design, deploy, and maintain secure, highly available Kafka clusters in AWS (MSK or self-managed).
• Perform capacity planning, performance tuning, and proactive scaling.
• Automate infrastructure and configuration using Terraform and GitOps principles.
• Implement observability: metrics, Grafana dashboards, CloudWatch alarms.
• Develop runbooks for incident response, disaster recovery, and rolling upgrades.
• Ensure compliance with security and audit requirements (ISO27001).
• Collaborate with development teams to provide Kafka best practices for .NET microservices and Databricks streaming jobs.
• Conduct resilience testing and maintain documented RPO/RTO strategies.
• Drive continuous improvement in cost optimization, reliability, and operational maturity.
Required Skills & Experience
• 4+ years in DevOps/SRE roles, with 2+ years hands-on Kafka operations at scale.
• Strong knowledge of Kafka internals: partitions, replication, ISR, controller quorum, KRaft.
• Expertise in AWS services: VPC, EC2, MSK, IAM, Secrets Manager, networking.
• Proven experience with TLS/mTLS, SASL/SCRAM, ACLs, and secure cluster design.
• Proficiency in Infrastructure as Code (Terraform preferred).
• Familiarity with CI/CD pipelines for cluster and topic configuration.
• Monitoring and alerting using Grafana, CloudWatch, and log aggregation.
• Disaster recovery strategies.
• Strong scripting skills (Bash, Python) for automation and tooling.
• Excellent documentation and communication skills.
• Kafka Certification.
Nice-to-Have
• Experience with AWS MSK advanced features.
• Knowledge of Schema Registry (Protobuf) and schema governance.
• Familiarity with Databricks Structured Streaming and Kafka Connect.
• Certifications: Confluent Certified Administrator, AWS SysOps/Architect Professional.
• Databrick data analyst/engineering and Certification
• Geo-data experience