Lead Data Engineer
ArdhiSasa- Led ArdhiSasa through a national audit by the Office of the Auditor-General covering documentation, user guides, development workflow and methodology, architecture and implementation, and security practices — receiving a positive assessment in the Auditor-General's report on the State Department for Lands and Physical Planning for the year ended 30 June 2025
- Designed and administer PostgreSQL 17 production clusters with hybrid master-replica topology using repmgr HA and PgBouncer connection pooling
- Scaled cluster throughput from 1,000 TPS to 10,000 TPS through architectural redesign and query optimization
- Implemented SCRAM-SHA-256 password encryption across PostgreSQL clusters so the database never handles raw passwords during user creation or rotation, paired with audit triggers and automated backup scripts across 31+ databases
- Implemented S3-based long-term archival of database backups, VM security logs, application logs, and database logs to meet Kenya's statutory data retention requirements
- Defended the platform against an active attack by enforcing end-to-end TLS across every service hop — client-to-server, database connections, and support tooling such as Mattermost and pgAdmin — and restructured credential access around HashiCorp Vault for system-critical authentication
- Built and maintain ETL/ELT pipelines in Python, Apache Airflow, and SQL for land transaction and records data
- Implemented PeerDB change data capture replication from production OLTP to data warehouse with custom schema mapping
- Architected the data warehouse on Data Vault 2.0 and grew data engineering from a non-existent BI function into a fully independent team delivering all of ArdhiSasa's BI, visualization, and reporting on system-generated data for decision makers
- Led digitization of 5 million+ land records through a custom EDMS including a bulk document ingestion ETL tool
- Designed and implemented CI/CD pipelines with GitHub Actions for automated data pipeline testing and production deployments; led the migration of the deployment platform from Docker Swarm to a fully self-hosted, on-premises Kubernetes cluster to improve service delivery
- Improved database observability by rolling out pghero, pgBadger, and pgwatch for query performance monitoring, log analysis, and error tracing across production clusters
- Served as Scrum Master for the data engineering team — sprint planning, backlog grooming, and agile coaching
- PostgreSQL 17
- PeerDB
- Airflow
- Python
- dbt
- Airbyte
- MinIO
- Docker Compose
- GitHub Actions
- Rocky Linux
- nginx
- Nagios
- ELK Stack
- TLS
- HashiCorp Vault
- Data Vault 2.0
- Kubernetes
- pgHero
- pgBadger
- pgwatch
- AWS S3