I am an analytical and self-driven Data Engineer and BSc IT graduate building production-grade data systems from the ground up. My work spans end-to-end pipelines, raw ingestion through transformation, NLP intelligence, ML forecasting, and BI dashboards, with deep focus on Kenya and Africa-wide domain problems in fintech, employment intelligence, real estate, price tracking, public-sector data, and document intelligence.
I have shipped 80+ open-source projects across the full data engineering stack: Airflow orchestration, dbt transformation layers, DuckDB/PostgreSQL storage, HuggingFace NLP models, XGBoost/Prophet forecasting, and React-based data products. Every project is containerised, tested, and built to production standards.
- Batch & Streaming Pipelines: End-to-end workflows using Apache Airflow 3.0 (Task SDK, SimpleAuthManager, asset-based scheduling) for orchestration and Apache Kafka for real-time streaming and decoupled ingestion.
- ELT Architecture & Data Modeling: dbt-driven staging/intermediate/mart layers with automated quality testing (54–72 tests per project) across PostgreSQL, DuckDB, BigQuery, and Snowflake.
- NLP & Intelligence Pipelines: Named entity recognition and sentiment analysis using spaCy, HuggingFace Transformers (FinBERT, zero-shot classification, token classification), and MLflow experiment tracking, applied to news, job postings, legislation, and financial documents.
- ML Forecasting & Predictive Analytics: Time-series forecasting with Prophet (29–109 models per project), regression and classification with XGBoost and SHAP explainability across Kenya, EAC, and Africa-wide datasets.
- Web Scraping & Price Intelligence: Production-grade scrapers using Playwright (session priming for Cloudflare-protected sites) and BeautifulSoup — e-commerce pricing, classified listings, tenders, and property data.
- Cloud & Infrastructure: Containerised solutions on GCP (BigQuery, GCS) and AWS (S3, Lambda, SNS) with Terraform IaC; all local projects run on Docker Compose with isolated environments.
- API Development & Data Serving: Production REST APIs with FastAPI and Pydantic schemas; vector search with pgvector for embedding-based job intelligence.
- Analytics Engineering: Interactive dashboards with Streamlit, Plotly Dash, Apache Superset, Evidence.dev, and Grafana — from fiscal reconciliation to real-time price tracking.
| Project | Stack | Scale | GitHub |
|---|---|---|---|
| Kenya Financial Document Intelligence | Airflow 3.0 · pdfplumber · HuggingFace · spaCy · dbt | 87 docs · 2,490 pages · 40,801 entities | → |
| JobSense Kenya | Airflow 3.0 · FastAPI · pgvector · HuggingFace · Streamlit | 604 jobs · 7 sources · 604 embeddings | → |
| LedgerSync | Airflow 3.0 · DuckDB · dbt · Evidence.dev | 1.5M rows · 9 dbt models · 2,961 alerts | → |
| Africa Health Investment Returns | XGBoost · SHAP · scikit-learn | 53 countries · 1,219 rows · R²=0.945 | → |
| CupCast 2026 | XGBoost · Poisson · Monte Carlo · React + Vite | 10k simulations · 7 JSON APIs · React UI | → |
| Kenya News NLP Intelligence | Airflow 3.0 · PostgreSQL · dbt · spaCy · Streamlit | 70 articles · 1,308 entities · 63 dbt tests | → |
| Ecommerce Analytics | Airflow 3.0 · DuckDB · dbt · Superset | 8/8 DAG · 15 models · 54 dbt tests | → |
| EAC Inflation Forecaster | Prophet · World Bank API · IMF API · dbt | 50 models · 5 countries · 10 indicators | → |
- Portfolio: ian-mwendwa.vercel.app
- LinkedIn: Ian (Dankan) Mwendwa
- Email: idankan571@gmail.com
- Instagram: @faust._