Data

Architect of high-performance data platforms, distributed telemetry, and enterprise-grade analytical infrastructure.

Overview

I am a Data Platform Architect specialising in deploying enterprise data service clusters. I architect the systems that transform raw telemetry into observable, high-fidelity intelligence. Outside of pure-engineering, I studied applied mathematics and statistics at university, and do understand data modelling.

Apache Ecosystem

Building resilient data factories using Airflow for ETL/state-driven orchestration and Kafka for real-time event streaming. I apply software engineering to the development of pipelines to ensure continuous deployment, validity, integrity and reliability across complex environments. I do MLOps using the DVC tools for machine learning projects.

Distributed Analytics

Specialist in high-velocity exploration using Trino and Apache Drill for federated querying across disparate silos. I use Superset for reporting and visualisation, and my enterprise Grafana dashboards for observability and metrics visualisations.

Python Programming

I do Python development - and run Python teams for the data stack. Python is the de facto language of data - and whilst I am heavily open source in my own shop, my tools and methodologies can be applied against the proprietary majors in Data Science - and I’ve probably got the Python and Airflow providers, clients, connectors to do so.

In addition to being competent at SQL, I do major abstracted pieces using SQLAlchemy, Marshmallow, GraphQL and a lot of JSON in RDBMS with Pydantic and tools that bring robust type-safety and object-relational agility to traditional SQL.

I share my expertise with Jupyter and Pandas to the business to lift and shift Excel spreadsheets into rigorous, engineered corporate artifacts. These are source code controlled, lifecycle managed and even trivially automated for execution in the afore-mentioned Airflow.