Our work
What we have built.
Engagements described by sector and scale. Clients are not named — see the note at the foot of each study.
470tables migrated to BigQuery
Retiring a Hadoop estate: 470 tables, one migration framework
A lift-and-shift of on-premise Hadoop, Netezza and Oracle into BigQuery — built as a reusable framework so adding tables cost almost nothing, validated table by table, targeting a 15% improvement in data availability SLAs.
4 min read
40engineers, 10+ projects
40 engineers, 10+ projects, 5 frameworks that held it together
A multi-year enterprise data engagement where the durable output was not any single project, but the shared frameworks — data quality, sync, testing, deployment — that every workstream was built on.
4 min read
3,700+features, from 5% of data to all
Recognising the customer before they leave the page
Real-time omnichannel pipelines — cross-channel identification for anonymous sessions, 40+ fulfilment KPIs, and a feature pipeline that took propensity models from 5% of the data to 3,700+ features.
4 min read
40+KPIs defined and delivered
40+ KPIs, four dashboards, and one agreed definition each
Real-time digital, store, payment and pathing analytics for a retailer — where the work that made them trustworthy happened before a single chart was drawn.
4 min read
2016founding platform architecture
One version of the truth, at Kafka scale
The founding architecture for a retailer's omnichannel analytics platform — mirrored event streams, real-time and batch KPIs into Redshift, and a bucketing design that keeps the numbers correct when brokers fail.
4 min read
6parallel workstreams
Multi-tenant data processing for a super-app
Discovery and delivery for a Southeast Asian ride-hailing platform — multi-tenant processing with governance, internationalising a scheduling system, and adding Spark and Beam beneath an in-house framework.
3 min read
18h → 3htraining time, 19,000 models
Six times faster model training, across 30 million customers
Rebuilding a retailer's propensity engine to run on the complete dataset instead of a 5% sample — cutting training from 18 hours to three across 19,000 models and lifting scoring accuracy 14%.
4 min read
17K/secevents ingested, 3 TB/day
17,000 events a second, live in one week
Replacing a legacy stack that had stopped scaling with a real-time pipeline ingesting 3 TB a day — delivered in a single week, and paid for by decommissioning what it replaced.
3 min read
+64%event stock accuracy
64% better stock accuracy by forecasting the town, not the store
A retailer kept running out of stock and nobody knew why. The answer was local events — and joining three years of external event data to inventory history lifted event-related stock accuracy 64% and revenue 10%.
3 min read
Analysing health records without ever seeing who they belong to
A risk and cost analytics platform for a US healthcare organisation, built so that patient records were anonymised before they left the source — putting the privacy boundary at the start of the pipeline rather than the end.
3 min read
3 → 1platforms consolidated
Three platforms, three teams, one data lake
Consolidating parallel Teradata, Netezza and Hadoop estates into a single core platform — ending duplicated engineering and giving merchandising, logistics, sales and marketing one set of numbers.
3 min read
+44%product tagging accuracy
The searches that returned nothing — and 44% better product tagging
Mining on-site and inbound search intent to find demand the catalogue could not answer, lifting product tagging accuracy 44% and sales 10%.
3 min read