Recognising the customer before they leave the page
Real-time omnichannel pipelines — cross-channel identification for anonymous sessions, 40+ fulfilment KPIs, and a feature pipeline that took propensity models from 5% of the data to 3,700+ features.
Omnichannel retailer
3,700+features, from 5% of data to all
Omnichannel is a word that hides a hard engineering problem. A customer browses on a phone at lunchtime, opens a laptop that evening, and walks into a store on Saturday. To the business that is one person with one intent. To the systems, it is three strangers.
We worked with a retailer on several fronts of that problem. The pattern underneath them all was the same: the data existed, but not in time and not joined up — and by the time it was both, the moment to act had passed.
Identifying a customer who has not logged in
The most valuable session is often anonymous. Someone views three products without signing in; a recommendation engine that only recognises logged-in users sees nothing.
The gaps were concrete: no real-time cross-channel identification for unauthenticated activity, nowhere capturing recent product views in a unified way across channels, and no route to serve those views back to the applications that needed them.
We built low-latency pipelines for cross-channel recent-product-view tracking, working across business teams on the identification logic itself — which is as much a definitional problem as a technical one. When are two sessions the same person? is a question engineering cannot answer alone, and getting it wrong in either direction is costly: too loose and you recommend a stranger's browsing back to someone, too strict and you learn nothing.
Forty-plus KPIs, defined before they were built
A second workstream covered real-time fulfilment reporting. During peak trading, teams could not see fulfilment data in time to act on it — packages needed tracking and intervening on while there was still time to ship.
The technically interesting part was not the pipeline. It was that each team had a different definition of the same KPI. Two dashboards showing different numbers for "orders shipped today" is not a data problem; it is an agreement problem wearing a data problem's clothes.
So the engagement started with collaborating with the business to define the KPIs — over forty of them — before building anything. Then the pipeline acquired the data and produced them consistently.
This is the part teams skip under time pressure, and it is the part that determines whether anyone trusts the output.
A feature pipeline that could use all the data
The propensity work is the clearest before-and-after.
The legacy implementation calculated propensity using descriptive statistics. It was slow, it could not keep pace with growing data, and — the number that matters — it ran against roughly five percent of the available data. Everything else was thrown away, not by choice but because processing it was not feasible.
We built an optimised feature engineering pipeline tuned to support 3,700+ features. The point was not the feature count for its own sake. It was that the constraint moved from what can we afford to compute to what is actually predictive — which is where a data science team should be spending its judgement.
A rule engine for offers
The final strand: dynamic offers based on individual behaviour rather than a single discount applied to everyone. The business wanted margin optimised at the level of micro-segments, or individuals, rather than uniformly.
We implemented a rule engine matching a customer's product views against defined business rules in real time — with the rules owned by the business and changeable without a deployment. That last property is what makes such a system survive contact with a marketing team.
The thread running through all of it
Every one of these had the same underlying shape: latency turning useful data into useless data. Fulfilment figures a day late cannot prevent a late delivery. A propensity model on five percent of data is a guess with extra steps. A recommendation for a customer you identified after they left is nothing at all.
The other recurring lesson is that the hardest parts were not the streaming infrastructure. They were the definitional questions — what counts as the same customer, what "shipped today" means, which of 3,700 features earn their place. Streaming makes it possible to act in the moment. It does not tell you what the right action is, and no amount of engineering substitutes for agreeing that first.
On numbers
We have deliberately not quoted revenue impact for this engagement. Figures appear in our internal material that we cannot independently substantiate, and an unsourceable number in a case study is worse than no number — it invites a question we cannot answer. The specifics above (forty-plus KPIs, 3,700+ features, five percent of data before the rebuild) are drawn directly from the project's own documentation.
We do not name clients. Engagements are described by sector and scale because confidentiality obligations outlast the work, and consent we cannot produce is consent we do not have.