
Machine learning to the rescue of campaigns
Vincent Prou
· 5 min
The data lost to consent can largely be made up by machine learning. On GA4, our client tests show the missing data reconstructed to within 10% of reality.
Swipe left to read
Keep reading
The data lost to consent can largely be made up by machine learning. On GA4, our client tests show the missing data reconstructed to within 10% of reality.
All user data will fall in 2023 with the nFADP/GDPR coming into force, and with companies bringing themselves into line on cookie consent. The digital ecosystem sees that loss as a disaster, because advertisers lose the ability to analyse their campaigns’ ROI and can therefore no longer allocate advertising budget properly. Without landmarks and without those precious tools, we are practically back to the stone age.
There are, however, a number of ways round the problem — machine learning, which we will talk about today. That loss of data can largely be made up by deploying machine learning algorithms to fill the gaps, as we will see in this article with the case of GA4. Another route does without cookies entirely: measuring each channel’s impact through media mix modelling.
The observation: cookies growing scarce have meant lost data for several years, and it will deepen in 2023
Web tracking retreats from 2018, app tracking from iOS 14.5, in 2021
Tracking capability remaining on the web and in apps, EMEA region, from 2017 to 2023. Heights are indicative, with no unit; 2023 shows what was expected when the article was published, in May 2023. Hover, tap or step through year event with the keyboard to read what it changes.
With the keyboard: Tab, then the left and right arrows to move from one event to the next.
Source: the article’s original figure, heights read pixel by pixel, approximate values; texts from the article · Chart: bright.swiss

Regulation on data collection (nFADP) carries important consequences for tracking digital campaigns.
Transparency about collecting and using data has become crucial to holding the trust of web and mobile users, customers and prospects.
Beyond their wish for more transparency and honesty, it is above all the European General Data Protection Regulation (GDPR), in force since 2018, that is the main cause of the consent banners now everywhere on the internet, and of the consequences that follow.
Bringing a significant fall in measured traffic, that regulation has affected the performance and the dependability of the reports coming from acquisition channels and campaigns, leading inevitably to an imprecise, even degraded, analysis of return on investment.
Companies have since faced a sizeable challenge: acquiring dependable data while respecting users’ privacy choices. Server-side tracking answers part of that challenge: it makes the collection of what may lawfully be collected more dependable, notably in the face of ad blockers.
Since the whole digital ecosystem grew up around cookies, when a user refuses them, traditional tracking solutions fail. It is nonetheless possible to get round that, through measurement technologies that distinguish the behaviour of users who allow cookies from those who refuse them.
That is the example of GA4 and Google Ads we will explore today.
Machine learning, used on the incomplete data captured when users refuse Analytics cookies, reconstructs the missing data through behavioural models built on similar users who accepted cookies. It therefore models the missing data so that it matches reality as faithfully as possible, while respecting users’ preferences on cookies.
What makes modelling the missing data possible?
Machine learning, in GA4’s case, rests on two observations:
- We have a smaller but sufficient sample of users who accept cookies. Even with the measurement lost through that channel, we go on collecting data useful for analysing a typical user’s behaviour — frequency, recency, journey.
- We can still measure the events, since they are not personal data. We can still count, for instance, the number of sales, the number of basket additions, of prospects and so on; what matters most is making sure each event is unique, hence the importance of giving every event a unique ID, so as to avoid the bias of double counting.
What does GA4’s machine learning do?
Simplifying to the extreme, machine learning runs a succession of rules of three, so as to extrapolate the missing data by crossing the counting data with the attribution and behavioural data from the cookies we have consent for. The machine learning Google applies to GA4 and Google Ads data is of course far more complex, so as to secure accuracy and eliminate potential bias. GA4 and Google Ads are therefore able, through machine learning, to recompose the missing data in the matrix.
That is also why GA4’s data takes so long to appear in the free version: the algorithm needs enough data to recompose it. That delay is shorter in the paid GA4 360 version, for two reasons. Partly commercial ones. And because data volumes are often larger in accounts holding that licence, which means extrapolated data appears far sooner.
For that modelling to work, however, two conditions have to be met:
- Put Google’s consent mode in place
- Load the GA4 tag in every case, and pass the user’s consent to GA4 so that the tag behaves correctly
With Consent Mode, Google counts 59 of the 62 real conversions, against 50 without
Conversions from 1,000 clicks on a Google Ads ad, by cookie consent: observed, modelled by Google, and real. Google’s worked example; move the sliders to change its rates.
The journey of the 1,000 clicks
The conversions, close up
Source: Google, the Consent Mode example quoted by the article (Google Marketing Platform, “Conversion modeling through Consent Mode in Google Ads”, 2021); outside the example, recalculated by Bright keeping its proportions · Chart: bright.swiss

In future, web data will be a combination of measured and estimated data, and through machine learning we will be able to estimate it as precisely as possible, close to 100% accuracy.
The diagram below shows the user data observed, in blue. We can see that part of it is missing — the top right corner — the data of users refusing cookies. Through machine learning, we can model that missing data, in yellow and orange.
Source: Google Marketing Platform, “Conversion modeling through Consent Mode in Google Ads” – 2021.

But are those calculations right? Is there bias?
We ran tests across several clients, and we were able to verify that machine learning can reconstruct the share of data to within 10% of reality. We therefore advise putting an internal protocol in place to validate the truth of GA4’s data, and to check regularly that the gaps are not too large. If you are within 10-15%, you are at the market average. At Vilebrequin, Consent Mode, combined with enhanced conversions and the conversion APIs, recovered the lost data with a margin of error of around 10%.
Conclusion
With GA4, Google has shown that machine learning is viable and gives data close to reality while respecting users’ privacy — which lets performance go on being measured and campaigns go on being optimised effectively.
The media, tracking and analytics platforms that have not begun that transition by 2023 risk suffering the consequences. Several players have taken a considerable lead, so as not to suffer too much from the loss of user data. As a marketing lead, it matters to ask how the tools you use will handle the disappearance of certain cookies, and the growing scarcity of others. Google and Meta are well on the way to taking that turn, through a combination of collection points and machine learning. But are you, and your technology solutions, ready to take the 2023 turn?
Going further
- Tracking and data collection→8 publications
- Data science→6 publications
- Data protection→2 publications



The key skills bright can bring to you






