Paid search in litigation
Abstract scattered dot illustration representing Click Fraud and Invalid Traffic

What it turns onData plus discoveryThe account shows part of it; the rest comes from the other side.

Click Fraud and Invalid Traffic

Short answer
The platform's invalid-click credits carry part of it; the rest is in the traffic logs
Where it comes from
Invalid clicks column, billing credits, server and CDN logs, vendor detection reports
Who holds it
The advertiser, the platform, and whoever sent the traffic
Retention window
Google daily data 37 months (Google Ads Help, Aug 14, 2026); server logs often shorter
What will not work
A credit shows traffic was filtered, not which clicks, not why, and not by whom
Applies to
Search and display campaigns billed per click, and the logs on the advertiser's side

The platform will tell you how many clicks it filtered — never which ones, or why, or who sent them

What this dispute turns on, and why the account answers half of it

A click fraud claim has three separable questions in it. Was some of the traffic invalid — meaning traffic the platform itself judged illegitimate, or traffic behaving in ways no human user would. Was the advertiser charged for it. And who sent it, on purpose or otherwise. The account record answers the first two, partially and on the platform's terms. It does not answer the third.

Google and Microsoft both filter traffic they classify as invalid before billing, both expose a count of what they filtered, and both issue credits. Neither exposes a click-level ledger tying a filtered click to an address, a device or a source. The advertiser sees that traffic was rejected and sees the money returned, without seeing what was rejected — which is why this page carries the verdict data plus discovery.

The other half sits in the advertiser's own server and CDN logs — the only traffic evidence a party controls end to end — and in whatever comes out of the traffic source itself. Whether any of it supports a claim is a question for counsel; I am not an attorney. What follows is what the data does and does not show.

"Click fraud" is a lay term; the standard uses two others

The technical vocabulary belongs to the Media Rating Council, whose Invalid Traffic Detection and Filtration Standards Addendum was finalized June 25, 2020. It splits invalid traffic, shortened to IVT, into two categories describing how hard the traffic is to detect.

General Invalid Traffic is what routine, list-based and parameter-based filtration catches. The enumerated categories include known invalid data-center traffic, bots, spiders and crawlers, non-browser user agents, pre-fetch or pre-rendered traffic where the ads were not subsequently accessed, and sessions from devices that cannot display images, including headless browsers.

Sophisticated Invalid Traffic requires advanced analytics, multi-point corroboration and significant human intervention to identify. The enumerated examples include automated browsing from hijacked or infected devices, incentivized human activity and click farms, falsification of impression, viewability, click, referrer and conversion attribution signals, misrepresentation of domains or apps, bots masquerading as human users, adware and malware, and cookie stuffing.

Two things follow. These are detection-difficulty categories, not severity levels — traffic classified as general is not less costly, only easier to catch. And much of it is innocent: a double-click, a crawler, a pre-fetch. Calling all of it fraud imports an intent the measurement does not carry, which is the fastest way to have an opinion taken apart.

What the platform gives you: a count and a credit, not a ledger

Google states that where its systems determine clicks are invalid, it tries to filter them automatically from reports and payments so the advertiser is not charged, and describes a multi-layered approach to protecting advertisers from invalid traffic. Its enumerated categories are manual clicks meant to increase an advertiser's costs or a site owner's profits, clicks by automated clicking tools, robots or other deceptive software, and accidental clicks of no value such as the second click of a double-click (Google Ads Help, Invalid clicks, read August 14, 2026).

What the advertiser gets is a column: adding Invalid clicks to a campaign report returns the number of clicks automatically filtered. Microsoft's structure is similar: three quality tiers, only one of them billed.

Four things are absent from both. Why any specific click was ruled invalid. Which clicks were filtered. Any complete pre-filtration record, so the filter's precision cannot be audited from the advertiser's side. And any basis inside the platform record for reconciling its number to anyone else's. A credit is documented as an automatic consequence of filtration, not an adjudication that someone acted with intent — "they refunded it, so they admitted fraud" does not survive contact with either platform's documentation.

The disclosure paradox, named on the record and still unresolved

The structural problem in these disputes was identified by the court-appointed independent expert in the first major click fraud class action against Google. Dr. Alexander Tuzhilin, a professor of information systems, put it this way: an operational definition cannot be fully disclosed to the general public because of the concern that unethical users will take advantage of it; however, if it is not disclosed, advertisers cannot verify or even dispute why they have been charged for certain clicks.

Tuzhilin concluded that Google's efforts to combat click fraud were reasonable. That finding and the paradox sit together, and both still hold: the platform cannot publish the rule without handing it to the people it is defending against, and while it does not, no advertiser and no retained expert can check the arithmetic against the definition behind it.

The matter settled in 2006, and the settlement was reported at the time as valued around ninety million dollars and paid in advertising credits rather than cash — but that figure comes from trade press rather than a docket I have pulled, and the court, case number and structure would need verifying before any of it went in a report. What survives the settlement is the paradox.

Why a detection vendor's number never equals the platform's

Every one of these matters produces two numbers that do not agree: the platform's invalid-click count and a detection vendor's fraud percentage. They are not two measurements of one thing.

DimensionPlatform filtrationVendor detection
Where it observesThe platform's own serversThe advertiser's site
When it observesBefore billingAfter the click lands
Definition appliedPlatform's own, unpublishedVendor's own
Visible to the advertiserA count and a creditA classified click list

A fourth difference is not technical. A vendor selling fraud protection has a commercial interest in a high number. That is not an accusation and not a reason to discard the data; it is a foundation problem an opinion should address before opposing counsel raises it.

The double-counting risk is concrete. Google's own documentation lists filtered invalid clicks among the reasons its click count runs lower than a third-party tool's — so part of what a vendor measured never entered the billed click count at all. Subtracting one number from the other is not a damages model. The defensible use of a vendor's output is as evidence of traffic characteristics — addresses, user agents, timing, session behavior — reconciled against the advertiser's own logs, rather than as a loss figure adopted whole.

The advertiser's own logs are the one record controlled end to end

Web server and CDN logs are the only part of this record that does not belong to somebody else. They carry addresses, user agents, timestamps, referrers, request sequences and the session that followed each click. They are the only place an examiner can observe the traffic independently rather than read a platform's summary, and unlike everything else here they can be preserved by the client alone.

What log analysis supports is a description of behavior. Requests arriving at machine-regular intervals rather than in a human distribution. Sessions terminating at the landing page across a population of visits. Repeated requests from a narrow set of addresses or one user-agent string. Referrer chains that do not match the placement the click was billed against. Stated as observations, those hold up.

What log analysis does not support is identification. An address is not a person, a user-agent string is trivially forged, and consumer connections rotate addresses routinely. The honest formulation is that the traffic shows characteristics consistent with automation, or inconsistent with its stated source — and discovery, not analysis, is what connects it to a party. That is where I stop and counsel continues.

The clock on each half of the record

Invalid-traffic figures have no retention schedule of their own; they age out on the schedule of whatever report carries them. These are the published windows, read August 14, 2026.

RecordRetentionHeld by
Google Ads reporting, hourly to weekly37 monthsGoogle
Google Ads change history2 yearsGoogle
Microsoft performance, daily or coarser36 monthsMicrosoft
Microsoft change history6 monthsMicrosoft
Microsoft hourly search query detail1 monthMicrosoft
Server and CDN logsSet by the advertiserThe advertiser

The Google figure of 37 months took effect June 1, 2026; monthly and coarser aggregates are kept eleven years, which is why a matter reaching back four years can show what was spent in a month but not on a day — and day-level detail is what an invalid-traffic analysis usually needs.

The last row decides most cases and has no published number, because there is none to give. Log retention is whatever the hosting and CDN configuration says, often days or weeks — shorter than every platform window above it. So I ask for that setting in writing at the start. A hold on those logs is the cheapest time-sensitive step available, and nothing in any platform's public documentation suggests a preservation letter extends a platform's own schedule.

What the data does not establish

The list is longer than most parties expect, and naming it keeps an opinion inside what the record supports.

  • Which clicks were filtered, and why. The advertiser sees a count, never a click-level ledger, and no platform discloses its operational definition.
  • Intent. Invalid traffic is a measurement category that includes double-clicks, crawlers and pre-fetch. A finding that traffic was invalid is not a finding that anyone meant it.
  • Identity. Logs describe traffic, not people. Connecting traffic to a party is discovery's work.
  • The loss. A vendor percentage applied to total spend is not a measured figure, and it double-counts clicks the platform already filtered and never billed.

One more, regularly overstated in both directions. Google Ads holds accreditation from the Media Rating Council for clicks and invalid clicks as reported in the interface, with sophisticated invalid traffic filtration in scope; the letter is dated March 31, 2026. That means an independent auditor examined the measurement and filtration process against a published standard (MRC Invalid Traffic Detection and Filtration Standards Addendum, June 25, 2020, read August 14, 2026). It does not mean the numbers can be audited click by click. Accreditation is process assurance, not a ledger and not a right of inspection.

What is established, what is an estimate, and when this is not worth an expert

Industrial-scale invalid traffic is a proven fact pattern, not a theory, and the clearest proof of it is a criminal conviction. In the Methbot and 3ve prosecution, Aleksandr Zhukov was convicted of wire fraud, money laundering and conspiracy counts, and sentenced on November 10, 2021 to ten years in prison with $3,827,493 in forfeiture ordered. Between September 2014 and December 2016 he ran a sham advertising network of more than 2,000 rented commercial servers programmed to simulate humans viewing ads, spoofing more than 6,000 publisher domains (U.S. Department of Justice press release, read August 14, 2026). That is the answer when invalid traffic is framed as speculation.

What I will not put in a report is a circulating industry loss percentage. The most cited baseline that is not vendor marketing — a trade association's bot baseline series — is a 2019 report on 2018 data, co-authored by a fraud-detection company, measuring display and video impression fraud rather than search clicks. Transplanting its percentages onto a search campaign is a guess wearing a citation.

Counsel may find Pulaski & Middleman, LLC v. Google, Inc., 802 F.3d 979 (9th Cir. 2015), useful on whether a platform's own pricing methodology can support a class-wide restitution calculation. That assessment is counsel's, not mine.

And the case for not retaining anyone: if the logs are already gone and only the platform's filtered-click count survives, an examination adds little beyond what the account report states on its face. The same is true where the disputed spend is small against the cost of log analysis. In those situations the record will not carry the claim, and saying so is worth more than an engagement.

Frequently Asked Questions

Can click fraud be proven from Google Ads data alone?

Not on its own. Google reports a count of clicks its systems filtered as invalid and issues credits, but it never discloses which clicks were filtered, why, or where they came from, and no complete pre-filtration record is exposed to the advertiser. So the account shows that some quantity of traffic was rejected and that money was returned. Connecting that traffic to a source generally requires the advertiser's own server and CDN logs, and identifying a party requires discovery. The account record supports part of the claim and is silent on the rest.

What is the difference between click fraud and invalid traffic?

Click fraud is a lay term that implies intent. Invalid traffic is a measurement category defined by the Media Rating Council, and it includes plenty of traffic nobody intended maliciously — the second click of a double-click, crawlers, pre-fetched pages. The standard splits it into general invalid traffic, which routine list-based filtration catches, and sophisticated invalid traffic, which requires advanced analytics and human investigation. Those are detection-difficulty categories rather than severity levels. Using fraud where the evidence supports only invalid traffic overstates the finding and invites a challenge to the whole opinion.

Does a Google or Microsoft invalid-click credit mean the platform found fraud?

No, and both platforms' own documentation says otherwise. Google describes automatic filtration of invalid clicks from reports and payments before billing. Microsoft states that only standard-quality clicks are billed and that credits are issued automatically where invalid clicks are suspected. Each is documented as an automatic consequence of a filtration system, not as an adjudication of anyone's conduct or intent. A credit supports the statement that the platform's systems classified some traffic as invalid. It does not identify who sent that traffic, or establish that anyone acted deliberately.

How long before the click data in a click fraud case is gone?

Google Ads keeps hourly, daily and weekly reporting for 37 months as of June 1, 2026, and monthly and coarser aggregates for eleven years — but day-level detail is usually what an invalid-traffic analysis needs. Microsoft keeps daily performance data 36 months and hourly search query detail one month. The advertiser's own server and CDN logs have no published window at all; retention is whatever the hosting and CDN configuration says, often days or weeks. That last one is usually the shortest clock in the matter and the only one a client can extend.

Are third-party click fraud reports reliable evidence?

They are useful as evidence of traffic characteristics and unreliable as a loss figure. A vendor measures on the advertiser's site after the click lands, using signals the platform never sees; the platform filters on its own servers before billing, using signals the vendor never sees. The definitions differ, the populations differ, and part of what a vendor flagged may never have been billed at all. A vendor also has a commercial interest in a high number. The defensible approach is to reconcile the vendor's raw traffic data against the advertiser's own logs.

Can an expert identify who was clicking the ads?

Log analysis supports statements about behavior, not identity. It can show requests arriving at machine-regular intervals, sessions ending at the landing page across an entire population of visits, a narrow set of addresses or a single user-agent string repeating, or referrer chains inconsistent with the placement billed. It cannot name a person. Addresses are shared, rotated and trivially spoofed, and user-agent strings are set by whoever sends the request. The finding is that traffic shows characteristics consistent with automation or inconsistent with its stated source; connecting it to a party is discovery's work.

What should be preserved first in a click fraud matter?

Web server and CDN logs, immediately, and the retention setting confirmed in writing — those logs are typically the shortest-lived record in the matter and the only one entirely within the client's control. After that, the platform reporting that carries the invalid-click counts, at daily granularity while it still exists, plus billing records showing credits applied. If a detection vendor is involved, its underlying per-click data matters far more than its summary percentage. No platform's public documentation indicates that a preservation letter extends the platform's own retention schedule.
Keep reading

The guides run the sequence

A page here covers one dispute, or one kind of record. A guide covers the order the work happens in — what has to be exported before access is lost, and which analysis is worth paying for at all.

Top