
At Quirk’s New York this July, almost every vendor on the show floor said some version of “our data is clean.” Not many of the 1,949 people walking that floor had a way to check the claim against anyone else’s.
Survey methodologists have long treated data quality as coming from multiple independent sources, not one. The academic term is Total Survey Error. What follows is a simpler, three-part practitioner version of that same idea:
- Design: does the questionnaire ask what it needs to ask, in a logical order, without misleading people toward a particular answer.
- Build: does the programmed script match the spec, question by question, wave by wave.
- Field: are the people answering real, paying attention, and free of fraud.
Get the first two wrong, and the third one absorbs the cost, in cleanup that happens later.
This year, that third claim finally got a public standard. It’s the one buyers hear about most and can verify least. The GDQ Data Quality Excellence Pledge spells out what a vendor has to prove. More than 250 organizations have signed it in the open. Next to it sits a public benchmark report. It’s built from real removal-rate data across 46 companies, 13 countries, and roughly 1.8 million survey records.
What The Pledge Requires
The Global Data Quality initiative, GDQ, is a coalition of the industry’s own standards bodies. Members include ESOMAR, the Insights Association, MRS, SampleCon, CRIC, and several national associations besides. Its Excellence Pledge asks a signing company for five things:
- Documented pre-survey identity checks on every project.
- Disclosure of sample type and quality metrics, not just a summary claim.
- Compliance with GDPR and local privacy law in every market surveyed.
- Staff and subcontractor training on quality practices.
- Active membership in at least one GDQ partner association.
Sign it, and the company’s logo lands on a public page. Anyone can check it before an RFP goes out.
Kalever, one of the two companies behind CodexMR, is on that list alongside Ipsos, Kantar, Dynata, NielsenIQ, and Cint. A signature puts the company’s data quality practices in front of a standard everyone can read. That’s different from a claim in a slide deck.
Worth being precise about what that signature covers. The Pledge’s five commitments sit inside the field claim only. They cover who’s answering, what the sample looks like, and whether data collection followed privacy law. None of that touches how a team designed the questionnaire or built the script. Those are earlier, different jobs, and no public pledge covers them yet.
What The Data Quality Numbers Show
The GDQ benchmark report’s Wave 2 is the fullest read yet. Globally, research agencies pulled 28.5% of records for quality issues, combining pre-survey and in-survey removal. Sample suppliers pulled 21.2%. That reverses Wave 1, when suppliers carried the higher rate. Fraud detection drives most of the gap: agencies caught 11.6% of records as fraudulent mid-survey, suppliers caught just 2.9%.
B2B remains the highest-risk audience by a wide margin. Combined removal for general B2B sits at 22.5% globally. After the survey closes, post-survey review still pulls 59.6% of the records that made it through, more than half. Incidence-rate gaps have narrowed since Wave 1 but haven’t closed: agencies sold 62.0% incidence and delivered 59.1%; suppliers sold 53.6% and delivered 45.9%.
The report doesn’t split that post-survey number by cause. That’s fair, it wasn’t built to. Fraud accounts for some of it. In our own production work, a real share of what’s caught this late traces back to something catchable much earlier. A skip pattern let the wrong respondents through. Someone set a quota once and nobody rechecked it. A translated question quietly changed meaning. None of that is fraud, but it’s still bad data. It still costs the same fix, just later and more expensively than an earlier catch would have.
Country-level numbers move around more than one global figure can show:
| Market | Research Agency Removal | Supplier Removal |
|---|---|---|
| United States | 24.6% | 18.3% |
| United Kingdom | 11.2% | 18.6% |
| Germany | 40.7% | 10.1% |
| Australia | 65.7% | 62.9% |
(Source for table: GDQ Data Quality Benchmarking Report, Wave 2, H1 2026)
What This Looked Like On The Show Floor
At Quirk’s New York, the AI conversation had visibly moved past adoption. NORC’s AmeriSpeak team ran a session on generating trustworthy insight. Survey fraud is getting harder to catch, and AI-written answers are now showing up in both quant and qualitative work. It’s the same problem the benchmark puts a number on. Bad data doesn’t announce itself anymore. Someone has to measure it, wave over wave, in public.
Where This Leaves You Before The Next RFP
Three checks are worth adding to a vendor evaluation now.
- First, is the vendor on the GDQ Pledge signee list? Does their quote match a public commitment?
- Second, ask where their removal rates land against the Wave 2 benchmark for the audience type you’re buying. “Clean data” means something different for general B2C than it does for B2B or healthcare.
- Third, and this is the one the Pledge doesn’t ask: what does the vendor check before the survey ever fields?
How CodexMR Approaches Data Quality
The public standard stops at the field claim. The design and build claims sit outside it. Together, they decide how much fraud-shaped cleanup shows up later.
It’s close to what we check before any of our own studies go to field.
As part of the workflow we have two powerful validation tools:
- ResearchReady reviews the questionnaire itself, across seven specialist areas including a dedicated data quality pass. It runs before a single line of code exists, when a fix is still just a document edit.
- QAReady checks the built script against that source questionnaire. It runs continuously, as fieldwork moves and waves repeat, not once before launch.
A tracker that drifts between waves is exactly what a single pre-launch check misses. It’s also what the benchmark’s post-survey cleanout numbers pick up later, once the fix is no longer cheap. Neither tool claims to catch fraud. That’s a different job, done by whoever’s running the fieldwork. It’s why Kalever’s name is on the Pledge, not the platform’s.
What The Numbers Don’t Prove
The three-claim split above, design, build, field, is how we read where data quality problems come from in our work. It isn’t a category the benchmark itself measures. The report doesn’t attribute its removal numbers to design, build, or field.
The benchmark is voluntary and self-reported, so it’s a data quality signal, not a data quality audit.
Forty-six companies submitted data, but thousands of research businesses didn’t. Coverage still skews toward large contributors in the US, UK, and Canada. Several country-level bases outside those three are too small to read as more than directional. The report says as much plainly about its own Canada research-agency figures. A vendor’s absence from the Pledge or the benchmark isn’t proof of bad practice. It’s one fewer thing you can check before you ask them directly.
The Pledge and the scoreboard won’t make fraud disappear from the sample pool. Neither will catch the next AI-written open-end on its own. That turns “trust us, our data is clean” into a public list and a public number you can check. Check it before the contract, not after the fieldwork report comes back wrong.



