Two QA Layers, Not One

QA Layers

CodexMR’s QAReady checks a built survey against its source questionnaire: every question needs to exist, with the right type and order. Wording, routing, and quotas need to match. The code needs to run without errors. You activate it yourself, on any build, on any wave of a tracker. That’s the automated QA layer. On its own, it works great, but it’s not the whole job.

Here’s the reality: no single QA layer catches everything survey QA needs to catch. A vendor’s “AI-powered QA” pitch usually describes just this first layer, the automated one. What it can’t do is decide whether what it flagged, or didn’t flag, matches what the research was supposed to measure. That’s the second layer, a human one, and a vendor running only the first is exposed on whatever the second would have caught.

What the Automated QA Does Well

In general, the automated layer is good at repetition. Qualtrics’ ExpertReview for instance checks a survey against expected patterns. QAReady applies that same logic directly against the source questionnaire, question by question, whether that’s the first build or the fourth round after a client comment reopened three sections at once.

But what it can’t do is judge intent. An automated check can confirm that question 14 appears exactly when question 3 equals yes. It cannot tell you whether question 14 should appear there at all. That’s a specification question, not a logic question, and a script has no way to know what the research was trying to measure.

Software engineering has had a name for that gap since 1984. Barry Boehm’s original distinction splits testing into verification, confirming a system matches its own specification, and validation, confirming the specification itself solves the right problem. An automated QA layer verifies. It cannot validate, because validation means checking the build against something the code has no access to: what the research was trying to measure. 

The MRS’s 2026 Data Analysis and AI Toolkit makes the same point about survey QA in particular. It positions AI as useful for the proofing and logic-routing work, but treats human oversight as still critical on everything AI touches, not optional.

Why You Absolutely Need an Automated QA Layer in 2026

Two reasons: speed and quality. Skipping the automated QA layer means betting that nothing will slip between builds, waves, or markets, and that bet gets riskier every year. 

  1. It runs the same checks every time, with no fatigue. QAReady’s seven checks run identically on build one and build forty.
  2. It’s available on your call, not someone else’s. No queue, no wait, no chain of emails sitting on somebody else’s calendar.
  3. It compares wave to wave on trackers. On a tracker or copy-over study, it checks copy to copy and questionnaire file to survey code. It flags what changed before the next wave reaches field, the kind of drift a person re-reading a familiar document tends to skim past.
  4. It frees people for the calls that need judgment, instead of burning that time re-checking questions.
  5. Programs keep getting bigger. The market research services industry is projected to reach $116 billion by 2030, and multi-country, multi-wave tracker work makes up a growing share of that. That’s exactly the kind of program where a missed check compounds across markets instead of staying contained to one.

What the Human Layer Catches That the Checklist Can’t

This is where Greenbook’s 2026 GRIT Insights Practice Report is worth a close look. It splits AI adoption by function: 73% of analytics professionals use AI to prep and organize data, versus 43% of researchers, the people designing and running studies. The gap is where judgment lives. Prepping data is repeatable. Designing a study, and deciding whether a programmed survey still reflects that design, is not.

The strongest evidence for what happens when a person stays in the loop comes from outside market research entirely. A Stanford and MIT study published in the Quarterly Journal of Economics gave customer support agents access to a generative AI assistant. Issues resolved per hour rose 14% on average, and 34% for newer, less experienced agents. Still, the AI didn’t replace the agents. It suggested responses. The agent still decided what to send. The productivity gain came from that combination, not from removing the person.

What Each QA Layer Catches, and What It Misses

Layer Catches Misses
Automated Consistency against a fixed spec, every time, at any scale Whether the spec itself reflects the research intent
Human Whether programmed behavior matches research intent, not just logic Consistency at scale, especially under deadline pressure

So neither layer substitutes for the other. An automated layer without a human layer ships a script that’s internally consistent and still wrong. A human layer without an automated layer is accurate but doesn’t scale past a handful of builds a week.

Buyers Are Sceptical. And Humans Are Very Much Still in the Game.

Look at ESOMAR’s 20 Questions to Help Buyers of AI-Based Services. Human oversight gets its own section, separate from the questions about what the technology can do. Buyers are being encouraged to ask a fairly basic question: who is checking the AI’s work, and what happens when it gets something wrong?

And that question goes well beyond the research industry. According to TrustRadius’s 2026 B2B Buying Disconnect Report, vendor marketing materials came last among the resources buyers consulted. Meanwhile, the share of buyers who always or very often verify AI-generated information rose from 58% to 72% in a year.

Buyers are using AI, but they are not taking everything it produces at face value. So when a vendor answers “How do you QA this?” with “Our AI checks it,” there is still a fairly important question left unanswered: Who checks the AI? 

How CodexMR Runs Both Layers

This is the split the platform is built around, not bolted onto afterward. QAReady runs the automated layer: seven checks against the source questionnaire, activated on your call, never skipped because a deadline is close. But the human layer runs differently depending on how much of it you want to keep in-house.

So your own team can run both layers in DIY mode, with QAReady handling the repetition. Move to DIT, and CodexMR’s team works alongside yours on the judgment calls, as part of the collaboration rather than a separate line item. In DIFM, that human layer runs on CodexMR’s own specialist services: survey programming that sets the logic up correctly the first time, and data validation that checks it against the questionnaire before anything moves to field, with your team keeping sign-off at every stage. Same two layers, same standard, but different hands on the human one.

The stakes on that judgment call rise with the complexity of the study. On a conjoint analysis project, the human layer does more than confirm that Q14 fires correctly. It decides whether the attribute levels and choice tasks still model the trade-offs the research was designed to test, the kind of call worth specialist time when the output feeds a pricing or product decision.

What This Doesn’t Solve

Two QA layers still only cover design and build. Neither one checks whether the people answering the survey are real, attentive, and fraud-free. That’s field QA, and it’s a different job entirely. A vendor who nails both layers here can still field a study that collects fraudulent responses. 

Want to see how QAReady and CodexMR’s DIT and DIFM teams handle both layers on a real survey build? Get in touch. We’d be happy to walk you through the process and answer your questions.