UX & AI

AI Moderated Interviewing

A new NN/g study found four ways it AI moderated Interviewing falls short but it seems .

AI Moderated Interviewing featured image
————
TLDR

AI is starting to run interviews. Not just transcribe or  analyze them, but actually run them. You hand it a research goal or an interview guide, and it asks the questions, listens to the answers, decides what to dig into, and follows up. 

That is worth sitting with for a second. An interview is one of the most personal things a person can engage in. You’re not just collecting information. You’re helping someone put their own story into words, and whether they say the true thing or the safe thing depends on whether they trust you. That trust gets built in small ways. It can look like a pause you leave open instead of filling, a follow-up that shows you were listening, a summary that clarifies. Interviewing is a craft. It has real skill in it. And the skill has always worked because a person with good judgment was moderating the process.

So what happens when you hand that skill over completely to something that isn’t a person at all? Nielsen Norman Group went and tested that idea with two AI-interviewer tools currently available for use, Marvin and UserFlix, in front of ten real participants. It’s not hypothetical anymore.

Three problems that can likely be addressed.

The thing that surprised me most was that summarizing what people said, in their own words, made them feel heard more than anything else the AI did. That’s a real active listening tactic called mirroring, and it’s usually a difficult human aspect of interviewing, so the machine was good at something you’d expect it to be bad at.

AI wasn’t flawless though. Neither tool paused to ask if the summary of the participant’s words was right. One person said she cut the AI off just to correct it. I’ve worked with AI long enough to confidently say that is not a limitation. AI can absolutely say “What I’m hearing is…” and then ask “Did I get that right, or is there more?” What is likely missing is a line in the system prompt.

The article also noted that the pacing of the interview was a bit rough. Interview length ran anywhere between 13 to 56 minutes. NN/g attributed this to AI’s inability to tell when a topic had been covered because it can’t read tone and body language. That’s fair, but it doesn’t need to read the room well in order to cover topics adequately. For example, if you give AI a list of objectives, let it track what topics are covered, and have it keep pulling a thread only while something’s still unanswered. And yes, I realize a non-answer is data. When someone dodges or can’t find the words, that gap is worth noticing. But saying something like “Anything else, or should we move on?” hands pacing to the participant instead of asking the model to guess at cues it can’t see.

The AI can also implement gentle probing techniques. AI can mostly detect if an answer is vague or closed based off a transcription. If that’s the case, it can rephrase the question from a different angle. If the participant deflects again, it accepts that, notes how they declined, and offers to move on. Honoring a refusal belongs in the prompt like everything else here.

The study also found the AI’s praise felt fake. Responses were so over-the-top that people stopped trusting it. That’s the easiest fix on the list. Ask AI to dial it back in the prompt. Self-disclosure also handles that more cleanly. A “tell” only works if something’s hidden. AI should begin interviews with opening statements that disclose they are talking to an AI and that some comments may feel fake. With that I would anticipate that the fake-warmth problem mostly dissolves. Go further and have AI frame itself as a thinking partner, not a person — you can’t get caught pretending to be human if you never pretended to be in the first place.

“None of these issues requires the AI to get smarter. They needed someone to write down what a good interviewer already knows to do.”

There was another finding that doesn’t seem to be addressed as cleanly. A couple of participants held back sensitive information because it wasn’t able to build trust with the person behind the technology. One wouldn’t say where she worked (even though it’s on her LinkedIn profile) because she didn’t feel comfortable telling the AI. She didn’t know how it would use the information, and pointed out that with a person, someone is accountable for what happens to it.

Part of that is fixable, and their study says so, AI interviews should start with a proper introduction, and a word about how information will be used. One participant stating it herself. Have the AI say it’s an AI, say who it’s working for, say the conversation is confidential. That’s mostly a prompt modification, same as the previous three concerns.

There is a second issue underneath that one that disclosure can’t fully resolve.  One participant stated it nicely “With a person, someone is acountable.” A statement about confidentiality can make promises, but it can’t produce the human who’s actually accountable. The researcher behind the AI, the man behind the curtain, should take intentional steps to build rapport with the participant. It’s about trusting that someone on the other end can be held to account.

NN/g finishes their article by concluding that AI can handle structured interviews well but not semi-structured ones. I’d argue that their line is drawn too firmly. Much of what looks like on-the-fly judgment calls to a human is really a just unprompted confirmation checks, pacing handoffs, gentle probes, and missing guardrails that hand control back to a participant.

I’d also push back on the test setup. The tools evaluated were purpose-built for structured interviews and then judged for lacking semi-structured skills, which it’s a bit like testing a low-clearance vehicle and concluding cars can’t go off roading. A fairer test would use a deeper agent or a multi-agent setup that checks answers for completeness, generates its own follow-ups, builds notes on the participant, and runs on a current frontier model.

That said, for now, my recommendation is to use AI for structured, well-scoped interviews that benefit from scale and consistency. Unless you’re building your own tool, the available options will struggle with the judgment semi-structured work demands. More research would give practitioners real guidance on how to close the gaps flagged above.

TAGGED:
————  Work with me ————

Your work matters. Let's tell that story.

Leave a Reply

Your email address will not be published. Required fields are marked *