RedFlag

How to analyze a podcast guest: body language, voice, and measured signals

← All articles · By the RedFlag team · September 24, 2026

Long-form video podcasts are, quietly, the best behavioral footage on the internet. A guest sits in one chair for two hours under steady studio light, lit and miked better than most television, answering questions they mostly did not script. That is exactly what a baseline needs. It is also a format with three traps that most "body language breakdown" videos fall straight into: the two-shot, the shared audio track, and the three-hour drift. This is how to handle each of them and read what's left honestly.

RedFlag is not a lie detector. Nothing on this page can tell you whether a guest is telling the truth. It measures change in the guest's own voice and face over the course of a video. Whether a change means anything is your judgment — made against what was actually said, and what is actually on the record.

Why podcasts are good material

Trap 1: the two-shot

Most video podcasts cut between a wide shot of host and guest together and individual close-ups. The wide shot is useless for scoring a single person, and RedFlag treats it that way on purpose: a stretch is left unscored when a second face is at least 55% of the subject's size and at least 13% of the frame height — the app labels it "multiple people — no single subject" rather than guessing which of the two it is reading. That is the correct behavior, not a failure. Your usable material is the guest's own close-up.

The face-size gates matter here too. Analysis runs when the face fills at least 10% of frame height, but the blink channel only records data at 22% or more. A tight podcast close-up clears both easily; a mid-shot of a guest leaning back in a lounge chair may only clear the first, in which case the read rides on voice alone and the app says so. The thresholds are published in full on the methodology page.

Trap 2: whose voice is it?

This is the one nearly everyone misses. A podcast has one mixed audio track, and RedFlag does not separate speakers in the audio. When the editor holds the guest's close-up while the host asks a long question — a very common cut, because reaction shots are good television — the pitch and voiced-continuity measurements in that stretch describe the host's voice, laid over the guest's face.

The rule is simple: only read a flag where the guest is the one talking. Rewatch every flagged stretch and check. Cross-talk, laughter over each other, and the host's "mm-hm, right" under a guest's answer all feed the same track, so a flag that lands on a lively back-and-forth is usually telling you about the conversation, not the guest.

Trap 3: the three-hour drift

People change over a long recording for boring reasons. Voices tire and drop in pitch; eyes dry out under studio lights and blink more; the coffee arrives, or wears off. A baseline built over three hours averages the fresh guest and the exhausted one together. For anything longer than about an hour, pick a coherent segment of 10–30 minutes that contains both comfortable and pointed questions, and analyze that. (It also fits the free tier: 1 video a day up to 10 minutes with no signup, 2 a day up to 30 minutes once you sign in.)

The method, step by step

  1. Use the full episode, not the viral clip. Clip channels cut to the most dramatic moment and remove the baseline around it. Find the timestamp in the full upload.
  2. Read the transcript around the moment first. What exactly was the question? Was it the first time the topic came up, or the fifth? A question the guest has answered on thirty other shows produces a rehearsed answer, and rehearsal is the most common cause of a flat voice.
  3. Choose a segment that includes the easy part. The warm-up — the origin story, the book plug, the small talk — is the guest's comfortable register. You want it inside the same run, because RedFlag measures each signal against that speaker's own median across the video.
  4. Run the segment and note which stretches are unscored. Two-shots and silence are gated out; that tells you how much of the segment actually measured the guest.
  5. Rewatch each flag with three questions: Is the guest the one speaking? What was asked? Did anything non-behavioral change (a camera switch, a sip of water, a laugh)? Only a flag that survives all three is worth describing — and then only as "their delivery changed here," never as a verdict.

What podcast delivery does to each signal

ObservationThe ordinary podcast explanationWhat it is not
Flat, monotone stretchA story told on a dozen previous shows; a practiced pitch for their productEvidence of a lie — repetition flattens delivery for honest and dishonest speakers alike
Pitch flattening on a pointed questionChoosing words carefully on a sensitive, legal or personal topicA confession — careful phrasing is what anyone does when the stakes go up
Long unbroken monologueThe format rewards it; hosts often deliberately let guests runEvasion — on a podcast, talking for four minutes straight is normal
Blink rate above their own baselineStudio lights, a long session, looking between host and screenGuilt — and it only counts if the face is large enough to measure
Filler words spikeA question they didn't anticipate, answered liveFabrication — composing a true answer produces the same fillers

That middle column is where the research lands. DePaulo et al.'s 2003 review of over a hundred cue studies found no behavior that reliably marks deception, and Bond & DePaulo's 2006 meta-analysis put human truth-judgment accuracy at about 54% — you can check your own number in a few minutes with the 54% test. What behavior does track, per Vrij, Fisher & Blank (2017), is cognitive load: effortful recall and careful wording, which show up whether the statement is true or not. For the signals individually, see voice pitch under stress, blink rate and cognitive load in speech.

Publishing what you find, responsibly

Podcast guests are frequently not public officials — founders, authors, athletes, people with one viral moment. Hold the analysis to the narrowest honest form: this moment, this measurement, this question. Never write that a measurement means someone lied; it cannot mean that, and saying so about a real person is both wrong and potentially defamatory. The interesting finding is almost always which question moved the guest off their own baseline — and the answer to that is in the transcript, not in the tool. For a worked example of what a modest, record-anchored analysis looks like, see our Lance Armstrong 1999 case study, and for single-camera interviews in general, how to analyze an interview.

Analyze a podcast segment yourself. Paste a public YouTube URL into the RedFlag analyzer — free with no signup (1 video a day, up to 10 minutes; 2 a day and 30 minutes once you sign in), and the analysis runs in your browser. RedFlag is not a lie detector: it shows you where a guest's signals diverged from their own baseline, and leaves the judgment to you.

Open the analyzer

How the scoring works · How to analyze a political speech · Test your own judgment

Sources