Related
Date
10/2/2026
10 minute read
Share
Every enterprise security architecture is built on the assumption that the person participating in a digital interaction is who they claim to be, but AI-powered impersonations are breaking that assumption.
Whether it’s a job interview, a wire transfer approval, a credential reset, a board meeting, or a customer support call, organizations can no longer rely on what they see and hear on audio/video calls.
And while some people may imagine a deepfake being an obviously manipulated video or poorly executed face swap, enterprise attacks rarely look like that. The most effective AI-powered deceptions are designed to infiltrate everyday business interactions and bypass existing security controls.
A recent hiring interview between Ilya Kroogman, founder of creative agency The Digital Panda, and what appeared to be a highly qualified software engineering candidate illustrates just how subtle these attacks have become.
On the surface, there was nothing particularly unusual. The resumé looked legitimate and the LinkedIn profile was well developed. The only detail that gave Kroogman pause was the candidate’s accent, which didn’t match the personal background he described during the interview.
Only later did suspicions emerge that the individual on the call may have been concealing their real identity using AI-powered face-swapping technology. Cases like this are becoming increasingly common as organizations face attempts by North Korean IT operatives and other threat actors to infiltrate companies through remote hiring.
One of the first questions we’re asked at GetReal Security is whether there was something obvious to the eyes that should have revealed the deception. In most cases, the answer is no.
Interviewers are evaluating a candidate’s experience, asking technical questions, taking notes, responding to answers, and often looking at their own screen rather than carefully inspecting another person’s face frame by frame. If something looks slightly unusual, the natural assumption is almost always poor lighting, limited bandwidth, or video-app glitch — not an AI-generated identity.
Recruiters are trained to remove bias from hiring decisions, not to make subjective judgments about whether someone looks or sounds “real.” Expecting them to determine whether a live participant is using AI to impersonate someone else or isn’t who they claim to be places them in an impossible position. And existing identity verification and background screening processes typically occur later in the hiring workflow and verify credentials and identity, not what’s happening onscreen.
This is why detecting AI-powered impersonation, face swaps or determining authenticity is both a security problem and a scientific problem rather than a visual one.
Why is black-box detection not enough for live video calls?
Many people assume that building a deepfake detector only involves collecting examples of authentic and manipulated videos, training a machine-learning model, and measuring its performance. While that may produce impressive benchmark results, it often struggles to encapsulate the complexities of real-world communications and so they tend to break when deployed in the real world.
Our research begins with a different question. Rather than asking whether a video “looks fake,” we ask what measurable evidence an AI-powered impersonation attack inevitably leaves behind, and how those signals differ from the ordinary transformations introduced by cameras, microphones, compression algorithms, networks, and video-conferencing platforms.
This philosophy shapes how GetReal’s Research Lab operates. Rather than treating deepfake detection as solely a real/fake classification problem, we study the adversary, the communication environment, and the interaction between the two.
Why start with the adversary?
Before you can defend against an attack, you need to understand how the attack works.
That means studying the software used by attackers, whether it is commercially available face-swapping applications, open-source voice-cloning models, or tools developed by sophisticated threat actors. We want to understand not only what these systems produce, but how they produce it.
Once we understand the mechanics of the attack, patterns begin to emerge.
Take face swapping as an example. Regardless of which software is used, every real-time face swap has to accomplish the same basic task. It must continuously transform one person’s face so that it matches another person’s head pose, facial expressions, and movements. That process requires image warping (the digital equivalent of how a rubber mask conforms to the face). Although different tools implement it in different ways, they all perform fundamentally similar operations.
That observation is important because it changes the problem entirely. Instead of asking whether a particular video was created by Tool A or Tool B, we ask what evidence those underlying operations inevitably leave behind. If we can identify signals that arise from the process itself rather than from a particular application, then our detection methods become more resistant to changes in individual tools.
Why does the communication pipeline matter?
Understanding the attacker is only half of the problem. The other half is understanding everything that happens to a video before it reaches the participant.
Enterprise video conference calls are complex and variable. The same conversation may take place on Zoom, Teams, Webex, or Google Meet. Participants may connect from laptops, mobile phones, or conference rooms using entirely different cameras and microphones. Network conditions fluctuate continuously throughout a meeting, changing compression rates, resolution, frame timing, and audio quality in real time.
All of those variables alter the underlying media.
If you don't understand how video conferencing systems naturally affect audio and video, you risk mistaking completely benign behavior for evidence of manipulation. A temporary bandwidth reduction, motion blur, or aggressive video compression can all resemble anomalies unless you understand the communication pipeline well enough to distinguish normal variation from adversarial behavior.
For that reason, a significant part of our research has nothing to do with AI at all.
We characterize how real communication systems behave under different conditions, build mathematical models describing those behaviors, and validate them across different devices, networks, lighting environments, and conferencing platforms. Only after understanding those natural transformations can we begin asking whether something remains that is better explained by AI manipulation.
Why does explainability matter?
Understanding the underlying signal also explains why explainability matters so much in security.
When a detection system produces a false positive or misses a real attack, we need to understand why. By understanding which signals we are measuring, we can determine whether a failure resulted from motion blur, an unexpected camera characteristic, a new attack technique, or some other factor. This allows us to improve the system in a principled way rather than simply retraining another model and hoping for better results.
Why isn’t one signal enough?
Modern identity attacks manipulate far more than facial appearance.
They involve voices, background consistency, environmental cues, timing, device characteristics, and behavioral patterns that emerge over the course of an interaction. Looking at only the face — or only the voice — creates blind spots.
Robust detection requires combining multiple independent sources of evidence while continuously evaluating them throughout the duration of the conversation.
Although detecting synthetic media during a live call is challenging, defenders possess one advantage that attackers do not.
An attacker must generate convincing synthetic media continuously, frame after frame, while maintaining a natural conversation in real time.
A defender is not limited to evaluating a single image. We can observe the interaction as it unfolds, accumulating evidence over seconds or minutes. Behavioral repetitions, temporal inconsistencies, environmental anomalies, and subtle correlations become much easier to identify over time than they are within any individual frame.
What does this mean for enterprise security?
The interview with Kroogman illustrates something much larger than a single attempted impersonation. It demonstrates that enterprise security can no longer assume that the person appearing on a video call is necessarily the person they claim to be. As these attacks continue to improve, successful detection will depend less on recognizing today’s deepfake tools and more on understanding the scientific principles that all of them inevitably share.
The tools will continue to evolve. The underlying physics, mathematics, and computational constraints that govern how they operate evolve much more slowly. Organizations that anchor their defenses in those enduring scientific principles will be better positioned to adapt as AI-powered impersonation continues to evolve.
See How GetReal Stops AI Impersonation