
The three terms get used interchangeably, and that costs security teams real money. They are not variations on the same attack. They run on different channels, exploit different reflexes, and fail against different controls. A workforce trained to spot a suspicious email will still approve a wire transfer for a cloned voice, and will still act on instructions from a synthetic face on a Teams call.
Here is the boundary between the three, and what each one demands from your awareness program.
Key takeaways
- Phishing is written deception. Vishing is voice deception. A deepfake attack is synthetic identity, and it can ride on either channel.
- Vishing is now the second most common initial infection vector at 11% of intrusions, and phone-based simulations are clicked 40% more often than email ones.
- A deepfake video call is not "vishing with video". It defeats the exact control most companies use to verify a suspicious request: get the person on camera.
- Training on one channel leaves the other two untested. Attackers chain all three in a single kill chain.
What is the difference between phishing, vishing, and deepfake attacks?
Phishing is deception delivered in writing, usually email, where the attacker fakes a sender identity to get a click, a credential, or a payment. Vishing is the same deception delivered by voice over a phone call, where the attacker fakes a caller identity and applies live pressure. A deepfake attack uses AI-generated voice or video to fabricate a specific, recognisable person, most often one of your own executives, so the target is not just trusting a role but trusting a face and a voice they know. Phishing and vishing describe the channel. Deepfake describes what the attacker is faking.
| Phishing | Vishing | Deepfake attack | |
|---|---|---|---|
| Channel | Email, chat, collaboration tools | Phone call | Video call, and voice calls when the voice is cloned |
| What is faked | A sender address and a document or page | A caller ID and a persona | A specific real person's face and voice |
| Typical ask | Click a link, enter credentials, open an attachment | Read out a code, approve MFA, install remote access | Approve a payment, grant access, confirm a transaction |
| Why it works | Volume, plausibility, routine | Urgency and real-time pressure | Visual and vocal recognition of a known colleague |
| What breaks it | Verifying the sender and the URL before acting | Hanging up and calling back on a known number | Out-of-band confirmation on a channel the caller does not control |
What is phishing, and how has AI changed it?
Phishing is a written social engineering attack in which an attacker impersonates a trusted sender to make a target click a malicious link, surrender credentials, or authorise a payment. It remains the highest-volume vector by a wide margin, but the version arriving in inboxes today is not the version most awareness programs were built to teach.
The old detection cues are gone. Generative AI removed the spelling errors, the awkward phrasing and the generic salutations that made a phishing email identifiable at a glance. It also removed the effort barrier on personalisation: an attacker can now build a message around a target's real projects, real vendors and real reporting line using nothing but public sources.
Two shifts matter most for training:
- Conversational attacks. Instead of a single email with a payload, the attacker opens a harmless-looking thread, builds rapport across several exchanges, and only introduces the malicious ask once the target has already replied twice. Nothing in the first message is flaggable, which is exactly the point.
- Multi-stage lures with no link. Techniques like ClickFix ask the target to paste a command or run a "fix" themselves, so there is no attachment to scan and no URL to block. The user becomes the delivery mechanism.
If your phishing simulations still test a single-message credential harvest, you are measuring resilience against a format attackers have largely moved past. See our guide to choosing the right phishing simulation type for how to map simulation formats to current attacker behaviour.
What is vishing, and why is it growing faster than email phishing?
Vishing, short for voice phishing, is a social engineering attack delivered by phone, in which an attacker impersonates a vendor, an internal service such as IT or finance, or a senior executive, and pressures the target into surrendering credentials, approving an MFA prompt, or installing remote-access software. No malware is required. A convincing caller and a moment of pressure are enough.
The growth is not marginal. CrowdStrike's 2025 Threat Hunting Report found H1 2025 vishing volume had already exceeded all of 2024, a 442% year-on-year increase. Mandiant's M-Trends 2026 puts vishing as the second most common initial infection vector, behind only exploits, at 11% of intrusions. And Verizon's 2026 DBIR found phone-based phishing simulations are clicked 40% more often than email ones.
The reason is structural. Email gives the target time. A phone call does not. The techniques are old and still effective:
- Caller ID spoofing to display a bank, a supplier or an internal extension.
- Manufactured urgency, usually a compromised account, a failed payment or a security incident that requires action right now.
- Help desk impersonation, where the attacker calls IT posing as an employee and requests an MFA reset, or calls an employee posing as IT and requests remote access.
- OSINT-built pretexts drawn from LinkedIn, press releases and org charts, so the caller already knows the names and processes that make the story credible.
Speed compounds the problem. Modern vishing crews move from an answered call to attempted data exfiltration in minutes, not days.
What is a deepfake attack, and how is it different from vishing?
A deepfake attack uses AI-generated audio or video to impersonate a specific, identifiable person rather than a generic role. This is the boundary that matters: vishing fakes someone from IT, while a deepfake fakes your CFO. Vishing asks the target to trust an unfamiliar voice claiming authority. A deepfake asks the target to trust a voice or face they personally recognise, which removes the moment of doubt vishing training is designed to create.
Deepfakes are not a separate channel. They are a capability layered onto existing ones. Dropped onto a phone call, a cloned voice makes vishing harder to refuse. Dropped onto a video call, a synthetic face defeats the standard verification advice: if the request looks strange, get them on camera.
The economics collapsed fast. A usable voice clone can be built from around ten seconds of public audio, which turns every executive interview, conference talk and podcast appearance into raw material. Signicat recorded a 2137% increase in deepfake-enabled fraud attempts between 2022 and 2025, and Surfshark attributes 25% of documented deepfake fraud losses to executive impersonation, mostly CEO and CFO-style transfer requests.
How does a deepfake video call attack work?
A deepfake video call attack replaces the attacker's webcam feed with a real-time synthetic version of a person the target knows, then uses that borrowed credibility to push an urgent request through in a single meeting. The video call is now the primary surface, not an exotic edge case: research from Resemble AI found 53% of corporate fraud incidents in H1 2026 were video-only.
The kill chain is consistent:
- Reconnaissance. The attacker scrapes public video and audio of a target executive from LinkedIn, webinars and press appearances, then builds the persona.
- The invite. A phishing email or calendar invitation impersonating the executive and the meeting platform pulls the target into an "urgent" call.
- The call. The deepfaked executive appears on camera, speaks, and makes the ask. A manufactured glitch or an in-call chat message often carries the payload.
- The payload. The call drops, a system prompt appears standing in for a malware download or a credential-harvesting page, and the target resolves it themselves.
The most cited case remains the Hong Kong incident in which a finance employee at an engineering firm authorised roughly 25 million dollars in transfers after joining a video conference where the CFO and every other participant was synthetic.
The uncomfortable finding is that employees cannot be trained to spot these visually. Human accuracy at identifying a deepfake sits close to a coin flip, and lower when nobody is expecting one. Liveness tricks like asking the person to turn fully sideways or wave a hand across their face still break many real-time face swaps today, but the models improve week over week, so these are prompts to verify, not proof.
What holds up is procedural. No face and no voice should be able to authorise money or access without a second confirmation on a channel the requester did not choose. The point of a deepfake simulation is not to teach people to see the seams. It is to make sure they have already met one, safely, and that the reflex to verify fires before the pressure does.
Which vector should you train against first?
The question is a trap. Attackers do not pick one.
A real kill chain now runs across all three: a spear-phishing email sets up the pretext, a voice call closes it, and a video call defeats the verification step in between. Testing email alone tells you how your workforce performs against roughly one third of the attack surface, and gives you no visibility on the two channels growing fastest.

That is the practical test for any awareness program: if it still lives entirely in the left-hand column, it is measuring resilience against threats the workforce has already outgrown. Coverage should follow the attacker, not the tooling:
| Priority | What to test | Why |
|---|---|---|
| Baseline | Email phishing, including conversational and multi-stage lures | Highest volume, and the entry point for most chains |
| High | Vishing across the whole workforce, not just the C-suite | Fastest-growing initial vector, and clicked 40% more than email |
| Strategic | Deepfake video calls with executive impersonation | Defeats your verification control, and carries the largest single-incident losses |
One engine, one dashboard, one report covering voice, video and text, because that is what a real attacker's kill chain looks like. Explore phishing simulation, vishing and smishing simulation, and deepfake simulation.
-
Yes. Vishing is phishing delivered by voice. The social engineering logic is identical, impersonate a trusted party and manufacture urgency, but the channel changes the defence. Email gives a target time to check a sender address. A live call does not, which is why voice resilience has to be trained as a reflex rather than taught as a rule.
-
No. Vishing is defined by the channel, a phone call. A deepfake is defined by what the attacker fabricates, a specific person's voice or face using AI. A cloned voice on a phone call is deepfake-enabled vishing. A synthetic executive on a Teams call is a deepfake video call attack, which is a distinct vector because it defeats video verification, the control most organisations rely on to check a suspicious request.
-
They are the same attack across four delivery channels. Phishing is email, smishing is SMS, vishing is voice calls, and quishing uses QR codes to move a target from a physical or on-screen code to a malicious page. Attackers routinely combine them, which is why testing one channel in isolation understates real exposure.
-
Roughly ten seconds of clean public audio is enough for a usable clone, and a few minutes produces a very strong one. Conference talks, earnings calls, podcast appearances and webinar recordings are all viable source material, which makes voice cloning a standing exposure for any public-facing leadership team.
-
Not reliably. Human accuracy at identifying synthetic video is close to a coin flip even when people are told to look for it. Physical challenges such as asking the person to turn fully sideways or pass a hand in front of their face still break some real-time face swaps, but they degrade as models improve. The durable defence is procedural: out-of-band confirmation for any payment or access request, on a channel the caller did not choose.
-
Yes, because they test different things. Phishing simulations measure whether people scrutinise written requests. Deepfake simulations measure whether the verification reflex survives contact with a familiar face and voice under time pressure. Most organisations discover the second number is far worse than the first.