Deepfake-driven social engineering means your awareness training is already behind

Most awareness training still teaches people to spot phishing as it looked five years ago. The highest-consequence attacks now arrive as a phone call in a familiar voice, and that is far harder to train against.

Most security awareness training still teaches people to spot phishing the way it looked five years ago: suspicious links, urgent subject lines, poor grammar. That training is not wrong, it is incomplete, because the highest-consequence social engineering attacks happening right now do not arrive as an email at all. They arrive as a phone call in a familiar voice, or a video call with a familiar face.

Why voice and video changes the calculation

Text-based phishing asks a person to evaluate content: does this read as legitimate. Voice and video cloning asks a person to override their own senses: this sounds and looks like someone I know and trust, so the identity check I would normally do feels unnecessary. That is a fundamentally harder thing to train someone to resist, because it exploits a form of trust verification humans are not used to consciously performing.

What the realistic scenario looks like

It is rarely a full deepfake video call sustained for minutes. It is a short voice clip, a few seconds pulled from a public earnings call or conference talk, used in a brief urgent phone call asking for a wire transfer, a password reset, or an out-of-cycle access grant. Short and urgent is the pattern, because it minimizes the window for the clone to break down or for the target to think carefully.

What training needs to add, not replace

  • A specific out-of-band verification step for any financial or access-granting request, regardless of how the request arrives or how confident the recipient feels about the caller’s identity.
  • A shared verification phrase or callback protocol for high-risk requests, established in advance rather than improvised in the moment.
  • Explicit permission, stated by leadership, for staff to delay or refuse an urgent request from someone senior until it is verified. This is the piece most training programs skip, and it is often the actual blocker: people know they should verify, and do it anyway because they are afraid of the social cost of questioning an executive.

The realistic goal

You cannot train people to reliably detect a good voice or video clone by ear. The goal is not detection. It is making verification the automatic next step for a defined category of request, regardless of how convincing the ask sounds, so the attack has to beat a process instead of a person’s judgment in the moment.

Free Download

Know what an attacker sees before they do.

A practical exposure checklist covering the gaps that cause most breaches, plus what POPIA actually requires you to have in place.

Free PDF · No spam · Unsubscribe anytime

Send me the checklist

Leave a Reply

Your email address will not be published. Required fields are marked *