Real-Time AI Voice Cloning Fuels New Wave of Vishing Attacks

by Oct 3, 2025ai, mfa, Phishing, security, software, spam, Technology, update0 comments

Cybersecurity researchers are sounding the alarm on a troubling advancement: artificial intelligence can now clone a person’s voice in real time, enabling scammers to carry out convincing voice phishing attacks that were nearly impossible until recently.

The breakthrough was revealed this week by NCC Group, whose researchers demonstrated how AI-driven voice cloning can be weaponized to trick employees into disclosing sensitive data or performing risky actions. In live simulations, attackers successfully impersonated trusted individuals to reset passwords, change email addresses, and request confidential information — all with a cloned voice responding on the fly.

From Novelty to Threat: How Real-Time Voice Cloning Works

For years, deepfake voice technology was limited to pre-recorded clips or text-to-speech (TTS) outputs, both of which were too rigid or too slow for live conversations. That changed when researchers routed an attacker’s microphone through a real-time AI voice modulator.

The process works like this:

  • The attacker speaks normally into their microphone.

  • A machine learning (ML) model instantly converts the audio into the cloned voice.

  • The output is routed into communication apps like Microsoft Teams or Google Meet.

  • On the victim’s side, it sounds as though they are speaking directly with a familiar colleague or supervisor.

To make matters worse, caller ID spoofing can pair the cloned voice with a trusted phone number, compounding the illusion of legitimacy.

Why It’s a Game Changer for Scammers

Traditional vishing has always relied on human impersonation skills. But with AI, fraudsters no longer need to sound convincing — the cloned voice does the work.

“Real-time voice cloning makes the scam more believable and increases the chances of success,” said Matthew Harris, senior product manager for fraud protection at Crane Authentication.

Experts note several reasons this development is so dangerous:

  • Dynamic response: Unlike prerecorded clips, AI voices can answer questions, escalate requests, and improvise naturally.

  • Authentic tone and cadence: The system adapts in real time, making conversations flow without the awkward pauses of earlier tech.

  • Scalability: The tools are now accessible, affordable, and don’t require elite technical skills.

“Real-time voice cloning is a force multiplier,” explained T. Frank Downs of BlueVoyant. “It sustains the illusion of authenticity throughout the call, dramatically increasing the success rate.”

From Proof of Concept to Proliferation

NCC Group stressed that the hardware and software used in its tests were not extraordinary — in fact, they were “good enough” off-the-shelf components. This means that the barrier to entry is low, and even small criminal groups could soon deploy these techniques.

That reality is troubling because victims often rely on three things to validate a call:

  1. The number displayed on caller ID.

  2. The voice they recognize.

  3. The content of the request.

All three can now be spoofed or cloned.

Brandon Kovacs, a senior consultant at Bishop Fox, noted that pairing real-time voice cloning with deepfake video could make scams on platforms like Zoom or Teams nearly undetectable. “Attackers can now handle questions and adjust their approach mid-call, just like a real human,” he said.

A Growing Market for Synthetic Voices

Real-time AI voice cloning is still emerging, but experts agree it is advancing rapidly. Generative AI has eliminated many of the flaws that once made synthetic voices sound robotic. As AI continues to improve its probabilistic pattern-matching, the gap between real and fake voices narrows further each month.

Roger Grimes, CISO advisor at KnowBe4, predicted that by 2026, most voice-based social engineering attacks will involve cloned voices rather than real humans. “Hacking via social engineering is getting ready to change forever,” he said.

This shift mirrors what enterprises are already experiencing with text-based scams. Alex Quilici, CEO of YouMail, explained that executives are often impersonated by SMS campaigns today. Voice deepfakes are expected to become the “next major attack vector” as tools spread.

What Organizations Should Do Now

Experts agree that technical defenses alone will not be enough. Instead, organizations should focus on identity as the new security perimeter.

Marc Maiffret, CTO of BeyondTrust, recommends several best practices:

  • Least privilege access: Limit what any employee can do, even if their credentials are compromised.

  • Identity monitoring: Watch for unusual access attempts or suspicious account changes.

  • Multi-factor authentication (MFA): Require more than just voice or credentials to authorize sensitive actions.

  • Employee training: Teach staff to verify unusual requests through alternate channels, such as direct callbacks.

Human vigilance, coupled with stronger identity controls, is currently the most effective way to blunt the impact of these attacks.

Looking Ahead: Audio and Video Together

While NCC Group’s demonstration focused on voice alone, its researchers are already exploring combined deepfake audio-video attacks. Synchronizing real-time speech with realistic facial animations is challenging, but they warn it’s likely only a matter of time before that barrier falls.

The implications are unsettling: business leaders may soon be confronted with fake video calls that look and sound exactly like their colleagues or executives.

Conclusion

Real-time AI voice cloning is no longer science fiction. It is a practical, accessible tool that cybercriminals can use today to launch highly convincing vishing attacks. By making impersonation more scalable, adaptive, and realistic than ever, it represents a fundamental shift in social engineering tactics.

Organizations must move quickly to update their defenses, train their staff, and prepare for a world where hearing a familiar voice on the other end of the line no longer guarantees trust.

PTSI Editorial Team

Support Line: Phone: +1 646-535-HELP (4357) Email: helpdesk@progressny.com Support web: helpdesk.progressny.com